Skip to content

Author

Nikita Kozodoi

We have 2 of 19 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Test-Time Augmentation for LLMs: When Input Diversity Beats Output Diversity at Matched Compute

Test-time scaling improves LLM accuracy but multiplies inference cost, making the accuracy gained per unit of compute the metric that matters in deployment. Self-consistency is one of the established approaches, which spends this budget entirely on the output side by sampling repeated reasoning paths. We study Test-Time Augmentation (TTA), which extends self-consistency by also perturbing the input, aggregating predictions across transformed versions of the input, and ask whether input-side diversity converts compute into accuracy more efficiently than output-side diversity. We perform a systematic, matched-compute comparison: we evaluate three simple input-side strategies (semantic rephrasing, lexical perturbations, and visual transformations) across six datasets covering general and multilingual knowledge, mathematical reasoning, multi-modal question answering, and sentiment classification, against chain-of-thought prompting and self-consistency. Semantic rephrasing delivers consistent and statistically significant accuracy gains while Pareto-dominating self-consistency on cost-effectiveness, delivering roughly 1.8X more accuracy per dollar and outperforming it on five of six tasks. We further analyze the number of augmentations, multi-modal strategies, and base model scaling, finding that TTA is most cost-effective for mid-tier models where a stronger model is unavailable or too expensive. Our findings indicate that for current mid-tier LLMs, varying the input converts inference compute into accuracy more efficiently than varying the reasoning path alone. The TTA implementation is available at https://github.com/aws-samples/sample-genai-reflection-for-bedrock.

Nikita Kozodoi, Zainab Afolabi, Jack Butler · 0 citations
Jul 2026

Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

A striking method-dependent pattern is revealed: simple averaging degrades sharply with overfitting, while sparsification-based methods achieve their best performance well past the validation optimum, suggesting that training duration and merging method should be chosen jointly rather than independently.

Nikita Kozodoi, Zainab Afolabi, Jack Butler · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.