Developmental Trajectories in a Sparse Language Model with 348M Active Parameters
Abstract
A LAMBADA accuracy difference discovered on 128 repeatedly inspected examples persists on the additional 5,025 examples of the same benchmark. At approximately four billion training positions, ConeML scores 30.51% versus 24.84% for Pythia-1B; at approximately ten billion positions, its later branch scores 39.92% versus 34.77%. The paired differences are +5.67 and +5.15 percentage points, with unadjusted bootstrap 95% intervals of [4.40, 6.97] and [3.82, 6.51]. Both pass the specified five-contrast Holm-adjusted testing procedure. Checkpoints and contrasts were selected after exploratory inspection, with the expanded evaluation specified before inference. A selected-training-inventory overlap screen finds four normalized full-passage matches; checkpoint-specific delivery is unresolved, and conservative common exclusions preserve the score ordering. This result is documented within one pretrained sparse-model lineage: six experts per layer with 820.23M total / 348.374M active parameters per token, followed by a nine-expert branch with 1,174.22M total / 348.466M active parameters. Both select two experts per token. A frozen five-task screen reveals an uneven profile: near the later LAMBADA comparison, both Pythia references lead on the other four tasks. Accuracy and likelihood select different early checkpoints. Python parseability and execution also separate: across nine-expert checkpoints at 75k, 78k and 80k, strict parses are 66, 57 and 96 out of 129, while successful executions are zero, one and zero. A shared-ancestor six-versus-nine-expert comparison at 75k shows mixed changes across public scores, parsing, binding and narrative retention. These findings establish benchmark-specific and non-monotonic developmental observations. They do not establish a causal curriculum benefit, matched-compute efficiency or robustness across training seeds. Training positions are tokenizer-dependent, and the proprietary training recipe is not independently reproducible from this report. Version 1.0 observation boundary: 6E-75k and 9E-80k. The expanded LAMBADA analysis retains its pre-specified 9E-78k comparison. The supplement includes text-free per-item records, pinned public-reference revisions, fixed benchmark configurations, paired statistics, overlap sensitivity, figures and an evaluation wrapper. Public-reference inference is reproducible; private ConeML training is outside the disclosure scope. Licensing: CC BY-NC 4.0 for the report, documentation, figures and authored results. The evaluation-wrapper source in supplement/public-evaluator-source.json is separately licensed under MIT, as stated in the archive. Third-party dataset, model and software terms remain applicable.