Developmental Trajectories in a Sparse Language Model with 348M Active Parameters
A LAMBADA accuracy difference discovered on 128 repeatedly inspected examples persists on the additional 5,025 examples of the same benchmark. At approximately four billion training positions, ConeML scores 30.51% versus 24.84% for Pythia-1B; at approximately ten billion positions, its later branch scores 39.92% versus...