Training a small projector between a frozen vision encoder and language model is an established approach to multimodal learning. As the parameter count of language models scales dramatically, we revisit which vision capabilities this approach can add while keeping their pretrained weights fixed. Here we train a 50M par...
Jaehoon Lee, Harry B. Partridge, M. Jayasekara et al.· 0 citations
This work asks how the optimal learning rate and batch size move with model scale, family, and data, and whether one selection rule transfers across them; what LoRA trades against full fine-tuning, and how its rank and alpha set what the adapter can learn; whether validation loss faithfully ranks downstream quality.
Charles O'Neill, M. Jayasekara, Harry B. Partridge· 1 citation
This work follows invented facts written into Qwen3 models from creation through sequences of twenty to one hundred later writes, using held-out questions of five types, and finds that facts can be behaviourally forgotten without being erased.
Charles O'Neill· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.