KronSAE is proposed, a design that factorizes the latent space into heads and forms post-latent features as pairwise compositions of lower-dimensional pre-latents using mAND, a differentiable AND-like interaction that imposes a compositional co-activation prior while remaining compatible with standard SAE objectives and variants.
Vadim Kurochkin, Yaroslav Aksenov, Daniil Laptev et al.· 1 citation
Layer-Integrated Memory (LIMe) is introduced, a lightweight extension that leverages existing key-value buffers and learns per-head, per-layer routing weights to integrate representations from previous layers to improve perplexity per FLOP and yield strong gains on synthetic tasks while preserving higher value-vector entropy and token separability.
This work introduces Soft Latent Thinking, a method that replaces the LM head during reasoning with a lightweight projector, enabling autoregressive rollout in embedding space where reasoning steps remain continuous rather than tokenized.
N. Koriagin, Yaroslav Aksenov, George Bredis et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.