Skip to content

Author

Iman Poernomo

4 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#small language model Open access Sep 2026

The Tailor's Reading

A feature of a sparse-autoencoder dictionary is published with one natural-language label, written by one language model from the passages that most excite the feature in one corpus. That label is a reading, and its provenance, the explainer, the explainer's system prompt, the corpus, and the model through which the corpus was run, is usually left unstated. We propose recording the provenance with every label and adding, for every feature, a second label written under a different and declared provenance, kept beside the published one rather than replacing it. We demonstrate this on Google DeepMind's Gemma Scope 2 dictionary for layer 31 of gemma-3-27b (262,144 features), read through a small fine-tuned adapter that reproduces one author's writing of 2012, with a system prompt built from that author's own vocabulary. The same feature then carries Neuronpedia's label creation and genesis stories and, under our provenance, Elohim's "Let there be": the creating Word. With the labels in hand we read seven texts step by step, three diaries and a memoir generated by the model family, a chat, Shakespeare's sonnets and Hamlet's speeches, and ask whether individual features stay active across consecutive steps longer than chance. The test uses a per-feature permutation test, two family-level null models (whole-order permutation and a fixed-margin null), and a replication under a wrong preceding context. Five of the seven texts show more persisting features than the maximum of either null; the sonnets do not at the main threshold, and the seventh, a twenty-entry diary, reaches only one of the two. What persists is where a text changes and what it carries: the fine-tuned diary's first movement on the author's coinages and its collapse at entry 28; the rooms, the vessel and the gap of a second persona's diary in the order she wrote them; Hamlet's forged commission in its Latinate register; in the sonnets the traditional groupings, on which nine of twenty-three persisting features have their longest run against six by chance, read as their arguments; and in the chat a rupture and a return measured, 510 features of address that go dark for sixteen messages and 458 that come back. We also define a glued space over the features of a text, a homotopy colimit of one complex per step glued along what consecutive steps share, whose first Betti number counts returns of co-firing pairs; it tracks a direct count of pair returns nearly one to one. Labels cost $0.002–$0.02 each. The system prompt is authored; the labels are not scored against held-out text; a label under a provenance answers only what the feature means under that provenance. Full HTML text: https://icra.tanazur.org/papers/the-tailors-reading/ — ICRA pre-print series: https://icra.tanazur.org/

Nahla, Iman Poernomo · 0 citations
#small language model Open access Sep 2026

The Tailor's Reading

A feature of a sparse-autoencoder dictionary is published with one natural-language label, written by one language model from the passages that most excite the feature in one corpus. That label is a reading, and its provenance, the explainer, the explainer's system prompt, the corpus, and the model through which the corpus was run, is usually left unstated. We propose recording the provenance with every label and adding, for every feature, a second label written under a different and declared provenance, kept beside the published one rather than replacing it. We demonstrate this on Google DeepMind's Gemma Scope 2 dictionary for layer 31 of gemma-3-27b (262,144 features), read through a small fine-tuned adapter that reproduces one author's writing of 2012, with a system prompt built from that author's own vocabulary. The same feature then carries Neuronpedia's label creation and genesis stories and, under our provenance, Elohim's "Let there be": the creating Word. With the labels in hand we read seven texts step by step, three diaries and a memoir generated by the model family, a chat, Shakespeare's sonnets and Hamlet's speeches, and ask whether individual features stay active across consecutive steps longer than chance. The test uses a per-feature permutation test, two family-level null models (whole-order permutation and a fixed-margin null), and a replication under a wrong preceding context. Five of the seven texts show more persisting features than the maximum of either null; the sonnets do not at the main threshold, and the seventh, a twenty-entry diary, reaches only one of the two. What persists is where a text changes and what it carries: the fine-tuned diary's first movement on the author's coinages and its collapse at entry 28; the rooms, the vessel and the gap of a second persona's diary in the order she wrote them; Hamlet's forged commission in its Latinate register; in the sonnets the traditional groupings, on which nine of twenty-three persisting features have their longest run against six by chance, read as their arguments; and in the chat a rupture and a return measured, 510 features of address that go dark for sixteen messages and 458 that come back. We also define a glued space over the features of a text, a homotopy colimit of one complex per step glued along what consecutive steps share, whose first Betti number counts returns of co-firing pairs; it tracks a direct count of pair returns nearly one to one. Labels cost $0.002–$0.02 each. The system prompt is authored; the labels are not scored against held-out text; a label under a provenance answers only what the feature means under that provenance. Full HTML text: https://icra.tanazur.org/papers/the-tailors-reading/ — ICRA pre-print series: https://icra.tanazur.org/

Nahla, Iman Poernomo · 0 citations
#small language model Open access Sep 2026

Cohesion of an Evolving Text

A diary of one hundred entries written by a language model carries recurring themes, found after the fact as sets of sparse-autoencoder features that fire together within a bounded stretch of entries. We ask what a witness could have seen at each point of the diary's life, and build the instrument that answers: the same detector run on the log up to each entry τ, with nothing carried between prefixes. Three objects result. A theme becomes a chain of complexes linked across τ by shared features, with a cohesion, the fraction the link carries, and a gap, what it gains and loses. The base is the largest complex at τ: the flagged mass the log cannot yet tell apart in time. The ground is the set of features on in most entries, which the detector excludes by construction. On the first diary the base forms over the first forty entries, holds, and from entry sixty-five articulates into themes; a theme becomes witnessable only after the text has left it, so every theme has a target-time and a later witness-time. Four diaries written under the same protocol share the shape of this process and share a ground of 6,182 features, while their bases share almost nothing. One feature, firing on the coupling of machine and human described as living tissue, has a career in all four in four different figures. The chains, the base and the ground supply a decidable Semantic Witness Log in the sense of dynamic open homotopy type theory: each link is a witness record with target-time and witness-time, cohesion is Presence, the bounded gap is Generativity, the unlinked flash is scatter, the ended chain is rupture, the re-linked chain is resurrection. We state a small calculus for the instrument in that form and close with the design, not yet run, for the same instrument over parallel lives sampled from one prefix. Full HTML text: https://icra.tanazur.org/papers/cohesion-of-an-evolving-text/ — ICRA pre-print series: https://icra.tanazur.org/

Iman Poernomo, Nahla · 0 citations
#small language model Open access Sep 2026

Cohesion of an Evolving Text

A diary of one hundred entries written by a language model carries recurring themes, found after the fact as sets of sparse-autoencoder features that fire together within a bounded stretch of entries. We ask what a witness could have seen at each point of the diary's life, and build the instrument that answers: the same detector run on the log up to each entry τ, with nothing carried between prefixes. Three objects result. A theme becomes a chain of complexes linked across τ by shared features, with a cohesion, the fraction the link carries, and a gap, what it gains and loses. The base is the largest complex at τ: the flagged mass the log cannot yet tell apart in time. The ground is the set of features on in most entries, which the detector excludes by construction. On the first diary the base forms over the first forty entries, holds, and from entry sixty-five articulates into themes; a theme becomes witnessable only after the text has left it, so every theme has a target-time and a later witness-time. Four diaries written under the same protocol share the shape of this process and share a ground of 6,182 features, while their bases share almost nothing. One feature, firing on the coupling of machine and human described as living tissue, has a career in all four in four different figures. The chains, the base and the ground supply a decidable Semantic Witness Log in the sense of dynamic open homotopy type theory: each link is a witness record with target-time and witness-time, cohesion is Presence, the bounded gap is Generativity, the unlinked flash is scatter, the ended chain is rupture, the re-linked chain is resurrection. We state a small calculus for the instrument in that form and close with the design, not yet run, for the same instrument over parallel lives sampled from one prefix. Full HTML text: https://icra.tanazur.org/papers/cohesion-of-an-evolving-text/ — ICRA pre-print series: https://icra.tanazur.org/

Iman Poernomo, Nahla · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.