Personality science compresses a person into a small readable coordinate system, and the content science of the companion papers does the analogous thing for writing, eight character axes that collapse to two interpretable dimensions, matter against manner and originality. This paper reports the coupling between those two low dimensional spaces, and reframes what kind of object it is. The whole world connects a person to effective content billions of times a day through recommenders, but as a black box with no readable variable in the middle. We find a legible one. The full mapping between the whole of personality and the whole of content character does not hold; reduced to the two personality metatraits and the two content dimensions, a coupling appears, on the production side, and it survives being read by two unrelated model families. The reviewable spine is established on public data. On a public corpus of journalists writing across editorial sections, the same person seen across rooms, disposition read from a single article is mostly performed room and section register, the stable trait share about an eighth, and the person's plasticity metatrait still predicts the originality of the content they produce at about 0.39 once performance is averaged out, up from about 0.18 read article by article, so the coupling and the trait against performance split both survive on data a reviewer can re run, at the attenuated effect professional house writing predicts. A cross site authorship corpus held internally, in which the same pseudonymous person is seen writing across many separate sites, corroborates the same structure at a larger effect and enters as a secondary internal validation rather than the load bearing evidence: there disposition read from one context is about half stable trait and half performed room state, and it firms toward the person, to about three fifths, once a whole room is read rather than a single line. That even split reproduces, by an unrelated method, the roughly half trait and half situation division that whole trait personality theory found in beeper studies of behaviour, a convergence reported in the series related work as a measured reduction rather than a resemblance. An earlier and much smaller sample of the internal corpus had read the trait share far lower, near a quarter; that figure was a small sample artefact and does not survive either the larger corpus or a second independent scorer, both of which land the split close to even. On the internal corpus the metatrait bridge survives the separation, the person's stable plasticity still predicting the character they produce at about a half once performance is averaged out, and holding across five languages. Culture, sought first in geography and not found there, is real at the level of the discourse community: after removing the author, the subject, the language and the structural form, about a fifth of content character still belongs to the site.
Jason Duke· Zenodo (CERN European Organi...· 0 citations
Paper 5 shows that for the commercial web, truth is a register join rather than a judgement, and that grounding in retrieved evidence is what makes verification work. This paper is the system that follows, and its contribution is threefold: the score, the rulebook, and the distillation boundary. **The Veracity Score.** We summarise a domain with V = F × H, fact integrity times honesty. F is a multiplicative gate over corroborated confirmed falsehoods, so no amount of true boilerplate buys back one confirmed lie; H is the consequence weighted calibration of assertions against retrieved evidence. Regulatory permission is held separate and never folded into V, because a claim does not become less truthful by crossing a border. Every result carries its tier (well formed, key resolves, entity matches) and a corroboration gate: a mismatch is not called a lie without two independent signals. Measured on a gate versus average ablation over 1,248 diverging domains, the gate scores flagged domains at a mean honesty of 0.167 where an average would score 0.326. **The rulebook.** The register join settles identity facts but not "this supplement prevents COVID" or "the election was stolen". The move that reaches those is not to judge each afresh but to recognise when a claim repeats one an authority has already adjudicated: a unified, embedded index of adjudicated claims (its register grounded core the layer the published fact check feeds do not carry), matched by a two stage bi encoder then cross encoder recogniser, returning the authority's own ruling and citation. **The distillation boundary, and it is a negative and a positive.** The judgement the earlier work showed only a language model could perform is a *learnable function*: a compact encoder trained on the regulator's published adjudications reaches an area under the curve of **0.794** on a disjoint held out slice of the rulings, above the **0.703** of the evidence grounded language model it learned from and far above the roughly 0.46 of a text quality instrument, and it needs no language model at the point of use. It generalises across regulators (trained on United Kingdom advertising rulings, it flags 92.1 per cent of United States drug enforcement claims it never saw while false flagging 3.8 per cent of ordinary factual text). The companion negative is the boundary: the claim *extractor* does not distil the same way, a small generative student fabricates, so extraction stays a faithful large model and only the judgement is compressed. This is where distillation works and where it breaks, not a blanket claim that it works. **Scale.** Fifty thousand copies of one claim collapse to one canonical claim, verdicted once against the evidence and cached with its citation, so the per page cost collapses to recognise, normalise and look up; the expensive grounded judgement runs only on genuinely new claims. The recogniser is a property of the page and bakeable; the verdict is a fact about the evidence and is a live join, re verdicted only when the evidence changes. The system returns the record that settles a claim and a pointer to where it lives, never a proprietary opinion. --- ### 10.8 The composition: character, verification, and the sincerity residual The two halves of this paper measure different properties of the same entity. The character instrument measures how a site presents itself: how rigorous, how confident, how commercial its prose is.
Jason Duke· Zenodo (CERN European Organi...· 0 citations
Paper 5 shows that for the commercial web, truth is a register join rather than a judgement, and that grounding in retrieved evidence is what makes verification work. This paper is the system that follows, and its contribution is threefold: the score, the rulebook, and the distillation boundary. **The Veracity Score.** We summarise a domain with V = F × H, fact integrity times honesty. F is a multiplicative gate over corroborated confirmed falsehoods, so no amount of true boilerplate buys back one confirmed lie; H is the consequence weighted calibration of assertions against retrieved evidence. Regulatory permission is held separate and never folded into V, because a claim does not become less truthful by crossing a border. Every result carries its tier (well formed, key resolves, entity matches) and a corroboration gate: a mismatch is not called a lie without two independent signals. Measured on a gate versus average ablation over 1,248 diverging domains, the gate scores flagged domains at a mean honesty of 0.167 where an average would score 0.326. **The rulebook.** The register join settles identity facts but not "this supplement prevents COVID" or "the election was stolen". The move that reaches those is not to judge each afresh but to recognise when a claim repeats one an authority has already adjudicated: a unified, embedded index of adjudicated claims (its register grounded core the layer the published fact check feeds do not carry), matched by a two stage bi encoder then cross encoder recogniser, returning the authority's own ruling and citation. **The distillation boundary, and it is a negative and a positive.** The judgement the earlier work showed only a language model could perform is a *learnable function*: a compact encoder trained on the regulator's published adjudications reaches an area under the curve of **0.794** on a disjoint held out slice of the rulings, above the **0.703** of the evidence grounded language model it learned from and far above the roughly 0.46 of a text quality instrument, and it needs no language model at the point of use. It generalises across regulators (trained on United Kingdom advertising rulings, it flags 92.1 per cent of United States drug enforcement claims it never saw while false flagging 3.8 per cent of ordinary factual text). The companion negative is the boundary: the claim *extractor* does not distil the same way, a small generative student fabricates, so extraction stays a faithful large model and only the judgement is compressed. This is where distillation works and where it breaks, not a blanket claim that it works. **Scale.** Fifty thousand copies of one claim collapse to one canonical claim, verdicted once against the evidence and cached with its citation, so the per page cost collapses to recognise, normalise and look up; the expensive grounded judgement runs only on genuinely new claims. The recogniser is a property of the page and bakeable; the verdict is a fact about the evidence and is a live join, re verdicted only when the evidence changes. The system returns the record that settles a claim and a pointer to where it lives, never a proprietary opinion. --- ### 10.8 The composition: character, verification, and the sincerity residual The two halves of this paper measure different properties of the same entity. The character instrument measures how a site presents itself: how rigorous, how confident, how commercial its prose is.
Jason Duke· Zenodo (CERN European Organi...· 0 citations
Personality science compresses a person into a small readable coordinate system, and the content science of the companion papers does the analogous thing for writing, eight character axes that collapse to two interpretable dimensions, matter against manner and originality. This paper reports the coupling between those two low dimensional spaces, and reframes what kind of object it is. The whole world connects a person to effective content billions of times a day through recommenders, but as a black box with no readable variable in the middle. We find a legible one. The full mapping between the whole of personality and the whole of content character does not hold; reduced to the two personality metatraits and the two content dimensions, a coupling appears, on the production side, and it survives being read by two unrelated model families. The reviewable spine is established on public data. On a public corpus of journalists writing across editorial sections, the same person seen across rooms, disposition read from a single article is mostly performed room and section register, the stable trait share about an eighth, and the person's plasticity metatrait still predicts the originality of the content they produce at about 0.39 once performance is averaged out, up from about 0.18 read article by article, so the coupling and the trait against performance split both survive on data a reviewer can re run, at the attenuated effect professional house writing predicts. A cross site authorship corpus held internally, in which the same pseudonymous person is seen writing across many separate sites, corroborates the same structure at a larger effect and enters as a secondary internal validation rather than the load bearing evidence: there disposition read from one context is about half stable trait and half performed room state, and it firms toward the person, to about three fifths, once a whole room is read rather than a single line. That even split reproduces, by an unrelated method, the roughly half trait and half situation division that whole trait personality theory found in beeper studies of behaviour, a convergence reported in the series related work as a measured reduction rather than a resemblance. An earlier and much smaller sample of the internal corpus had read the trait share far lower, near a quarter; that figure was a small sample artefact and does not survive either the larger corpus or a second independent scorer, both of which land the split close to even. On the internal corpus the metatrait bridge survives the separation, the person's stable plasticity still predicting the character they produce at about a half once performance is averaged out, and holding across five languages. Culture, sought first in geography and not found there, is real at the level of the discourse community: after removing the author, the subject, the language and the structural form, about a fifth of content character still belongs to the site.
Jason Duke· Zenodo (CERN European Organi...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.