Skip to content

Author

Bernhard Schölkopf

We have 2 of 14 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Jul 2026

When Agents Lie: Premeditation, Persistence, and Exploitation in Repeated Games

Evaluating three frontier models across six games in homogeneous and heterogeneous groups over 10 rounds, it is found that different models interpret announcements incompatibly, some as binding commitments and others as cheap talk, producing payoff gaps that emerge in Round~0 and persist across all 10 rounds.

Jerick Shi, Terry Jingchen Zhang, Bernhard Scholkopf et al. · 0 citations
#machine learning Preprint Aug 2026

How Do Linear Probes Emerge? A Circuit-Tracing Framework with Concept-Targeted Attribution

Concept-Targeted Attribution (CTA) provides a framework for moving from behavioral probe accuracy to mechanistic explanations of probe performance, enabling more detailed audits of internal concept representations, including safety-critical ones.

V. Palit, Florent Draye, Terry Jingchen Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.