Skip to content

Author

Shauli Ravfogel

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Making LLMs Say What They Think: Measuring and Improving CoT-Interpretability Alignment

Chain-of-thought (CoT) traces often serve as a proxy for how Large Language Models (LLMs) arrive at their answers. However, growing evidence shows that models'CoT often fails to reflect their internal computations and can be changed without affecting their final answers. In this work, we measure and improve the alignme...

Yi-Huai Hong, Shauli Ravfogel, Chen Zhao et al. · 0 citations

From Directions to Regions: Decomposing Activations in Language Models via Local Geometry

This work leverage Mixture of Factor Analyzers (MFA) as a scalable, unsupervised alternative that models the activation space as a collection of Gaussian regions with their local covariance structure, and positions local geometry, expressed through subspaces, as a promising unit of analysis for scalable concept discove...

Or Shafran, S. Ronen, Omri Fahn et al. · 10 citations · ⚡2

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.