This work introduces OptimismBench, which detects directional bias with inverted pairs: each scenario elicits both P(success) and P(failure), and asymmetry between the two framings yields a signed bias score without ground truth.
Seonglae Cho, A. Koshiyama· arXiv.org· 0 citations
This work analyzesparse autoencoder features across six models and three SAE families and zero-ablate at full layer depth, finding cross-family claims are sensitive to training methodology, not just activation function or scale.
Seonglae Cho, Zekun Wu, Kleyton Da Costa et al.· arXiv.org· 1 citation
Behavioral topology is shaped more by the deployment harness than by the LLM, providing a model-agnostic structural primitive for safety auditing and runtime monitoring, and addresses both prediction goals.
Seonglae Cho, F. Fernandez, Umar Mohammed et al.· 0 citations
This work tracks quantization across 16 models from 8 families under round-to-nearest, seven under AWQ, two under GPTQ and one under GGUF, at 8 down to 2 bits, and measures the margin, the picked option's score minus its best alternative's, which removes the protection a large margin affords.
Web agents observe a browser through text, pixels, or both, and the choice is usually fixed once for all tasks, so a stronger agent can overturn the result, and the rerun noise bands and the full measurement protocol are reported.
Jiaming Wei, Zekun Wu, A. Koshiyama et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.