LLM judges are widely used to evaluate model outputs, but their verdicts can be unreliable: a judge may favor the worse answer for its position, length, or other surface features. When a judge is wrong, is the information needed to judge correctly absent from the model, or present in its internal representations but no...
This work introduces interactional cultural markers, measurable patterns of doctor-patient interaction grounded in cross-cultural clinical communication, and uses them to compare real, simulated, and synthetic consultations from Indian and US clinical contexts to find distinct patterns of participation and control.
Krithi Shailya, Siddharth D. Jaiswal, A. Makani et al.· 0 citations
This work introduces Pluralis v0.1, a novel multimodal, multi-regional, and multilingual dataset built from a culture-first perspective and calls upon the research community to utilize this foundation to advance the science of multilingual, multicultural evaluation to better support AI cultural alignment globally.
Alicia Parrish, Rajat C. Shinde, Sanket Badhe et al.· arXiv.org· 0 citations
A single trace-extraction regex, not the model, manufactured a multilingual failure: two worked examples raise one model's measured accuracy twenty-sixfold while its accuracy on readable outputs barely moves.