Skip to content

Author

Samee Arif

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

One Example Is Enough to Pass Fairness Benchmarks: Rethinking Fairness Evaluation for Aligned LLMs

It is argued that BBQ-style multiple-choice abstention benchmarks measure a single structural cue, and a model that solves them does not thereby become fair, and is called for evaluation suites that cover a broader spectrum of fairness alignment.

Naihao Deng, Samee Arif, Shuai-Chen Chang et al. · 0 citations

The Age of Curiosity Meets the Age of AI: Benchmarking Child Safety in Large Language Models

KIDBench, a benchmark for evaluating child-facing LLM safety for ages 7-11 using a LLM-as-a-Judge rubric grounded in developmental-psychology, is introduced and KIDGuardLlama, a child-safety evaluator, and KIDLlama, a child-safe response model are introduced.

Samee Arif, Angana Borah, Rada Mihalcea · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.