Skip to content

Author

Abhishek Nagaraj

We have 3 of 6 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Review Sep 2026

ORQA: An Occupation-Realistic Question and Answer Framework for LLM Professional Knowledge

We present ORQA, a method for testing occupation-level knowledge in large language models. Prior methods either map abstract LLM skills to occupations via task definitions or utilize expert knowledge which is difficult to obtain at scale and expensive. ORQA complements both of these methods by connecting O*NET occupations to trusted occupation-specific websites (such as regulatory agencies, licensing bodies, professional organizations, and government publications) and converting these into source-traceable question-answer pairs. A combination of an automated pipeline and human review produces a set of high quality questions about occupations. The question set created via our method covers 116 occupations from all 21 major groups in the SOC, with 480 questions sourced from 187 different websites. Each question is designed to probe a real-world skill question that is relevant to the occupation in question. We test 15 state-of-the-art frontier and open-weight models via this method. Claude Opus 4.6, GPT-5.4 and Claude Sonnet 4.6 all perform the best at approximately 58-62% while smaller open-weight models achieve approximately 33-41% performance. Performance varies significantly across occupations. Healthcare-related occupations achieve the highest performance (78%) while Office and Administrative Support achieve approximately 40%. Performance on individual occupations (e.g. Sheet Metal Workers and Fish and Game Wardens) is essentially zero. We also find that open-ended questions and weighting by wage bill do not significantly affect the ranking of models on this benchmark. We believe that leveraging existing trusted occupation-specific information to test LLM knowledge in professional domains may be a scalable and useful method for evaluating occupation-level AI performance in the future. Results and data are available at orqabench.org.

Shreyas R. Krishnan, Serina Chang, Abhishek Nagaraj · 0 citations
Open access Jul 2026

Simulating strategic interactions with AI agents

This article introduces a framework for designing and running simulated experiments with LLM‐powered agents and applies the framework to the exploration–exploitation dilemma and shows that LLM‐based experiments reproduce patterns observed among human participants.

Matteo Tranchero, Cecil-Francis Brenninkmeijer, Arul Murugan et al. · 0 citations
#artificial intelligence Preprint Aug 2026

CentaurBench: Benchmarking LLM Capabilities on Augmenting vs. Automating Real-World Work Tasks

A unified framework that evaluates the capability of models to automate and augment another agent's performance, motivating benchmarks that evaluate models according to the roles they play in human-AI and multi-agent systems.

Pattaraphon Kenny Wongchamcharoen, K. Gulati, Min Min Fong et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.