Skip to content

Author

Karmendra Pandey

21 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#explainable ai Open access Oct 2026

PREreview of "What AI Benchmarks Actually Measure: Adapting Convergent and Discriminant Validity to Interrogate Fifty-Six AI Benchmarks"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23180288. ## Summary Benchmark scores shape which models get funded, bought, and regulated — yet whether benchmarks measure what they claim to measure is rarely tested. This paper i...

Karmendra Pandey · 0 citations
#generative ai Open access Oct 2026

PREreview of "COMED: The Missing Middle Between Routing and Collaboration in Multi-LLM Inference"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23177551. ## Summary Routing picks one model and stops; dense collaboration invokes peers on every query. COMED argues both ends of this spectrum are wrong and formalizes the middle...

Karmendra Pandey · 0 citations
#generative ai Open access Oct 2026

PREreview of "Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23197136. ## Summary This paper studies Leni, a production enterprise AI business-analyst agent, and asks a question most vendor papers avoid: where does the system's measured relia...

Karmendra Pandey · 0 citations
#generative ai Open access Oct 2026

PREreview of "Memory Control Signals Emerge Before Action in Long Horizon Agents"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23197372. ## Summary This paper asks whether a language model already represents its memory needs — when to compress history, when to recall earlier evidence — in its hidden state b...

Karmendra Pandey · 0 citations
#generative ai Open access Oct 2026

PREreview of "Beyond Static Sandboxing: Learned Capability Governance for Autonomous AI Agents"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23197319. ## Summary This paper names a problem every agent deployer has felt but few have measured: open-source agent runtimes like OpenClaw expose every tool to every session by d...

Karmendra Pandey · 0 citations
#generative ai Open access Oct 2026

PREreview of "TRACER: Trace-Based Adaptive Cost-Efficient Routing for LLM Classification"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23179235. ## Summary This paper tackles one of the most economically significant problems in production LLM deployment: classification endpoints where every query currently pays ful...

Karmendra Pandey · 0 citations
#generative ai Open access Oct 2026

PREreview of "Where Does Agent Reliability Come From? A Cross-Benchmark Decomposition of Verification Loops, Specialist Models, and Scaffolding in a Production Enterprise Agent"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23197406. ## Summary This paper studies Leni, a production enterprise AI business-analyst agent, and asks a question most vendor papers avoid: where does the system's measured relia...

Karmendra Pandey · 0 citations
#generative ai Open access Oct 2026

PREreview of "Towards a Science of AI Agent Reliability"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23179310. ## Summary This paper argues that compressing agent behavior into a single success metric obscures the operational flaws that matter in deployment. Grounded in safety-crit...

Karmendra Pandey · 0 citations
#small language model Open access Oct 2026

PREreview of "Measured Joules, Learned Routes: Learning to Route for Energy-Efficient LLM Serving"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23177620. ## Summary Routing papers usually optimize proxy costs — parameter counts, API prices, model counts. This one measures the thing itself: per-query GPU energy. The authors...

Karmendra Pandey · 0 citations
#small language model Open access Oct 2026

PREreview of "Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23129468. Summary This paper attacks what it calls the compliance fiction: the industry practice of treating regulatory conformity as a binary verdict declared at deployment time, w...

Karmendra Pandey · 0 citations
#explainable ai Open access Oct 2026

PREreview of "Who judges the judges? Governance from metrics: a runtime framework for continuous LLM compliance monitoring"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23129468. Summary This paper attacks what it calls the compliance fiction: the industry practice of treating regulatory conformity as a binary verdict declared at deployment time, w...

Karmendra Pandey · 0 citations
#artificial intelligence Open access Oct 2026

PREreview of "You Cannot Pick a Provider From the Price List: Market-Aware Routing for Open-Weight LLM Inference"

This Zenodo record is a permanently preserved version of a PREreview. You can view the complete PREreview at https://prereview.org/reviews/23147062. ## Summary This paper identifies a genuinely under-appreciated routing axis: after a model router picks Llama-3.3-70B, the client must still choose *which provider serves...

Karmendra Pandey · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.