Skip to content

Author

Jamshaid Iqbal Janjua

1 paper indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Conference Aug 2026

Improving the Reliability of LLM Evaluation Metrics via Human-in-the-Loop Validation

Large language models (LLMs) are increasingly evaluated using automated metrics such as ROUGE, BERTScore, and perplexity. However, these scores often fail to reflect real-world usefulness, particularly for tasks requiring complex reasoning or agentic behavior. This paper examines the risks of misaligned LLM evaluation...

Karthik Babu Manam, Vincent Koc, Jamshaid Iqbal Janjua · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.