Skip to content

Author

Nigel Collier

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Oct 2026

VeriHarness: Scaling Agentic Verification for Long-Horizon Tasks

As LLM agents undertake increasingly complex, long-horizon tasks, verifying their outputs becomes increasingly challenging. We study how verification capability can be strengthened with a fixed base model, without access to reference answers or grading rubrics at test time. Repeated sampling yields multiple rollouts th...

Caiqi Zhang, Ru-Jun Han, Zifeng Wang et al. · 0 citations

When Calibration Rankings Reverse: Accuracy-Controlled Evaluation for Fair Comparison of LLMs

It is shown that raw global calibration metrics are not robust for cross-model comparison, and that fair calibration comparison requires accuracy-aware evaluation, and proposed ACE, an accuracy-controlled evaluation framework, is proposed.

Zhichao Yang, Caiqi Zhang, Ruihan Yang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.