Modeling Memory-Dependent Reliability of LLMs: A Hidden Markov Model
A hierarchical Bayesian framework for LLM reliability assessment is extended by relaxing the assumption of independent task outcomes and introducing a Hidden Markov Model to capture sequential dependence in benchmark-constructed interaction sessions, suggesting that ignoring sequential dependence may lead to overconfident reliability estimates.