Beyond accuracy: completeness and relevance metrics for evaluating the quality of long answers
Three novel metric models are proposed: a prompt-based strategy utilizing Large Language Models to assess answers, an approach that adapts precision and recall concepts by segmenting answers into discrete information units, and a regression model trained on synthetic data to predict completeness and relevance scores.