Peut-on faire confiance aux juges ? Validation de méthodes d'évaluation de la factualité par perturbation des réponses

Gharsallah, Sarra; Robaldo, Adele; Tokareva, Mariia; Gatti Pinheiro, Giovanni; Guendouz, Ilyana; Troncy, Raphaël; Papotti, Paolo; Michiardi, Pietro
CORIA-TALN 2025, 20e Conférence en Recherche d’Information et Applications, & 32e Conférence sur le Traitement Automatique des Langues Naturelles, (Actes de l’atelier Évaluation des modèles génératifs (LLM) et challenge 2025 (EvalLLM)), 30 June-4 July 2025, Marseille, France

Can WeTrust the Judges? Validation of Factuality Evaluation Methods via Answer Perturbation Evaluating the factual correctness of large language models (LLMs) is vital for many applications. But are our evaluation tools themselves trustworthy? Despite the rise of factuality-based metrics, their sensitivity and reliability remain underexplored. This paper introduces a meta-evaluation framework that systematically tests these metrics using controlled corruptions of gold standard answers. Our method generates ranked outputs with known degrees of degradation to probe how metrics capture nuanced changes in truthfulness. Our experiments reveal that pipeline-based methods, such as the RAGAS’sfactual correctness metric, better track degradation than LLM-as-judge approaches. We also propose a new variant of the factual correctness metric that provides a competitive and cost-efficient.


HAL
Type:
Conference
City:
Marseille
Date:
2025-06-30
Department:
Data Science
Eurecom Ref:
8942
Copyright:
Creative Commons Attribution 4.0 License (CC-BY)

PERMALINK : https://www.eurecom.fr/publication/8942