THOMAS REED; GEORGE MASON. Hallucination Detection and Confidence Calibration for Large Language Model Outputs: Reproducible Experiments on HaluEval. Journal of Artificial Intelligence Review, [S. l.], v. 6, n. 4, p. 1–17, 2025. DOI: 10.69987/. Disponível em: https://learnedvertex.com/index.php/JAIR/article/view/321. Acesso em: 10 oct. 2026.