Thomas Reed and George Mason (2025) “Hallucination Detection and Confidence Calibration for Large Language Model Outputs: Reproducible Experiments on HaluEval”, Journal of Artificial Intelligence Review, 6(4), pp. 1–17. doi:10.69987/.