ContextFidelity-Bench: A Three-Paradigm Framework for Evaluating Multi-Turn Hallucination Behavior in Language Models

Authors

  • Parsa Bakhtary Google, Mountain View, CA, USA Author

DOI:

https://doi.org/10.66834/0azb3w20

Keywords:

Large Language Models, hallucination evaluation, multi-turn evaluation, benchmarks, context fidelity, source fidelity

Abstract

We introduce ContextFidelity‑Bench, a reproducible benchmark and evaluation framework for studying hallucination‑related behavior in multi‑turn language model conversations. The framework comprises three paradigms: numeric accuracy across a 20‑turn financial analysis that ends with a data correction (P1), constraint satisfaction in a 20‑turn scheduling task whose requirements accumulate, are revised, and may become jointly unsatisfiable (P2), and source fidelity in a multi‑document synthesis with planted contradictions, information gaps, and authority manipulations (P3). The released artifact includes scenario generators with pre‑computed ground truth and verified feasibility labels, scoring scripts with a documented protocol version, prompt templates, and 550 conversations from five API‑accessible models, of which 4,250 P1 answers and 800 P2 checkpoints are scored automatically and 100 P3 syntheses are judged. An initial case study illustrates the framework’s diagnostic use. Numeric accuracy and feasible‑plan success rank the five models similarly (Spearman \(\rho=0.90\)), whereas source‑fidelity abstention ranks them differently: the model with the lowest numeric accuracy (0.880) has the highest gap abstention (0.975) and no fabrication, while the model that produces complete valid plans most often (24 of 29 feasible post‑revision checkpoints) has the joint‑lowest gap abstention (0.875). When requirements are jointly unsatisfiable, models differ in whether they say so and in whether they still supply a plan; when a new requirement merely appears to conflict, some models raise false alarms. These observations come from a five‑model, single‑run snapshot and are presented as hypothesis‑generating; the framework is designed for replication and extension.

References

1. S. Lin, J. Hilton, and O. Evans, “TruthfulQA: Measuring how models mimic human falsehoods,” in Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (ACL). Association for Computational Linguistics, 2022, pp. 3214–3252, doi:https://doi.org/10.18653/v1/2022.acl-long.229. [Online]. Available: https://aclanthology.org/2022.acl-long.229/ DOI: https://doi.org/10.18653/v1/2022.acl-long.229

2. J. Li, X. Cheng, W. X. Zhao, J.-Y. Nie, and J.-R. Wen, “HaluEval: A large-scale hallucination evaluation benchmark for large language models,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 2023, pp. 6449–6464, doi:https://doi.org/10.18653/v1/2023.emnlp-main.397. [Online]. Available: https://aclanthology.org/2023.emnlp-main.397/ DOI: https://doi.org/10.18653/v1/2023.emnlp-main.397

3. S. Min et al., “FActScore: Fine-grained atomic evaluation of factual precision in long form text generation,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 2023, pp. 12076–12100, doi:https://doi.org/10.18653/v1/2023.emnlp-main.741. [Online]. Available: https://aclanthology.org/2023.emnlp-main.741/ DOI: https://doi.org/10.18653/v1/2023.emnlp-main.741

4. P. Liang et al., “Holistic evaluation of language models,” Transactions on Machine Learning Research, 2023. [Online]. Available: https://openreview.net/forum?id=iO4LZibEqW

5. A. Srivastava et al., “Beyond the imitation game: Quantifying and extrapolating the capabilities of language models,” Transactions on Machine Learning Research, 2023. [Online]. Available: https://openreview.net/forum?id=uyTL5Bvosj

6. J. Xie et al., “TravelPlanner: A benchmark for real-world planning with language agents,” in Proceedings of the 41st International Conference on Machine Learning (ICML), ser. Proceedings of Machine Learning Research, vol. 235, 2024, pp. 54590–54613. [Online]. Available: https://proceedings.mlr.press/v235/xie24j.html

7. K. Valmeekam, M. Marquez, A. Olmo, S. Sreedharan, and S. Kambhampati, “PlanBench: An extensible benchmark for evaluating large language models on planning and reasoning about change,” in Advances in Neural Information Processing Systems 36 (NeurIPS), 2023, pp. 38975–38987. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2023/hash/7a92bcdede88c7afd108072faf5485c8-Abstract-Datasets_and_Benchmarks.html DOI: https://doi.org/10.52202/075280-1693

8. Z. Ji et al., “Survey of hallucination in natural language generation,” ACM Computing Surveys, vol. 55, no. 12, pp. 1–38, 2023, doi:https://doi.org/10.1145/3571730. DOI: https://doi.org/10.1145/3571730

9. N. F. Liu et al., “Lost in the middle: How language models use long contexts,” Transactions of the Association for Computational Linguistics, vol. 12, pp. 157–173, 2024, doi:https://doi.org/10.1162/tacl_a_00638. [Online]. Available: https://aclanthology.org/2024.tacl-1.9/ DOI: https://doi.org/10.1162/tacl_a_00638

10. L. Zheng et al., “Judging LLM-as-a-judge with MT-Bench and Chatbot Arena,” in Advances in Neural Information Processing Systems 36 (NeurIPS), 2023. [Online]. Available: https://proceedings.neurips.cc/paper_files/paper/2023/hash/91f18a1287b398d378ef22505bf41832-Abstract-Datasets_and_Benchmarks.html

11. K. Chen, Q. Chen, J. Zhou, Y. He, and L. He, “DiaHalu: A dialogue-level hallucination evaluation benchmark for large language models,” in Findings of the Association for Computational Linguistics: EMNLP 2024. Association for Computational Linguistics, 2024, pp. 9057–9079, doi:https://doi.org/10.18653/v1/2024.findings-emnlp.529. [Online]. Available: https://aclanthology.org/2024.findings-emnlp.529/ DOI: https://doi.org/10.18653/v1/2024.findings-emnlp.529

12. V. Rathore, S. Aneesh, and H. Singh, “Temporal graph network: Hallucination detection in multi-turn conversation,” arXiv preprint, 2026, arXiv:2601.03051. doi:https://doi.org/10.48550/arXiv.2601.03051. [Online]. Available: https://arxiv.org/abs/2601.03051

13. Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and C. Zhu, “G-Eval: NLG evaluation using GPT-4 with better human alignment,” in Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Linguistics, 2023, pp. 2511–2522, doi:https://doi.org/10.18653/v1/2023.emnlp-main.153. [Online]. Available: https://aclanthology.org/2023.emnlp-main.153/ DOI: https://doi.org/10.18653/v1/2023.emnlp-main.153

14. W.-L. Chiang et al., “Chatbot Arena: An open platform for evaluating LLMs by human preference,” in Proceedings of the 41st International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 235, 2024, pp. 8359–8388. [Online]. Available: https://proceedings.mlr.press/v235/chiang24b.html

15. M. G. Kendall and B. Babington Smith, “The problem of m rankings,” The Annals of Mathematical Statistics, vol. 10, no. 3, pp. 275–287, 1939, doi:https://doi.org/10.1214/aoms/1177732186. DOI: https://doi.org/10.1214/aoms/1177732186

16. K. Deshpande et al., “MultiChallenge: A realistic multi-turn conversation evaluation benchmark challenging to frontier LLMs,” in Findings of the Association for Computational Linguistics: ACL 2025. Association for Computational Linguistics, 2025, pp. 18632–18702, doi:https://doi.org/10.18653/v1/2025.findings-acl.958. [Online]. Available: https://aclanthology.org/2025.findings-acl.958/ DOI: https://doi.org/10.18653/v1/2025.findings-acl.958

17. D. Wu, H. Wang, W. Yu, Y. Zhang, K.-W. Chang, and D. Yu, “LongMemEval: Benchmarking chat assistants on long-term interactive memory,” in Proceedings of the Thirteenth International Conference on Learning Representations (ICLR), 2025. [Online]. Available: https://arxiv.org/abs/2410.10813

18. P. Laban, H. Hayashi, Y. Zhou, and J. Neville, “LLMs get lost in multi-turn conversation,” arXiv preprint, 2025, arXiv:2505.06120. [Online]. Available: https://arxiv.org/abs/2505.06120

19. T. Gebru et al., “Datasheets for datasets,” Communications of the ACM, vol. 64, no. 12, pp. 86–92, 2021, doi:https://doi.org/10.1145/3458723. DOI: https://doi.org/10.1145/3458723

20. A. Paullada, I. D. Raji, E. M. Bender, E. Denton, and A. Hanna, “Data and its (dis)contents: A survey of dataset development and use in machine learning research,” Patterns, vol. 2, no. 11, p. 100336, 2021, doi:https://doi.org/10.1016/j.patter.2021.100336. DOI: https://doi.org/10.1016/j.patter.2021.100336

21. D. Kiela et al., “Dynabench: Rethinking benchmarking in NLP,” in Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. Online: Association for Computational Linguistics, 2021, pp. 4110–4124, doi:https://doi.org/10.18653/v1/2021.naacl-main.324. [Online]. Available: https://aclanthology.org/2021.naacl-main.324/ DOI: https://doi.org/10.18653/v1/2021.naacl-main.324

22. T. Byrt, J. Bishop, and J. B. Carlin, “Bias, prevalence and kappa,” Journal of Clinical Epidemiology, vol. 46, no. 5, pp. 423–429, 1993, doi:https://doi.org/10.1016/0895-4356(93)90018-V. DOI: https://doi.org/10.1016/0895-4356(93)90018-V

23. K. L. Gwet, “Computing inter-rater reliability and its variance in the presence of high agreement,” British Journal of Mathematical and Statistical Psychology, vol. 61, no. 1, pp. 29–48, 2008, doi:https://doi.org/10.1348/000711006X126600. DOI: https://doi.org/10.1348/000711006X126600

cover page

Additional Files

Published

2026-09-30

Issue

Section

Full Length Articles/Research articles

How to Cite

ContextFidelity-Bench: A Three-Paradigm Framework for Evaluating Multi-Turn Hallucination Behavior in Language Models. (2026). BenchCouncil Transactions on Benchmarks, Standards and Evaluations, 6. https://doi.org/10.66834/0azb3w20