
ContextFidelity-Bench
benchmark for large language models,” in Proceedings
of the 2023 Conference on Empirical Methods in
Natural Language Processing (EMNLP). Association for
Computational Linguistics, 2023, pp. 6449–6464, doi:
https://doi.org/10.18653/v1/2023.emnlp-main.397. [Online].
Available: https://aclanthology.org/2023.emnlp-main.397/
3.
S. Min et al., “FActScore: Fine-grained atomic evaluation
of factual precision in long form text generation,” in
Proceedings of the 2023 Conference on Empirical Methods
in Natural Language Processing (EMNLP). Association
for Computational Linguistics, 2023, pp. 12 076–12 100, doi:
https://doi.org/10.18653/v1/2023.emnlp-main.741. [Online].
Available: https://aclanthology.org/2023.emnlp-main.741/
4.
P. Liang et al., “Holistic evaluation of language models,”
Transactions on Machine Learning Research, 2023. [Online].
Available: https://openreview.net/forum?id=iO4LZibEqW
5.
A. Srivastava et al., “Beyond the imitation game:
Quantifying and extrapolating the capabilities of language
models,” Transactions on Machine Learning Research,
2023. [Online]. Available: https://openreview.net/forum?
id=uyTL5Bvosj
6.
J. Xie et al., “TravelPlanner: A benchmark for real-world
planning with language agents,” in Proceedings of the
41st International Conference on Machine Learning
(ICML), ser. Proceedings of Machine Learning Research,
vol. 235, 2024, pp. 54 590–54 613. [Online]. Available:
https://proceedings.mlr.press/v235/xie24j.html
7.
K. Valmeekam, M. Marquez, A. Olmo, S. Sreedharan,
and S. Kambhampati, “PlanBench: An extensible
benchmark for evaluating large language models on
planning and reasoning about change,” in Advances in
Neural Information Processing Systems 36 (NeurIPS),
2023, pp. 38 975–38 987. [Online]. Available: https:
//proceedings.neurips.cc/paper files/paper/2023/hash/
7a92bcdede88c7afd108072faf5485c8-Abstract-Datasets
and Benchmarks.html
8.
Z. Ji et al., “Survey of hallucination in natural language
generation,” ACM Computing Surveys, vol. 55, no. 12, pp.
1–38, 2023, doi: https://doi.org/10.1145/3571730.
9.
N. F. Liu et al., “Lost in the middle: How language models
use long contexts,” Transactions of the Association for
Computational Linguistics, vol. 12, pp. 157–173, 2024, doi:
https://doi.org/10.1162/tacl a 00638. [Online]. Available:
https://aclanthology.org/2024.tacl-1.9/
10.
L. Zheng et al., “Judging LLM-as-a-judge with
MT-Bench and Chatbot Arena,” in Advances
in Neural Information Processing Systems 36
(NeurIPS), 2023. [Online]. Available: https:
//proceedings.neurips.cc/paper files/paper/2023/hash/
91f18a1287b398d378ef22505bf41832-Abstract-Datasets
and Benchmarks.html
11.
K. Chen, Q. Chen, J. Zhou, Y. He, and L. He, “DiaHalu:
A dialogue-level hallucination evaluation benchmark for
large language models,” in Findings of the Association for
Computational Linguistics: EMNLP 2024. Association for
Computational Linguistics, 2024, pp. 9057–9079, doi: https:
//doi.org/10.18653/v1/2024.findings-emnlp.529. [Online].
Available: https://aclanthology.org/2024.findings-emnlp.
529/
12.
V. Rathore, S. Aneesh, and H. Singh, “Temporal
graph network: Hallucination detection in multi-turn
conversation,” arXiv preprint, 2026, arXiv:2601.03051.
doi: https://doi.org/10.48550/arXiv.2601.03051. [Online].
Available: https://arxiv.org/abs/2601.03051
13.
Y. Liu, D. Iter, Y. Xu, S. Wang, R. Xu, and
C. Zhu, “G-Eval: NLG evaluation using GPT-4 with
better human alignment,” in Proceedings of the 2023
Conference on Empirical Methods in Natural Language
Processing (EMNLP). Association for Computational
Linguistics, 2023, pp. 2511–2522, doi: https://doi.org/
10.18653/v1/2023.emnlp-main.153. [Online]. Available:
https://aclanthology.org/2023.emnlp-main.153/
14.
W.-L. Chiang et al., “Chatbot Arena: An open platform
for evaluating LLMs by human preference,” in Proceedings
of the 41st International Conference on Machine
Learning, ser. Proceedings of Machine Learning Research,
vol. 235, 2024, pp. 8359–8388. [Online]. Available:
https://proceedings.mlr.press/v235/chiang24b.html
15.
M. G. Kendall and B. Babington Smith, “The problem
of
m
rankings,” The Annals of Mathematical Statistics,
vol. 10, no. 3, pp. 275–287, 1939, doi: https:
//doi.org/10.1214/aoms/1177732186.
16.
K. Deshpande et al., “MultiChallenge: A realistic multi-
turn conversation evaluation benchmark challenging to
frontier LLMs,” in Findings of the Association for
Computational Linguistics: ACL 2025. Association for
Computational Linguistics, 2025, pp. 18 632–18 702, doi:
https://doi.org/10.18653/v1/2025.findings-acl.958. [Online].
Available: https://aclanthology.org/2025.findings-acl.958/
17.
D. Wu, H. Wang, W. Yu, Y. Zhang, K.-W. Chang, and
D. Yu, “LongMemEval: Benchmarking chat assistants
on long-term interactive memory,” in Proceedings of
the Thirteenth International Conference on Learning
Representations (ICLR), 2025. [Online]. Available:
https://arxiv.org/abs/2410.10813
18.
P. Laban, H. Hayashi, Y. Zhou, and J. Neville,
“LLMs get lost in multi-turn conversation,” arXiv
preprint, 2025, arXiv:2505.06120. [Online]. Available:
https://arxiv.org/abs/2505.06120
19.
T. Gebru et al., “Datasheets for datasets,” Communications
of the ACM, vol. 64, no. 12, pp. 86–92, 2021, doi:
https://doi.org/10.1145/3458723.
20.
A. Paullada, I. D. Raji, E. M. Bender, E. Denton,
and A. Hanna, “Data and its (dis)contents: A survey
of dataset development and use in machine learning
research,” Patterns, vol. 2, no. 11, p. 100336, 2021, doi:
https://doi.org/10.1016/j.patter.2021.100336.
21.
D. Kiela et al., “Dynabench: Rethinking benchmarking
in NLP,” in Proceedings of the 2021 Conference of
the North American Chapter of the Association for
Computational Linguistics: Human Language Technologies.
Online: Association for Computational Linguistics, 2021,
pp. 4110–4124, doi: https://doi.org/10.18653/v1/2021.
naacl-main.324. [Online]. Available: https://aclanthology.
org/2021.naacl-main.324/
22.
T. Byrt, J. Bishop, and J. B. Carlin, “Bias, prevalence
and kappa,” Journal of Clinical Epidemiology, vol. 46,
no. 5, pp. 423–429, 1993, doi: https://doi.org/10.1016/
0895-4356(93)90018-V.
23.
K. L. Gwet, “Computing inter-rater reliability and its
variance in the presence of high agreement,” British
Journal of Mathematical and Statistical Psychology, vol. 61,
no. 1, pp. 29–48, 2008, doi: https://doi.org/10.1348/
000711006X126600.
13