ConflictBank: A Benchmark for Evaluating the Influence of Knowledge Conflicts in LLM

Zhaochen Su, Jun Zhang, Xiaoye Qu, Tong Zhu, Yanshu Li, Jiashuo Sun, Juntao Li, Min Zhang, Yu Cheng

Advances in Neural Information Processing Systems (NeurIPS), 2024

Abstract

ConflictBank systematically evaluates knowledge conflicts in large language models across retrieved evidence, encoded model knowledge, and interactions between the two. The benchmark covers misinformation, temporal discrepancies, and semantic divergence to support the development of more reliable models.

Paper  |  Code