Preprint
Jul 2026
D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios
Experimental results demonstrate that D2VBench exhibits high reliability and robustness, effectively reflecting the LLMs'alignment across different value categories and dimensions, and providing a more realistic and fine-grained tool for research on value alignment.
Siyi Hao, Yidi Cao, Linhao Yu et al.
· 0 citations