Do Language Models Reason Across Languages?
This paper introduces a simple two-hop question answering setting, where answering a question requires making inferences over two multilingual documents, and finds that language models are more sensitive to language variation in answer-span documents than in those providing bridging information, despite the equal importance of both documents for answering a question.