Large Language Models for Code Generation from Multilingual Prompts: A Curated Benchmark and a Study on Code Quality
The results show that English prompts do not consistently produce the best functional correctness or code quality, the impact of prompt language depends on both the programming language and the LLM, and generated code frequently mixes English with the prompt language in comments and string literals.