Does Finetuning with Scientific Data Increase Hallucinations? A Multi-domain Factuality Evaluation of LLMs
SciFactCheck, a benchmark of 2,500 prompts across five scientific domains, is paired with a modular evaluation framework targeting three factuality hallucination types: unverifiability, overclaim, and attribution, and fundamentally challenge current methods of domain-specific fine-tuning for factuality and call for developing improved verification infrastructure for scientific content.