Benchmark Validity in Financial Language-Model Evaluation
Benchmark results are increasingly used to compare language models for financial decision-support tasks, but such comparisons can confound domain learning with compliance with task-specific output contracts. This study examines that measurement problem using FinVector-Market-4B, a LoRA adaptation of Qwen3.5-4B for stru...