The proposed MIMIC framework fundamentally transforms algorithms into verifiable reasoning trajectories through narrative fusion, code-guided test synthesis, and dynamic code instrumentation, demonstrating that the procedural rigor of executable code can effectively unlock and enhance the generalized reasoning capabili...
Jin-Yang Zhang, Wei-Bin Liao, Ke-Qin Bao et al.· 0 citations
FinIndices is a large-scale benchmark evaluating data-processing fidelity over uncropped financial statements (up to 32K tokens) and yields substantial zero-hint gains, validating that structured logic can be partially restored via data-centric alignment.
These findings suggest that, even as long-context evaluation shifts from simple retrieval toward complex reasoning, accurate grounding in relevant evidence remains an indispensable capability with substantial room for improvement.