CritICL: Inference-Time Weak-to-Strong Generalization from Small Language Model Failure Modes
Experimental results show that CritICL consistently outperforms standard in-context learning and achieves performance competitive with or superior to test-time scaling methods, while requiring significantly fewer generations and lower token cost.
Yu-Fan Wu, Ying-Hui He, Zheng-Yi Hu et al.
· 1 citation