Skip to content

Author

Ming-Kai Zheng

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#machine learning Preprint Sep 2026

Broken Symmetry in BF16 Attention: Why FlashAttention Gradients Blow Up Late in Training

BF16 is now standard in large-scale pretraining, including in fused attention kernels such as FlashAttention, and these kernels are widely trusted. When we used FlashAttention-3 to pretrain a 450M-parameter transformer on 50B tokens, however, we ran into a problem: training was healthy for 25B tokens, then the gradient...

Jun-Lin Chen, Daize Dong, Huan-Wei Di et al. · 0 citations
#artificial intelligence Preprint Sep 2026

IndustrialVLA-Bench: A Traceable Multi-Axis Evaluation of Open Robot Policy Models

IndustrialVLA-Bench is presented, an evidence-aware evaluation of six released VLA and WAM systems under a unified reporting schema that evaluates clean capability on LIBERO, non-language robustness on LIBERO-Plus, instruction sensitivity on LIBERO-Para, and observed execution cost.

Yi-Qi Wang, Zhi-Feng Rao, Jia-Qi Zhang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.