Skip to content

Author

Zuhao Yang

We have 3 of 10 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

StreamOPD: A Post-Training Recipe with Spatio-Temporal Cue Gating for Streaming Video Understanding

Streaming video understanding demands direct responses from the causally observed prefix of an unfolding video. Existing systems add inference-time memory, retrieval, and compression, yet a training-free sliding-window baseline already matches them. We therefore fix a memory-free recent-window protocol and ask how far post-training alone can go. Reinforcement learning with verifiable rewards fits this regime poorly, encouraging long ``think-then-answer''generations, while on-policy distillation (OPD) supplies dense token-level teacher supervision on student trajectories but is stable only when both models train in thinking mode. These observations lead to \textsc{StreamOPD}, a recipe combining verifiable streaming-video data, thinking-mode OPD, and instruct-mode deployment. It raises StreamingBench from $77.9\%$ to $83.9\%$---within $0.3$ points of the 9B teacher---and improves OVO-Bench excluding its hallucination-detection subtask (HLD) by $9.1$ points under unchanged inference. As a teacher-privilege extension, \emph{Spatio-Temporal CueGate (ST-CueGate)} aggregates cue-versus-no-cue teacher likelihood ratios into a group-relative response score that reweights OPD. It reaches $71.9\%$ on OVO-Bench (excluding HLD) and $64.9\%$ on Video-MME, and is the only variant that stays above the base model on all four benchmarks. Replacing the teacher with a frozen copy of the student's initial policy---on-policy self-distillation---retains most of these gains and lifts HLD to $57.0\%$, above both the untrained student and the 9B teacher, so abstention loss is not intrinsic to the recipe. We provide a transparent and reproducible reference for open-source streaming-video research.

Keming Wu, Baoyi Wang, Kaichen Zhang et al. · 0 citations
Jul 2026

Demonstrating TOFFEE: A Learned System for Synthesizing Data Agent Trajectories at Scale

This work introduces TOFFEE, a system for synthesizing high-quality data agent trajectories from given data environments via Monte Carlo Tree Search (MCTS) with adaptive model selection and cross-task prefix reuse, and shows that TOFFEE can effectively generate scalable trajectory data for complex analytical tasks across heterogeneous environments.

Ziting Wang, Yin Li, Zuhao Yang et al. · 0 citations
Jul 2026

PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models

PerceptionBench provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs, by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks and constructing an error taxonomy whose perception branch defines ten atomic perceptual capabilities.

Zichao Lin, Yifeng Xie, Bowen Qu et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.