Skip to content

Author

Yanzhi Wang

We have 6 of 20 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#computer vision Preprint Oct 2026

SpectralCache: Accelerating Diffusion-Based World Models via Spectral Feature Caching

Diffusion-based world models enable high-quality interactive environment generation but suffer from substantial inference overhead due to repeated Transformer evaluations during denoising. Existing caching methods mainly exploit temporal redundancy at the feature or token level, leaving the underlying mathematical stru...

Zhen-Dong Mi, Pu Zhao, Zi-Yu Hu et al. · 0 citations
#machine learning Preprint Oct 2026

From Patching to Pruning Visual Computation in Vision Language Models

Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a tr...

Rahul Chowdhury, Timothy Rupprecht, Xuan Shen et al. · 0 citations
#artificial intelligence Preprint Sep 2026

LEAP: Learned Block-wise Evidence Retrieval for Long Audio-Video Perception

Hour-scale audio-visual question answering is constrained by a context dilemma: dense whole-recording encoding rapidly exhausts context limits, whereas uniform temporal compression severely dilutes fine-grained acoustic and visual evidence. We introduce LEAP, a framework where the model retrieves its own evidence witho...

Ju-Yi Lin, Zhi-Qiang Lao, Jia-Li Cui et al. · 0 citations
Preprint Aug 2026

From Retrieved Context to Runtime Control: Adaptive Compression for Edge-based RAG

This paper proposes a vision for telemetry-informed adaptive compression in edge RAG, grounded in experimental evidence, and argues for runtime policies that dynamically manage compression, guided by workload features and edge telemetry.

Z. Feric, Amir Taherin, Yan-Zhi Wang et al. · 0 citations
Preprint Aug 2026

Fine-Tuning Qwen3-27B for C-to-Rust Code Translation: A Three-Stage Curriculum of Pretraining, Debugging-Aware SFT, and Task-Specific SFT

A three-stage fine-tuning curriculum applied to Qwen3-27B is described that is designed to progressively specialize the model for the C-to-Rust (C2Rust) translation task, and the resulting model is evaluated using the agentic, static-analysis-guided verification framework of SACTOR.

Pu Zhao, Chang-Di Yang, Yi-Xiao Chen et al. · 0 citations
#artificial intelligence Preprint Aug 2026

Mechanistic Interpretability of Structure-Aware Numerical Reasoning in LLaMA 3.1 8B

Recent work has shown that large language models (LLMs) exhibit strong numerical sequence modeling capabilities and show promise in time-series prediction. While LLMs display in-context learning capabilities, the mechanisms with which they accomplish time-series prediction remain unclear. Specifically, whether they tru...

Rahul Chowdhury, Timothy Rupprecht, Senhao Cao et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.