Skip to content

Author

Yike Guo

18 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

Learning Ophthalmologist Clinical Reasoning for Glaucoma Diagnosis from Fundus Images

A reasoning-driven vision-language framework that explicitly models the ophthalmologist's diagnostic workflow by generating structured clinical reasoning prior to diagnosis is developed, demonstrating that explicitly modeling expert clinical reasoning simultaneously improves interpretability and diagnostic performance.

Kai-Chen Zhou, Yuzhen Chen, Elif Yildiz et al. · 0 citations
Jul 2026

ReBind: Multi-Reference Video Editing via Structured Instructions with Explicit Reference Relationships

This work proposes ReBind, a systematic framework that introduces semantic instructions with embedded reference tokens as the intermediate representation for multi-reference image-conditioned video editing and develops ReBind-Instruct, a specialized MLLM that learns to establish explicit bindings between visual attribu...

Xin-Yu Liu, Shi-Hao Li, Weihong Lin et al. · 2 citations
Preprint Aug 2026

SkillProx: Self-Evolving Agent Skills via Proximal Textual Gradient Descent

SkillProx is introduced, a proximal-gradient-inspired forward--backward framework that couples closed-loop diagnostic evolution with utility-aware proximal refinement and demonstrates the complementary effects of closed-loop diagnosis and proximal refinement.

Mingxuan Zheng, Yu-Jin Zhou, Chuxue Cao et al. · 2 citations
Preprint Aug 2026

UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding

Experimental results indicate that fully automated data curation combined with imbalance-aware training yields non-trivial improvements, but models still struggle to capture fine-grained acoustic features, indicating a gap between surface-level alignment and deep musical comprehension.

Zi-Ya Zhou, Shangda Wu, Shenyang Xu et al. · 0 citations
Preprint Aug 2026

Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

FISA is proposed, a framework for MLLM self-improvement that constructs augmented images from the model's own failure cases that generates visually challenging yet answer-preserving image complications, verifies their utility through self-examination, and applies dual fidelity filtering to avoid semantic distortion.

Chun-Yang Jiang, Pingping Zhang, Yuzhi Zhao et al. · 0 citations
Jul 2026

A Control Theory of Predictability in Latent World Models

It is proved that the planner's suboptimality is bounded by twice this discrepancy between the predicted and the true plan-cost at the plan the planner commits to, whereas the data-averaged prediction error neither bounds nor tracks it.

Hanzhe You, Yonggang Zhang, Maohao Ran et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.