Skip to content

Author

Shimin Di

We have 5 of 45 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

CoRe-VLA: Preserving Cross-View Coordination in VLAs under Camera Shifts

VLAs combine pretrained vision-language representations with action generation to enable language-guided control across diverse tasks, becoming a mainstream paradigm in embodied intelligence. However, multiple studies have reported VLA's substantial declines in task success under camera shifts, revealing a key vulnerab...

Tian-Hang Pan, Xuan Wang, Yi-Wen Pang et al. · 0 citations
#artificial intelligence Preprint Sep 2026

TokenCast: Forecasting Token Consumption During LLM Agent Execution

TokenCast is proposed, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces, and captures the extra input cost incurred when context from earlier segments is re-read by every later call.

Chao-Qian Ouyang, Ling Yue, Li-Bin Zheng et al. · 0 citations
Jul 2026

DREvo: Distilling Recalibrated Historical Experience for Harness Self-Evolution

A new harness self-evolution method, named DREvo, is proposed, which integrates function-level evidence anchoring, state-dependent evidence recalibration, and role-conditioned search intent distillation to determine which historical evidence remains valid and where the harness should evolve next.

Hanghui Guo, Wei-Jie Shi, Zhangze Chen et al. · 4 citations
Preprint Aug 2026

UrbanAgent: A Tool-Augmented Agent for Cross-System Urban Tasks

This work proposes Urban-Agent, a tool-augmented agent framework for cross-system urban tasks that couples the cognitive and reasoning capabilities of a large language model with a tool-set supporting code execution, API calls, and Model Context Protocol.

Jiayu Cao, Xing-Yuan Zeng, Fei-Yue Li et al. · 0 citations
Preprint Jul 2026

AI-generated Images Challenge Visual Trust in High-risk Scenarios

SafeIMG is introduced, a safety-oriented benchmark spanning 12 public- and individual-safety scenarios generated using GPT Image 2.0 that provides human annotations that localise suspicious regions and explain local artefacts and higher-level commonsense or physical inconsistencies.

Yizhi Wang, Yichen Xiao, Linan Yue et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.