Skip to content

Author

P. Zhou

We have 8 of 31 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

CUA-Sandbox: Efficient Environments for Computer-Use Agent Reinforcement Learning

CUA-Sandbox is introduced, which separates private state capsules from shared runtimes through state-scoped execution and transactional lifecycle operations, including resets and branches, while retaining the original software interfaces and task evaluators.

Xin Yan, Zheng-Bo Jiao, Jia-Qi Liu et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Beyond One-Size-Fits-All: Sample-Adaptive Strategy Routing for Vision Token Pruning in MLLMs

Multimodal large language models (MLLMs) process hundreds or thousands of visual tokens per image, incurring prohibitive inference costs. While existing vision token pruning methods mitigate this overhead, they implicitly assume that a single fixed pruning strategy can be applied uniformly across all inputs. Our analys...

Hai-Jin Liang, P. Zhou, Zheng-Lin Wan et al. · 0 citations

Perceive Before Reasoning: A Pre-Reasoning Perception Framework for Efficient and Reliable Proactive Mobile Agents

The Pre-Reasoning Perception Framework (PRPF) is proposed, a two-stage framework built on perceiving before reasoning that substantially reduces false trigger rates (FTR) while improving success rates (SR) and inference efficiency over the ProactiveMobile baseline.

Zhi-Jie Ding, Wei-Nan Hong, Zichen Zhu et al. · 0 citations
Preprint Aug 2026

Agentic Game Development as a Verifiable Trajectory Data Engine for Scaling World Models

This work proposes Reinforcement Learning with Human-Engine Verification (RLHEV), a post-training paradigm that combines dense engine signals with implicit human acceptance feedback from the development process to support RL post-training.

P. Zhou, Hesong Wang, Zhengfeiyang Zhang et al. · 0 citations
Preprint Aug 2026

Improving Generalization Robustness of Multimodal RLVR

Reinforcement Learning with Verifiable Rewards (RLVR) makes Multimodal Large Language Models more accurate, but the gains are brittle: simply paraphrasing a question or changing the prompt template can degrade them, which challenges reliable deployment in high-stakes scenarios like medical VQA. We trace this to two iss...

P. Zhou, Zhiwei Tang, Xiaopeng Peng et al. · 0 citations
#artificial intelligence Preprint Aug 2026

SpikeOPD: Stable On-Policy Distillation for Autoregressive Spiking Language Models

Spiking neural networks (SNNs) offer a path to energy-efficient language modeling through sparse encoding and event-driven computation, but training capable spiking language models from scratch remains difficult. A practical alternative is ANN-to-SNN migration through knowledge distillation (KD), where a pretrained art...

En-Qiao Lu, Xingrui Yu, Yi-Wei Fu et al. · 0 citations
Preprint Aug 2026

SCOPE-Router: Cost-Aware Open-Set VLM Routing for Execution-Oriented Tasks

CRM+RCCR, an architecture-agnostic cost-aware objective that encodes cost preference into continuous relevance targets through per-pair independent scoring, eliminating multi-positive dilution while regularizing queries with similar routing preferences to be closer in the routing space.

Tao Yu, Yi-Fei Qu, Zhi-Qing Cui et al. · 1 citation
Preprint Jul 2026

Unified Hallucination Fuzzing for Multimodal Large Language Models

This work introduces UniHall, a fine-grained dataset grounded in a unified taxonomy spanning Object, Instruction, and Knowledge dimensions, and proposes Self-Adaptive Multimodal Fuzzing (SAMF), a self-adaptive framework that employs evolutionary mutation strategies to explore the boundaries of model hallucinations.

P. Zhou, Jiajun Song, Zhiwei Tang et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.