Skip to content

Author

Xiaocheng Lu

We have 4 of 26 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Reducing Pretraining-Generation Mismatch in Diffusion Language Models

Autoregressive language models align training and use: generation conditions on a clean prompt, and training predicts future tokens from clean left context. Diffusion language models offer parallel denoising, but native dLLM pretraining can randomly corrupt prompt and continuation tokens together, weakening the clean-prefix interface needed for prompt-conditioned generation. We identify this mismatch for prompt continuation and propose PCD (Prefix-Conditioned Diffusion), a pretraining objective that combines AR prefix supervision with no-shift suffix denoising. At the training-objective level, PCD changes the attention mask, corruption mask, and label construction in continued pretraining; it does not require an autoregressive decoder, verifier, or new inference mode. By supervising the clean-prefix side autoregressively and applying diffusion only to the unknown continuation, PCD makes the local training interface resemble how block-diffusion models are queried at evaluation time. We further separate intra-sample prefix conditioning from inter-sample objective mixing, allowing us to identify the local alignment signal separately from the optional batch-level mixing knob. Across LLaDA2-Mini and Qwen-1.7B backbones, PCD consistently improves over same-family native dLLM stable baselines, reaching a 4.2% relative gain on the main LLaDA2-Mini six-benchmark average (+2.56 points) and a 14.2% relative gain in the primary Qwen mechanism comparison (+4.86 points). These results suggest that aligning the pretraining context distribution with prompt-conditioned generation can recover a measurable part of the dLLM continuation gap without changing inference.

Xiaocheng Lu, Huabin Liu, Song Guo et al. · 0 citations
Preprint Aug 2026

WorldCycle: Self-Verifiable Reinforcement Learning for Long-Horizon Video World Models

WorldCycle is introduced, a self-verifiable RL framework that constructs closed action cycles and their repeated executions from ordinary action sequences, and optimizes two complementary rewards: a spatial closure reward enforcing symmetry between mirrored forward and reverse segments, and a temporal consistency reward aligning states across repeated cycle executions.

Bohai Gu, Yueyang Yuan, Taiyi Wu et al. · 2 citations
Preprint Aug 2026

OPTD: On-Policy Transition Distillation with Consistency-Guided Adaptive Compression for Few-Step Diffusion Language Models

This work proposes OPTD, On-Policy Transition Distillation with consistency-guided adaptive compression, which consistently improves the quality--efficiency trade-off and attains the strongest overall quality-constrained AUP among the evaluated few-step baselines.

Xiaocheng Lu, Hualei Zhang, Shuhan Guo et al. · 0 citations

Edge-Efficient Compositional Recognition via Disentangled Prompt Tuning of Frozen Vision-Language Models

This work proposes Progressively Disentangled and Recurrent Prompt Tuning (PDRPT), an edge-efficient framework that decouples object and state updates before joint refinement, suppresses traction force from highly-entangled prompts, and preserves alignment with the natural language space of CLIP.

Xiaocheng Lu, Chuan He, Ziming Liu et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.