Skip to content

Author

Yiming Zhong

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

KeyID: Decoupled Drafting and Keyframe Editing for Identity-Preserving Video Generation

Identity-preserving video generation (IPVG) requires synthesizing videos that are faithful to both reference subjects and text prompts. Existing methods are often hindered by high tuning costs or limited input-level enhancements, struggling to maintain rigid identity consistency during complex, long-sequence actions. To address these limitations, we propose KeyID, a training-free IPVG framework that decouples the synthesis of video dynamics from the injection of identity. Specifically, KeyID comprises two components: (1) Reference-Aware Video Generation, which produces an identity-agnostic video draft aligned with multiple references, and (2) Identity-Preserved Keyframe Editing, which integrates the target identity via sparse keyframe correction and subsequent motion interpolation. By shifting from dense frame-level supervision to sparse keyframe-level refinement, KeyID effectively resolves the capacity conflict between prompt adherence and identity fidelity. Crucially, our modular design allows seamless extension to multi-subject references and complex sequential action generation without additional training. KeyID outperforms prior works and is validated by automatic and human evaluations on the official challenge benchmark, ultimately securing the runner-up position in the Track 2 (Sequential Action) of the ACM Multimedia 2026 IPVG Grand Challenge. Source code is available at https://github.com/WISLab-GDUT/KeyID.

Jianjie Luo, Yiming Zhong, Hao Shen et al. · 0 citations
Conference Aug 2026

Boosting defocus deblurring via learning from all-in-focus images

Defocus deblurring is a challenging task due to spatially varying blur and limited aligned training data. Existing datasets suffer from insufficient scene diversity and misalignment between defocused and all-in-focus images, restricting network performance. Additionally, single-stage autoencoders often fall into local optima, causing under-recovery and artifacts. To address these problems, we propose a novel multi-stage restoration framework guided by information from a single all-in-focus image. First, rendering synthesis adds defocus attributes to all-in-focus images, solving data alignment and consistency issues. Second, a stacked autoencoder guided by defocus degree maps handles spatially varying blur hierarchically. Finally, Feature Selection and Feature Attention Modules discriminatively select valuable information and transmit first-stage features to later stages for better region-wise deblurring. Extensive experiments on multiple test sets validate that our method achieves state-of-the-art performance both quantitatively and qualitatively. Specifically, on the DPDD dataset, our method achieves 29.41 dB PSNR and 0.886 SSIM, outperforming the previous best method by 0.19 dB; on the RealDOF dataset, it achieves 23.65 dB PSNR, surpassing IFANet by 0.89 dB.

Yiming Zhong, Jifeng Chen · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.