Skip to content

Author

Chi-man Pun

5 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

An Interactive Agent for Requirement-Driven Candidate Sourcing

Finding people from a natural-language description (``ML engineers transitioning to research roles in biotech'') is increasingly delegated to LLM agents and framed as information retrieval. We argue that it is fundamentally a requirements engineering task: such a request is an under-determined requirement with implicit constraints, many valid answers, and no acceptance criterion, so useful answers require eliciting, validating, and verifying the requirement before search can matter. We present \sys{}, to our knowledge the first interactive, requirements-driven candidate-sourcing agent (it elicits, validates, retrieves, and verifies a vague people-request into a justified slate through bounded elicitation, workflow templates, a two-stage commit protocol, and bidirectional termination guards) and \bench{}, a benchmark that runs the requirements lifecycle (criteria-anchored validation, multi-model evidence-grounded oracle construction, and cost-aware verification). Across $21$ systems and all $691$ requirements, \sys{} dominates breadth ($100%$ coverage at $2.5\times$ the yield) and is \emph{near-orthogonal} to the field, with $90%$ of the people it returns are surfaced by \emph{none} of $20$ strong LLM-plus-web baselines combined. Beyond breadth, an evidence-grounded judging of every system shows \sys{} \emph{recalls} the most relevant real people: $0.241$ of the union pool, $1.9\times$ the next system, with a bootstrap $95%$ interval disjoint from every baseline. \sys{} is thus the strongest \emph{sourcing} engine (the deepest real, reachable candidate pool), while precision-ranking LLMs serve as~complementary verifiers.

Yuanpeng He, Fan Li, Xiangyu Ru et al. · 0 citations
Preprint Aug 2026

ETHead: Generating Expressive 3D Facial Animation and Head Movement from Speech

Generating expressive 3D talking heads solely from speech remains a significant challenge due to the scarcity of high-fidelity 3D data, which limits the modeling of complex emotional motion patterns. In this paper, we introduce \textbf{E}xpressive \textbf{T}alking \textbf{Head} (ETHead), a method for generating 3D facial and head motions that vividly align with the emotional content of input speech. To overcome the data limitations, we design a self-distillation framework that leverages large-scale 2D talking videos to pre-train a specialized speech encoder. By incorporating a novel emotion-modulated probabilistic masking mechanism, this framework aligns speech representations with expressive visual dynamics, allowing the encoder to extract features highly correlated with facial and head motions directly from audio. These features are then leveraged to guide 3D generation, enriching input cues and providing explicit supervision through a joint speech-motion latent space. Extensive experiments demonstrate that ETHead substantially outperforms state-of-the-art methods. Furthermore, our motion-aligned speech encoder can serve as a transferable module, offering a general solution for enhancing expressiveness in other 3D talking head animation frameworks. The project page is available at https://verdure-oss.github.io/ETHead.github.io/.

Jiucheng Xie, Jiwang Zheng, Yongkang Xia et al. · 0 citations
Aug 2026

Interpretation Before Integration: LLM-Guided Multimodal Completion and Fusion Network for Survival Analysis With Incomplete Data

Multimodality survival analysis for nasopharyngeal carcinoma (NPC) holds great potential for improving prognosis prediction and clinical decision-making. However, it is challenged by structural and semantic misalignments across heterogeneous data. Structural misalignment arises from incomplete clinical records, where missing data introduce uncertainty in prediction. Semantic misalignment stems from the gap between structured modalities (e.g., clinical and radiomic features) and unstructured data such as 3-D magnetic resonance imaging (MRI), hindering effective feature integration. Existing methods often ignore missing data or compress multimodal information into scalar representations, failing to capture complex modality interactions and solve the problem of semantic misalignment. Furthermore, current completion techniques typically lack interpretability and overlook joint modeling of inter- and intra-sample correlations when dealing with structural misalignment, limiting their reliability in clinical settings. These issues are further exacerbated by over-parameterized models prone to overfitting in small-sample scenarios. To address these challenges, we propose LMCF, a large language model guided multimodal completion and fusion (LMCF) network tailored for survival analysis with incomplete data. LMCF consists of two core components: a lightweight dual-branch multimodality enhanced feature encoding (LDME) layer, which incorporates an interpretable multisource cross-modality completer (IMCC) for explainable reconstruction of missing data to resolve structural misalignment; and a large language model (LLM)-guided structure-semantic two-stream fusion (LSTF) layer, equipped with a quaternion convolution-based cross-domain adaptive attention fusioner (QCAAF) to effectively integrate features across modalities and mitigate semantic misalignment. Extensive experiments on the Cancer Genome Atlas (TCGA) and two proprietary NPC datasets [postradiation nasopharyngeal necrosis (PRNN) and nasopharyngeal carcinoma dataset (NCD)] from Sun Yat-sen University Cancer Center demonstrate LMCF’s superior performance in survival prediction and risk stratification, particularly under conditions of incomplete modalities and limited data resources.

Fen Ling, Haoming Zeng, Ming Li et al. · 0 citations
Preprint Aug 2026

EditaLive! Unified Character Video Editing for Live Streaming

Conventional video editing primarily focuses on scene-level content, whereas live streaming places greater emphasis on the human subject. However, directly applying existing video-editing methods to human-centric live streaming remains challenging, as they may introduce facial-expression inconsistencies and typically depend on multiple offline inference steps, making them unsuitable for real-time interaction. We propose EditaLive, a novel framework for real-time streaming character video editing. In detail, we start from a pretrained image animation model (Wan-Animate), which naturally decouples appearance from motion, and repurpose it as the base model for instruction-based human-centric video editing by reference frame editing and video reconstruction via the collected CharEdit-50K dataset. Besides, we adapt the model from offline bidirectional to causal streaming generation, and design an aligned self-rollout distillation strategy that compresses the model into a two-step sampler, where fixed RoPE and align forcing reduce training--inference discrepancies, and first-frame preserved sparse attention filters redundant historical information to mitigate appearance drift. Extensive experiments demonstrate that EditaLive delivers state-of-the-art editing performance with faithful preservation of facial expressions and low-latency real-time streaming inference.

Zhiyuan Li, Chi-Man Pun, Peng-Tao Jiang et al. · 0 citations
Preprint Aug 2026

Fast Test-Time Refinement for Robust Learned Image Compression

This study reveals an Asymmetric Adversarial Trajectory (AAT) property in LIC systems: transitioning from adversarial to benign regions is significantly easier than the reverse process, where adversarial examples can often be roughly recovered within only 1-2 steps.

Jiaming Liang, Chi-Man Pun, Weisi Lin · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.