Skip to content

Author

Yin-Yuan Zhao

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Open access Jul 2026

PAST: Prior-Aware Sparse Transformer for Micro-Expression Recognition

Micro-expression recognition (MER) has a lot of applications in lie detection, education, healthcare, etc., as involuntary micro-expressions (MEs) may provide subtle facial cues associated with affective responses. With the development of deep learning, many studies have recently employed Vision Transformers (ViTs) to investigate MER, since ViTs show promising performance in various visual domains due to their excellent local–global modeling ability. However, such methods confront two fundamental challenges: First, fine-grained visual features are needed to capture the subtle facial movements of MEs, which ViTs relatively fall short on due to coarse patch resolution constrained by their quadratic complexity. Second, the data-intensive nature of ViTs impedes effective learning given the limited scale of ME data. To overcome the aforementioned limitations of using ViTs for MER, we propose the Prior-aware Sparse Transformer (PAST), a novel Transformer-based architecture integrating spatial and semantic prior knowledge synergistically into a sparse attention mechanism, enabling linear-complexity processing of large amounts of fine-grained features. Specifically, we first designed an extraction algorithm to generate a representative set of motion-intensive Principal Anchors, which are used to guide the model’s focus on biologically critical regions during sampling. Second, we introduced the Semantic Dictionary, which was trained with a carefully designed self-contrastive loss to embed task-invariant discriminative semantics of the anchors. Such global semantics further modulate patch sampling and attention weighting in the sparse attention procedure, achieving better training performance with limited ME data. Extensive evaluations on MEGC and CD6ME protocols demonstrate state-of-the-art performance, validating PAST’s efficacy for MER.

Jiateng Liu, Tianchen Zhou, Hengcan Shi et al. · 0 citations
Preprint Aug 2026

Residual Flow Matching with Dynamic Cross-Interaction for 3D Multi-Person Motion Prediction

3D multi-person motion prediction requires modeling both individual kinematics and inter-person interactions. While Flow Matching is effective for multi-hypothesis generation to improve prediction accuracy, directly predicting skeletal sequences from pure noise often compromises structural consistency and introduces unreliable cross-agent interactions during early noise-dominated integration steps. To address this, we propose a Prior-Guided Residual Flow Matching framework. First, a Deterministic Coarse Prior (DCP) establishes a kinematic anchor, formulating the generative process as a conditional flow over motion residuals to simplify the generative objective and preserve structural stability. Second, a Dynamic Cross-Interaction (DCI) mechanism temporally synchronizes inter-agent message-passing with the integration progress, ensuring the extraction of reliable social contexts and improving multi-person motion fidelity. Finally, a decoupled joint-motion architecture with bidirectional fusion effectively preserves fine-grained kinematic coherence. Extensive experiments demonstrate that our approach achieves state-of-the-art prediction accuracy across multiple datasets. Code is available at https://github.com/Wei-Wei-a/Residual-Flow-Matching-with-Dynamic-Cross-Interaction-for-3D-Multi-Person-Motion-Prediction.

Wei Wei, Yin-Yuan Zhao, Ruixuan Yu · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.