Skip to content

Author

Donglin Di

2 papers indexed here

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

MOJITO: Modal Joint Learning for Unified End-to-End Autonomous Driving

End-to-end autonomous driving systems commonly follow a cascaded two-stage pipeline where a perception stage compresses multi-modal sensor inputs into a compact context and a downstream planner predicts trajectories conditioned on this context. We argue that this one-way perception-to-planning interface forces sensor inputs into a compact representation, losing the fine-grained details critical for planning. Moreover, by constraining the planner to this compressed context, it is difficult to leverage the rich representations offered by modern vision foundation models. To address these issues, we propose MOJITO, a unified sensor-to-action framework for end-to-end autonomous driving built on modal joint learning. MOJITO removes the cascaded interface and instead performs block-wise Modal Joint Attention that simultaneously updates action, image, and LiDAR features, allowing the planner to directly access multi-modal features during action generation. MOJITO achieves 88.9 PDMS on the NAVSIM v1 dataset and 88.4 EPDMS on the more challenging NAVSIM v2 dataset, setting a new state-of-the-art. Extensive experiments further demonstrate strong scalability, instruction following, and diverse trajectory generation. Code and models are available at https://github.com/mumucc01/MOJITO.

Zhijing Cheng, Xuan Zhang, Donglin Di et al. · 0 citations
Open access Jul 2026

STAG: Biologically guided spatial transcriptomics prediction via hypergraph learning

Spatial transcriptomics (ST) enables spatially resolved gene expression profiling within intact tissue sections. However, its widespread adoption is constrained by the high cost and low throughput of current sequencing-based protocols. This has motivated growing interest in computationally predicting gene expression directly from routinely acquired histology images. Existing methods are largely restricted to isolated 2D tissue slices and fail to capture richer spatial relationships or structured dependencies among spot-level gene expression profiles. In this paper, we propose STAG, a dual-branch framework for gene-aware expression prediction and spatial context modeling. A Query branch predicts ST expression for an individual target spot, while a Neighbor branch acts as an auxiliary branch to model structured relationships among multiple spots. By leveraging hypergraph learning, the Neighbor branch captures higher-order spatial and molecular dependencies, enabling unified modeling of both intra-slice and inter-slice relationships. This design supports standard 2D settings (a single slice) and naturally extends to 3D scenarios when adjacent tissue sections are available. Moreover, STAG leverages gene semantic information as biological guidance by encoding gene names with a foundation model, enabling coordinated gene-aware interactions beyond independent gene prediction. STAG achieves an average gain of 5.16% in PCC@250 across six datasets. Under highly variable gene selection, STAG maintains the lowest RMSE and highest PCC@50 across three datasets. The effectiveness of the learned representations is further demonstrated in pseudo-3D prediction and downstream cancer classification tasks. Code is available at https://github.com/MCPathology/STAG.

Mingcheng Qu, Yuchuan Zhao, Guang Yang et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.