Skip to content

Author

Lianjie Cao

We have 3 of 12 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Book Open access Aug 2026

Beyond Monoliths: Enabling Flexible and Composable AI Systems via Memory Disaggregation

Modern state-of-the-art AI systems are increasingly built as monolithic supernodes integrating large numbers of specialized accelerators with proprietary high-bandwidth interconnects. These systems provision compute, memory, and networking resources in fixed ratios at design time. As AI workloads evolve, their resource demands increasingly diverge from these static configurations, leading to underutilization, limited scalability, and high operational cost. Composable systems based on disaggregated resources offer a more flexible alternative by allowing memory and compute capacity to be scaled independently without replicating an entire supernode. We introduce CAISA, a composable AI systems architecture that enables disaggregated memory expansion for large-scale AI workloads using CXL-based shared memory. CAISA goes beyond conventional memory disaggregation by jointly designing hardware support, runtime mechanisms, and workload mapping to provide contention-free shared-memory access across CPUs, accelerators, and CXL memory devices. Its key mechanism is workload-aware memory isolation, which maps shared memory through reserved address spaces to control data visibility, avoid hardware coherence overheads, and enable fine-grained pipelined data movement between compute and memory devices. By coupling memory isolation with software-managed data orchestration, CAISA decouples compute and memory resources while preserving efficient data movement, providing a scalable and cost-efficient foundation for future AI infrastructure.

Divya Kiran Kadiyala, Lianjie Cao, Jinsun Yoo et al. · 0 citations
Book Open access Jul 2026

Scaling Attention Beyond GPUs for LLM Inference

Beyond is presented, a drop-in runtime that integrates a smart offloading scheme to selectively identify and retain salient KV entries across continuous decoding sessions, together with a hybrid CPU–GPU attention mechanism for scalable inference.

Weishu Deng, Yujie Yang, Peiran Du et al. · 1 citation
Book Open access Aug 2026

CCSwitch: A Scalable Data Plane for Non-Blocking In-Network Collective Communication

This work presents CCSwitch, a modular switching fabric built from 4×4 non-blocking Collective Engines, a modular switching fabric built from 4×4 non-blocking Collective Engines (CEs) that combines spatial and temporal parallelism to perform reductions without accumulation buffers.

Sumukh Pinge, Hardik Soni, Bob Lantz et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.