Aurora: A Disaggregated GPU-PNM-PIM System for High-Throughput Mixed-Length LLM Inference
The proposed Aurora is a GPU–PNM–PIM disaggregated system designed to efficiently serve mixed-length LLM inference under ILGA, which introduces an ILGA-aware multi-path PNM-PIM pipeline that explicitly accounts for block-level heterogeneity and request-length diversity, improving pipeline utilization without overprovisioning tensor parallelism.