EdgeAgent: Orchestrating On-Device LLM inference for End-User Multi-Agent Systems on CPU-GPU Unified Memory Architectures
This work presents EdgeAgent, a cross-layer inference system explicitly co-designed for edge UMA and multi-agent workloads, and demonstrates that the UMA-aware execution alone contributes a 1.29x speedup over batched speculative decoding and adding the agent-aware scheduling lifts the full EdgeAgent system to a 1.77x s...