Automated Tensor Scheduling for Hybrid CPU-GPU LLM Inference on Consumer Devices
ATSInfer combines static tensor placement with load-aware dynamic transfer and introduces asynchronous CPU-GPU coordination to efficiently schedule hardware storage, data movement, and computation across heterogeneous backends and can substantially improve the user experience of local LLM deployment on personal consumer devices.