FusionML: Prefill, Not Decode - Mechanism and Boundaries of CPU+GPU Co-Execution on Unified-Memory Apple Silicon
A per-layer, contention-aware CPU+GPU row split for transformer prefill built on a fix for MLX's lazy-graph scheduler that accelerates Llama-shaped decoder-block prefill and saves time-to-first-token on a real Qwen2.5-7B checkpoint.