What actually runs: a measurement study of language model placement and decode speed on the Apple Neural Engine
This work sweeps a 64-shape matrix of LLM primitives that varies how a computation is expressed while holding what it computes fixed, recording per-operation device support and finds that placement is a property of how a computation is expressed, not of what it computes.
A. ShahirM
· 0 citations