Pre-Compiled Pipeline Shards for Distributed LLM Inference on Intel AI PC Fleets
This paper shows that a handful of AIPCs, working together over an ordinary network, can serve models beyond the capability of any single one, and leverages speculative decoding on stateful OpenVINO models.
Tate Berenbaum, Muthaiah Venkatachalam
· 0 citations