Communication-Efficient Decentralized LLM Inference over Low-Bandwidth Distributed Nodes
BandwidthLLM, a communication-efficient framework that integrates three techniques: a bandwidth-aware layer placement algorithm that minimizes boundary-level transfer cost according to link bandwidth and node reliability; a lightweight activation compression scheme combining adaptive quantization with outlier-aware clipping and error feedback; and a semantic preservation check that automatically falls back to higher precision when compressed activations deviate beyond a calibrated threshold.