HotBa: A Heterogeneous Mamba Accelerator with Δ-Guided Early Rejection for Speculative Decoding
HotBa is presented, a heterogeneous Mamba accelerator that reduces per-token weight transfer by 82% and redundant computation by 59% via Δ-guided early rejection for wide-tree speculative decoding in Mamba, while a heterogeneous INT8/FP16 core and tree management unit achieve 40.4× area efficiency and 5.18× SSM speedup.