Hardware-Adaptive Inference Orchestration: Zero-Overhead Local LLM Agents on Resource-Constrained Edge Devices
Autonomous AI agents typically rely on multi-turn ReAct loops that demand repeated system-prompt evaluation, persistent state tracking, and frequent tool selection. On constrained edge hardware—especially CPU-only devices with approximately 8 GB of RAM—this style of orchestration creates two compounding failure modes:...