Beyond LLM Serving: Characterizing Vision-Language-Action Workloads for Embodied AI System Design
This work describes four representative VLA models on an edge GPU server and two onboard SoCs, using single-inference profiling and 43,200 closed-loop episodes, and guides joint design of VLA model architectures, hardware, and runtime policies.
Seonghun Jung, Sieun Moon, Jiyoung Jeong et al.
· 0 citations