AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture Assembly
Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult. Furniture assembly requires not only recovering step-level operations from diagrammatic manuals, but also translating semantic attachm...