EgoAfford: Task-Oriented Affordance Grounding via Egocentric Referring Segmentation
EgoAfford is introduced, a benchmark designed to connect the semantic roles of participating objects with task-state-aligned visual observations and multi-step planning and EgoLens, a 3B multimodal large language model with role-specific mask decoders, as an in-domain reference model for this joint task.