Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent's observations, cooperative settings require this ability to remain effective after independentl...
Zhi-Kun Zhou, Kun-Yu Peng, Run-Yi Yang et al.· 0 citations
Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning. However, handling diverse vision tasks -- spanning dense and sparse predictions -- remains challenging due to their inherently varying output structures. In this paper, we...
Mohammad Mahdi, Nedyalko Prisadnikov, Yu-Qian Fu et al.· 0 citations
Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoff...
Ego-Forge is a caption-free and bias-free framework for exo-to-egocentric video generation that achieves state-of-the-art performance on Ego-Exo4D, runs faster end-to-end, requires no external annotation at inference, and generalizes to in-the-wild scenes, including cases where over-reliance on geometry blocks appearan...
Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring...
Yue-Dong Tan, Lei-Tao Qi, Yu Liu et al.· 0 citations
This work presents GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script, achieving state-of-the-art performance on object class, predicate, and triplet recall on the 3DSSG benchmar...
Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar et al.· 0 citations
iARCS is presented, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to naturallanguage task requirements and shows that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable sce...
SceneBench is introduced, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects that provides a realistic testbed for developing and evaluating models capable of...
Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al.· 0 citations
Kairos is introduced, a video dataset for video-language modeling with time-resolved annotations that supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation.
Ruibo Ming, Lei Sun, De-Heng Zhang et al.· 0 citations
Experiments show that FRAMEWORKERS outperforms strong LLM planners in routing accuracy, recovers reliably from runtime failures, generalizes to unseen sub-agents without retraining, and achieves higher end-to-end video quality and broader task coverage than fixed pipelines, single-agent systems, and prior multi-agent a...
Zhen-Dong Li, Lei Sun, Le-Tian Shi et al.· 1 citation
Oriented Distance Field Motion Estimation (ODF Motion Estimation), which replaces this optimization with a single averaging step over a precomputed field of event distance vectors, combined with an adaptive event-count selection strategy and a parameter-free trail filter, reaches sub-pixel accuracy at the lowest latenc...
A cross-domain egocentric video question answering benchmark designed to evaluate whether multimodal large language models can generalize beyond common daily-life scenarios, and two official Codabench tracks.
Yu-Qian Fu, Tianwen Qian, Yanjun Li et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.