Skip to content

Author

D. Paudel

We have 13 of 205 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Sep 2026

CoRef-GS: Cooperative Referring Gaussian Splatting for Multi-Agent Scene Understanding

Referring scene understanding for embodied robots requires grounding object- and relation-centric language queries from a designated viewpoint. While a local semantic Gaussian map can support such grounding within one agent's observations, cooperative settings require this ability to remain effective after independentl...

Zhi-Kun Zhou, Kun-Yu Peng, Run-Yi Yang et al. · 0 citations
Preprint Sep 2026

AHMAD: Adaptive Hybrid Multi-task Vision Learning with Assisted Distillation for Keypoint Detection

Generalist multitasking vision models aim to unify multiple vision tasks within a single framework, enabling more efficient and versatile learning. However, handling diverse vision tasks -- spanning dense and sparse predictions -- remains challenging due to their inherently varying output structures. In this paper, we...

Mohammad Mahdi, Nedyalko Prisadnikov, Yu-Qian Fu et al. · 0 citations
#artificial intelligence Preprint Oct 2026

iADD: Improving Alignment and Diversity in Diffusion Policy Optimization

Reinforcement learning based post training of diffusion models, such as Denoising Diffusion Policy Optimization (DDPO), optimizes a reverse diffusion process under a reward function. However, current approaches to reward optimizations do so at the cost of diversity and quality. In this paper, we provide better tradeoff...

Ashok Prasad Neupane, Saugat Adhikari, Pramish Paudel et al. · 0 citations
Preprint Sep 2026

Ego-Forge: Text and Geometric-Attention Free Exo-to-Egocentric Video Generation

Ego-Forge is a caption-free and bias-free framework for exo-to-egocentric video generation that achieves state-of-the-art performance on Ego-Exo4D, runs faster end-to-end, requires no external annotation at inference, and generalizes to in-the-wild scenes, including cases where over-reliance on geometry blocks appearan...

Mohammad Mahdi, L. Gool, D. Paudel · 0 citations
#artificial intelligence Preprint Sep 2026

Does Local Video Understanding Transfer Across Encounters? The EgoGears Benchmark

Embodied systems must make knowledge acquired during one encounter usable in another despite changes in viewpoint, motion, and illumination. Yet aggregate cross-video accuracy conflates failures of local perception with failures to preserve observation identity, establish correspondence, and compose evidence, obscuring...

Yue-Dong Tan, Lei-Tao Qi, Yu Liu et al. · 0 citations
Preprint Sep 2026

GraphWrit3R: End-to-End 3D Scene Graph Writing

This work presents GraphWrit3R, a simple end-to-end method that takes a 3D point cloud, Gaussian Splats, or a combination of both as input, and directly outputs a complete scene graph as a structured JSON script, achieving state-of-the-art performance on object class, predicate, and triplet recall on the 3DSSG benchmar...

Luka Milivojevic, Nikola Popovic, Sayan Deb Sarkar et al. · 0 citations
Preprint Aug 2026

iARCS: Iterative Agentic RL for Controllable 3D Scene Generation

iARCS is presented, an iterative agentic reinforcement learning framework that adapts a pretrained scene generator to naturallanguage task requirements and shows that data generated by iARCS improves a base generator, supporting its value as a practical synthetic data generation tool rather than only a controllable sce...

Saugat Adhikari, Ashok Prasad Neupane, Pramish Paudel et al. · 1 citation
#artificial intelligence Preprint Sep 2026

SceneBench: A Hierarchical Benchmark for Vision-Language Understanding of 3D Scenes

SceneBench is introduced, a benchmark of 966 photorealistic 3D scenes reconstructed with Gaussian Splatting and densely annotated with hierarchical semantics spanning scenes, rooms, functional areas, object groups, and individual objects that provides a realistic testbed for developing and evaluating models capable of...

Anubhav Khanal, Prabigya Acharya, Roshni Poudel et al. · 0 citations
#artificial intelligence Preprint Sep 2026

Kairos: A Dataset for Fine-Grained Video-Language Modeling over Space, Time, and Dynamics

Kairos is introduced, a video dataset for video-language modeling with time-resolved annotations that supports fine-grained evaluation, long-range modeling and reasoning, instruction data construction, representation learning, and video generation.

Ruibo Ming, Lei Sun, De-Heng Zhang et al. · 0 citations
#artificial intelligence Preprint Aug 2026

FRAMEWORKERS: A Dynamic Multi-Agent Framework for AI-Generated Video Production

Experiments show that FRAMEWORKERS outperforms strong LLM planners in routing accuracy, recovers reliably from runtime failures, generalizes to unseen sub-agents without retraining, and achieves higher end-to-end video quality and broader task coverage than fixed pipelines, single-agent systems, and prior multi-agent a...

Zhen-Dong Li, Lei Sun, Le-Tian Shi et al. · 1 citation
Preprint Aug 2026

Event-Based Motion Estimation via Oriented Distance Fields

Oriented Distance Field Motion Estimation (ODF Motion Estimation), which replaces this optimization with a single averaging step over a precomputed field of event distance vectors, combined with an adaptive event-count selection strategy and a parameter-free trail filter, reaches sub-pixel accuracy at the lowest latenc...

Lei Sun, Yuqin Ma, Weilun Li et al. · 0 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.