Skip to content

Author

Sandeep P. Chinchali

We have 3 of 115 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Jul 2026

ViewMind3D: Modular View-Aware Inference for Training-Free 3D-QA

Recent advances in large language models (LLMs) and vision-language models (VLMs) have enabled new possibilities for 3D question answering (3D-QA), a key capability for embodied AI and robotic perception. However, most existing methods rely on 3D-specific training or fine-tuning with costly annotations, limiting their scalability and real-world applicability. We present \textbf{ViewMind3D}, a fully training-free and modular framework for 3D spatial reasoning over multi-view observations of a scene without requiring complete 3D reconstruction. The framework decomposes the 3D-QA task into four interpretable components: (1) question-driven multi-view selection, (2) guided visual grounding with language-conditioned object cues, (3) spatial context encoding via a bird's-eye-view (BEV) viewpoint indicator, and (4) structured answer generation through role-based reasoning. This design enables structured, robust, and interpretable reasoning without requiring model tuning. Experimental results on ScanQA and SQA3D show that ViewMind3D achieves competitive performance compared to prior training-free and fine-tuned 3D-LLMs. In particular, our method improves performance on spatially grounded question types, such as ``What''questions in SQA3D, while maintaining strong overall accuracy (50.8\%) and achieving 73.4 CIDEr on ScanQA. These results demonstrate that effective 3D reasoning can be achieved through modular orchestration of general-purpose LLMs and VLMs for robotic perception in real-world environments.

Ping-Kun Chiang, Kun-Ru Wu, Po-han Li et al. · 0 citations
Jul 2026

Incentivizing Vision Language Models to Search for Long Video Question Answering

VSeek is introduced, an agentic framework that transforms long-video question answering (LVQA) from a passive, single-pass perception task into a multi-turn retrieval process and proposes a novel neuro-symbolic approach that bridges open-ended natural language with discrete visual verification.

Harsh Goel, S. Sharan, Sahil Shah et al. · 1 citation
Preprint Aug 2026

Bayesian Partner Modelling enables Adaptive Replanning for LLM Coordination

BayesBeliefAgent is introduced, which pairs a hierarchical LLM planner with a Bayesian tracking module and evaluates performance using replanning efficiency and the belief-action gap: the fraction of total decisions where an agent with a correct partner estimate executes a non-complementary skill.

Harsh Goel, A. S. Ellendula, Vaishnav Tadiparthi et al. · 1 citation

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.