The results demonstrate the feasibility of a stateful contactless sensing-to-action architecture for long-term home health monitoring and integrates sensing, temporal state, reasoning, and action into a unified and auditable loop.
Xu-Wen Zhang, Zi-Jian Lu, Yi-Cheng Lei et al.· 0 citations
Vision-language models often use descriptions of earlier visual states to make decisions about the current scene. When the scene changes, stale language can redirect an otherwise correct visual judgment toward an outdated answer. We study this failure as visual lock-in in a controlled grounding setting where only the v...
VITAL-RAG is introduced, which organizes evidence by canonical code object, keeps one query-relevant companion only when it adds semantics not already represented, and renders selected evidence under per-object and global token budgets.