Wet-lab experimentation serves as the gold standard for hypothesis verification in scientific discovery; yet it is inherently labor-intensive, costly, and safety-critical. Embodied agents hold the promise of automating these tedious workflows, but their development is hindered by the scarcity of real-world training dat...
Chen-Xi Li, Hai-Yuan Wan, Rui Li et al.· 0 citations
Current evaluations and training of multimodal models predominantly focus on multi-image tasks, largely overlooking interleaved text-image scenarios. In such multi-image tasks, text typically serves merely as task instructions, lacking deep semantic interaction with the visual content. In contrast, realworld applicatio...
Zi-Hao Wang, Xinyue Xiang, Yuwen Sun et al.· 0 citations
4DHumanDiff is presented, a diffusion framework that directly generates dynamic humans represented by 4D Gaussian Splatting from text prompts and achieves better temporal and multi-view consistency, and reduces inference time by more than 10x.