Language models are moving beyond generating answers to pursuing long-horizon goals in interactive environments. Post-training these agents requires long, heterogeneous trajectories, and synchronous systems leave learner engines idle until rollout and verification finish. To squeeze out these pipeline bubbles, asynchro...
Seil Kang, Hangoo Kang, Tarun Suresh et al.· 0 citations
Decoding and interpreting naturalistic speech from the brain increasingly relies on alignment to pretrained speech and language representation spaces. However, current CLIP-style brain-speech alignment ground neural activity to a single anchor modality-audio or text-despite the brain's inherently multimodal speech proc...
Su-hyun Kim, Jin-Mo Han, Danny Dongyeop Han et al.· 0 citations
QWERTY is introduced, a training-free framework that enables flexible motion control in pretrained image-to-video DiTs via user-defined object warping and optical flow and achieves the most effective motion control among existing training-free approaches on a recent image-to-video DiT, with performance comparable to fi...
K. Choo, Young Min Kim, Hyunkyung Han et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.