Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has...
Yuta Oshima, Ku Onoda, Yusuke Iwasawa et al.· 0 citations
SAIL is a framework that reframes robot imitation as an iterative refinement problem capable of scaling with test-time compute, and utilizes Monte Carlo Tree Search, where each node is a complete trajectory and edges correspond to trajectory refinements.
Large language models (LLMs), when acting as agents, are expected to take observed data in context, infer the latent state space underlying the world, and leverage it for downstream prediction. However, prior work demonstrated that LLMs struggle to use representations learned in context on a graph tracking task, where...
Kohsei Matsutani, Gouki Minegishi, C. Park et al.· 0 citations
This work evaluates cross-embodiment transfer on RoboTwin 2.0 in a controlled setup, two bimanual robots demonstrate disjoint task sets, a policy is trained on all the demonstrations, and each robot is evaluated closed-loop on the tasks only the other demonstrated.
Maxime Alvarez, Renzo Caballero, T. Matsushima et al.· 0 citations
Robot demonstration datasets used to train vision-language-action policies can contain a subtle but harmful failure mode: trajectories that are behaviorally correct but paired with the wrong language instruction. We study post-hoc auditing of these Instruction–Trajectory Mismatches (ITMs). Unlike failed rollouts, ITMs...
Simon Holk, Ryosuke Takanami, T. Matsushima et al.· IEEE Robotics and Automation...· 0 citations
DREAM is presented, a framework that generates fine-tuning data for a pretrained VLA from a captured workspace and a language instruction, without requiring a task-specific human demonstration, and whether it can serve as a scalable data-collection system for the deployment workspace.
Makoto Sato, T. Matsushima, Yutaka Matsuo et al.· 1 citation
Recent work on looped language models suggests that many reasoning problems benefit from greater computational depth rather than from additional independent parameters. Existing studies, however, focus almost exclusively on Transformer backbones, leaving open whether this principle also applies to state-space language...
Zhen-Xuan Yu, Takeshi Kojima, Yutaka Matsuo et al.· 0 citations
Vision-language-action (VLA) models with billions of parameters now dominate the LIBERO manipulation benchmark, but the model capacity actually required by the benchmark remains unclear. We introduce MINERVA (MINimal Efficient Robotic Vision-Action policy), a family of deliberately compact visuomotor policies designed...
Kohei Sendai, T. Matsushima, Yusuke Iwasawa· 3 citations· ⚡1
A training-free adaptive pruning method designed specifically for batched inference in LRMs, built on two components: periodic top-k selection over the aggregated importance scores, unaffected by the shift that aggregation induces in the activation distribution, and based on the observation that important neurons re-fi...
Yongmin Kim, Shota Takashiro, Yusuke Iwasawa et al.· 0 citations
A taxonomy of CoT is proposed consisting of Explicit CoT, which outputs all operations without aggregation, Composed CoT, which combines multiple operations into a single step, and Implicit CoT, which omits intermediate operations.
Kohsei Matsutani, Gouki Minegishi, Takeshi Kojima et al.· arXiv.org· 1 citation
On HealMed, performance declined most in low-resource languages, although the size of the gap varied markedly across languages and models, whereas many open-source and medically specialized models showed larger and less consistent gaps.
Yingjian Chen, Fan Gao, Sherry T. Tong et al.· 0 citations
This study establishes a reliable approach of data generation, training, and benchmarking, paving the way toward further bootstrapping the quality of many-to-many translation for programming languages.
Kouki Yuki, Jie Zeng, Kyoko Ogawa et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.