Skip to content

Author

Yuta Oshima

We have 4 of 12 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

VIF-Bench: Evaluating Visual Instruction Following in Multi-Reference Image Generation

Recent multimodal image generation models can take multiple images and textual instructions as input, enabling reference-based generation guided not only by text but also by visual instructions such as layouts, arrows, and pose cues. However, existing benchmarks do not evaluate the joint setting in which multiple refer...

Yuta Oshima, Masakazu Yoshimura, Masahiro Suzuki et al. · 0 citations
#artificial intelligence Preprint Sep 2026

AutoRef: Harness Optimization for Agentic Multi-Reference Image Generation

Recent image generation models can take multiple reference images as input and combine them into a new image. However, multi-reference image generation remains challenging: models may omit or duplicate subjects from the references, or produce images in which multiple subjects appear unnaturally pasted. Recent work has...

Yuta Oshima, Ku Onoda, Yusuke Iwasawa et al. · 0 citations

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation

This work formalizes target-mode coverage as target-mode coverage, the coverage of a predefined set of semantically specified modes, and proposes multi-axis max@K, a group-based reinforcement learning objective for improving it in diffusion-based T2I models.

Ku Onoda, Paavo Parmas, Hiroki Furuta et al. · 3 citations
Open access Mar 2024

SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces

The ablation study shows that when using SSMs for temporal modeling, incorporating bidirectionality and selective scans enhances video generation performance, and SSM-based models incur lower computational cost to achieve the same Fréchet Video Distance as attention-based models.

Yuta Oshima, Shohei Taniguchi, Masahiro Suzuki et al. · 13 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.