Existing screenshot-to-code systems face a trade-off between flexibility and controllability. Direct multimodal generation can hallucinate visible details, whereas structured pipelines reduce such errors through component-wise decomposition, predefined templates, and customized intermediate representations. These structures, however, introduce additional generative orchestration and restrict outputs to designs covered by the representation. We investigate whether selective tool grounding can improve the fidelity--efficiency trade-off of direct widget-to-code generation. We introduce \textbf{WidgetGen}, a lightweight tool-grounded framework that extracts observable text and color evidence, performs high-level layout and optional chart reasoning, and directly generates executable JavaScript XML (\emph{JSX}). This design reduces reliance on component-wise generation while avoiding a fixed UI schema. Across six multimodal models and \(1{,}000\) held-out widgets, WidgetGen outperforms direct prompting and the structured Widget2Code pipeline on most visual reconstruction metrics, with consistent gains in area, legibility, and style. Finally, reconstruction-derived image-code pairs improve six Qwen-family open-weight models across every reported metric through supervised fine-tuning. These results establish WidgetGen as a strong lightweight baseline and show that selective evidence grounding offers an effective alternative to extensive representation constraints.
Houston H. Zhang, Tao Zhang, Li Gu et al.· 0 citations
Mainstream motion predictors have achieved low average forecasting errors on large-scale benchmarks, yet rare scenarios still exhibit high uncertainty due to the coexistence of multimodal behavioral ambiguity and sparse long-tail supervision. Existing augmentation-based long-tail strategies improve robustness through loss rebalancing or unconstrained trajectory synthesis, but some pipelines decouple three dependent decisions: where to augment, which synthesized trajectories are behaviorally valid, and how augmented signals should be assimilated in multimodal training. Such decoupling can increase augmentation cost and reduce efficiency, while introducing support mismatch between generated samples and supervision assimilation. We propose ProtoAug, a training-time augmentation framework that jointly optimizes these decisions. ProtoAug first performs prototype-based uncertainty mining in a prototype-structured latent space to allocate augmentation budget to high-value low-confidence regions. It then conducts consistency-aware augmentation via trajectory-vocabulary retrieval, map and interaction feasibility filtering, and contrastive behavior scoring for scene-conditioned candidate selection. Finally, it introduces query-wise multisupervision with warm-up-controlled optimization, assigning original and augmented trajectories to matched prediction queries to improve mode coverage while preserving optimization stability. ProtoAug is model-agnostic and removes all auxiliary components at inference, thus keeping the original deployment complexity unchanged. Experiments on Argoverse 2 and Waymo with multimodal baselines show improved overall forecasting quality and stronger long-tail robustness.
Ziheng Lu, Yingfeng Cai, Hai Wang et al.· IEEE Internet of Things Jour...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.