Novel view synthesis methods, such as neural radiance fields and 3D Gaussian splatting, offer a promising solution for photorealistic rendering. However, they remain challenged in few-shot settings, where models tend to overfit the limited supervised views, leading to artifacts such as quality fluctuations, degradation in distant views, and geometric inconsistencies. To address these issues, we introduce Human Perceptual Preference Optimization (HuPPO), a framework that incorporates human perceptual guidance into model training. HuPPO mitigates distortions by regularizing training dynamics with perceptual preference cues, thereby reducing the reliance on extensive supervised views. Specifically, HuPPO leverages human perception to identify and select candidate novel views, and introduces a corresponding objective function that steers optimization toward perceptually preferred outcomes. In addition, a meta-learning pipeline is integrated to promote the learning of generalizable representations. The framework is flexible and can be seamlessly applied to a wide range of neural rendering models without incurring additional inference overhead. Extensive experiments and analyses demonstrate that HuPPO achieves consistent improvements over state-of-the-art baselines.
Xiaoyu Xu, Jiebin Yan, Sheyang Tang et al.· IEEE Transactions on Visuali...· 0 citations
High-resolution images and long videos provide vision-language models with rich context for multimodal reasoning and fine-grained perception, but the resulting long visual token sequences make large language model-side computation and memory costly. Existing visual token reducers often operate at prescribed rates, while recent methods adapt token counts across inputs using method-specific learned thresholds or importance predictors. We introduce RUTA, a principled Rate-Utility Token Allocation method that performs pre-LLM reduction by jointly learning which tokens to retain and how many to allocate to each image-query pair. RUTA constructs query-conditioned candidate tokens and predicts a retention probability for each candidate. During training, these probabilities parameterize independent Bernoulli gates, while their sum provides a differentiable training-time estimate of the token count for each pair. Retained tokens serve as anchors that aggregate information from non-retained tokens according to semantic affinity and spatial proximity. RUTA is optimized with a penalized rate-utility objective that balances downstream task loss against expected token usage. Averaged across five benchmarks and measured relative to each backbone's full-token baseline, RUTA uses only $2.0\%$ and $4.2\%$ of visual tokens while preserving $88.2\%$ and $94.4\%$ of task performance on LLaVA-NeXT-7B and Qwen3-VL-8B, respectively.
Jiangyu Zou, Xiaoyu Xu, Zhihua Wang et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.