A world model learns to forecast how a physical system evolves from recorded trajectories, yet the systems it imitates obey physical laws that are neither fully supplied nor reliably respected. The model may create energy, drift or diverge over long rollouts, and answer a changed law query using the law observed during...
Yu-Feng Wang, Parivesh Priye, Lu Wei et al.· 0 citations
How many directions in weight space does training need? The intrinsic dimension answers this with the smallest number of random directions in which training still reaches a target accuracy, and small values have motivated parameter-efficient methods such as LoRA. We measure it for variational Monte Carlo (VMC), which t...
Lu Wei, Yu-Feng Wang, Chen-Feng Cao et al.· 0 citations
A selective predictor acts as a safety gate: it returns an output only when the prediction appears sufficiently trustworthy. Deployments increasingly require this reliability to be certified at a target precision for every reporting unit of interest, such as a tool, policy label, or patient subgroup. The main difficult...
Parivesh Priye, Yu-Feng Wang, Hai-Bin Ling et al.· 0 citations
A concrete design principle for physical world models: long-horizon stability and changed-law generalization arise from distinct structural commitments, and each can be imposed deliberately without requiring the other.
Yu-Feng Wang, Parivesh Priye, Lu Wei et al.· 0 citations
GRPO-QPS is introduced, a target-preserving framework in which GRPO learns proposal behavior and an exact Metropolis correction preserves the posterior after training, which combines target-preserving Bayesian inference with broad gains over learned transport baselines and a sampling advantage when efficient exploratio...
Yu-Feng Wang, Parivesh Priye, Lu Wei et al.· 0 citations
These results are preliminary and use LLM judges rather than human domain experts, but they support a narrow scientific-discovery claim: explicit derivation constraints are a promising mechanism for making LLM-generated scientific questions more auditable.
The diagnosis prescribes the fix: keep the goal out of the dynamics and supervise the \emph{read} path, recovering genuine, instruction-independent grounding, and the detection protocol and remedy apply to any goal-conditioned world model whose instruction names the scored quantity.
Yufeng Wang, Lu Wei, Haibin Ling· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.