Results indicate that with zero curated data, OPT-Zero matches state-of-the-art data-dependent methods while exhibiting substantially stronger generalizability, establishing self-play training as a highly scalable paradigm for advancing LLM reasoning in modeling and solving optimization problems.
Xia Jiang, Yao-Xin Wu, Chen-Yu Zhou et al.· 0 citations
Predicting the remaining useful life (RUL) is essential for effective predictive maintenance. Spatio-Temporal Graph Neural Networks (ST-GNNs), which can model both temporal and spatial relationships by representing time series data as a sequence of graphs, have shown exceptional performance in RUL prediction. However,...
Ya Song, Laurens Bliek, Yao-Xin Wu et al.· 0 citations
Conservative Discrete Quantile Actor-Critic (CDQAC), an offline RL algorithm that learns effective scheduling policies directly from static, suboptimal datasets, and is highly sample efficient, requiring only 1 to 5% of the original dataset to learn high-quality policies.
Jesse van Remmerden, Z. Bukhsh, Ying-Qian Zhang· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.