A seller and a buyer with independent private values can trade only at a posted price. We determine the worst case of this mechanism exactly: the best posted price always guarantees a $\beta_*=0.738024\ldots$ fraction of first-best welfare, where $\beta_*$ is given in closed form by the root of an explicit equation, th...
Ting-Yi Lin, Yi-Chen Shi, Ke Dong et al.· 0 citations
Imitation learning for robotics depends on human demonstrations, some of which people may later ask to remove. Retraining without them is the natural reference, but its cost grows with policy and dataset scale, motivating cheaper operators that edit a trained policy. Metrics inherited from machine unlearning, such as f...
Jia-Zhuo Li, Yu Zhang, Yi-Ming Fei et al.· 1 citation
The Dreamer-SAC framework, which integrates a recurrent state-space world model with an off-policy soft actor-critic algorithm trained directly in latent space, uses a combination of real interactions and short-horizon generated trajectories with n-step target estimation and multi-objective supervision.
Jia-Zhuo Li, Lin-Jiang Cao, Qi Liu et al.· 0 citations
In latent dynamics, subtracting a prediction's mean effect over actions cancels whatever the actions share--the action-independent variation where distractors live--leaving a clean, controllable channel, with no reward, no reconstruction, and no distractor-specific auxiliary loss.
Jiazhuo Li, Yiming Fei, Zhiruo Zhou et al.· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.