Reducing Scalar Rewards to Binary Success: General Off-Policy Learning with Success Functions
It is shown that any discounted-reward problem can be recast as a modified, reward-free problem with a single success state and an equivalent optimal policy, and any value-learning problem can be trained by classification with cross-entropy, in a way that is in principle exact and suffers no loss of precision.