Propensity Straight-Through Gradients for Discrete Stochastic Systems
This work exploits the affine state update to obtain the exact one-step conditional-mean sensitivity by differentiating normalized reaction propensities, and defines the propensity straight-through (PST) estimator, a temperature- and Gumbel-free path to scalable gradient-based learning through exact stochastic trajectories.