Skip to content

Author

Abolfazl Hashemi

We have 2 of 11 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

Preprint Aug 2026

Learning to Control Coupled-Dynamics Environments with Joint Markov Decision Processes

Coupled-dynamics environments expose the one-step outcomes that would follow from several possible counterfactual actions under a common realization of exogenous randomness. The ordinary Markov decision process formalism allows one to reason about the marginal law of each action but discards dependence across these counterfactual outcomes. The Joint Markov decision process (JMDP) formalism preserves that dependence. Prior work established the formalism and solved the fixed-policy joint moment evaluation problem in JMDPs. This paper develops optimal-control methods. We define a nonparametric distributional Bellman optimality operator for JMDPs, and prove that when the induced marginal MDP has a unique optimal policy, its iterates converge in Wasserstein distance to the optimal joint return law. For the first two moments, we establish convergence under a weaker condition that permits several mean-optimal actions as long as their tie resolutions share a second-moment fixed point. We also derive sampled targets for neural approximation.

Ege C. Kaya, Aliasghar Pourghani, Mahsa Ghasemi et al. · 0 citations

A Finite-Iteration Theory for Asynchronous Categorical Distributional Temporal-Difference Learning

Finite-iteration behavior of the exact asynchronous recursions used by categorical distributional temporal-difference methods is studied and analogous finite-iteration guarantees for horizon-stacked categorical methods under episodic sampling are established.

Ege C. Kaya, Abolfazl Hashemi · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.