Skip to content

Author

Shafiq Joty

We have 11 of 66 papers

We haven’t gathered this author’s papers yet. Follow them and we’ll fetch their work.

Not the right person? Other researchers publish under this name.

#artificial intelligence Preprint Sep 2026

Understanding the Synergy between SFT, RLVR, and OPD in LLM Post-Training

Modern LLM post-training composes supervised fine-tuning (SFT), reinforcement learning with verifiable rewards (RLVR), and on-policy distillation (OPD) into multi-stage pipelines, yet these stages are typically designed and evaluated in isolation. We show that this composition is consequential: a stage that improves th...

Emre Can Acikgoz, Yang Li, Z. Liu et al. · 0 citations
Preprint Sep 2026

Flattening Every Memory Peak in Long-Context Mixture-of-Experts Training

Four parallelism plans are left unbounded by the parallelism plans in common use, and each grows differently: expert dispatch with the routing matrix, the vocabulary projection with tokens times vocabulary, gradient checkpoint boundaries with depth times sequence length, and optimizer state with parameter count.

Shrey Pandit, Xuan-Phi Nguyen, Yiran Zhao et al. · 0 citations
#natural language process... Preprint Sep 2026

EvoHarnessBench: Can Your Agents Keep Pace with an Evolving Harness?

Modern LLM-based agents operate through a harness of tools, reusable skills, and specialist agents that shapes what they observe and what they can do. In practice, this harness continually evolves as new capabilities are added. We introduce EvoHarnessBench, a benchmark for evaluating agents under controlled harness evo...

Zi-Xuan Ke, Vaidehi Patil, Hai-Zhou Shi et al. · 1 citation
Preprint Aug 2026

VisEditBench: Can Vision-Language Models Edit Visualization Code from Multimodal Feedback?

This work introduces VisEditBench, a benchmark of 1,395 human-annotated visualization code-editing tasks grounded in realistic visualization workflows and failure cases, and proposes VisEditAgent, a render-grounded editing framework that iteratively generates, executes, validates, and refines candidate edits.

Mizanur Rahman, Arshia Azimlu, Shadikur Rahman et al. · 3 citations · ⚡1
Preprint Aug 2026

DSAgentBench: Can Agents Automate End-to-End Data-Science Workflows in Real Computer Environments?

DSAgentBench is introduced, the first benchmark to evaluate whether agents can automate full data-science workflows inside real computer environments, and reveals a substantial capability gap between current agentic systems and real data-science workflows.

Mizanur Rahman, Mohammed Saidul Islam, Ridwan Mahbub et al. · 0 citations
#machine learning Preprint Aug 2026

Learning Generalizable Behaviors for Terminal Agents

River, a simple training recipe that improves reward quality by filtering low-quality environments and augmenting outcome rewards with process-level behavior regularization is proposed, which achieves the best performance among evaluated open-source RL-trained 8B models across four terminal-agent benchmarks.

Yi-Fan Yao, Bo Pang, Xuan-Phi Nguyen et al. · 2 citations
Preprint Jul 2026

Procedural Memory Distillation: Online Reflection for Self-Improving Language Models

Procedural Memory Distillation is proposed, which converts crossepisode signals into reusable procedural memory and distills it into the policy's weights during training, yielding a memory-free model at inference.

Ye Liu, Srijan Bansal, Bo Pang et al. · 2 citations

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.