Skip to content
#software testing Book Open access

When Delivery Meets Error: Exploring Email Delivery Retry Strategies and Defects

Oct 2026 · Proceedings of the 2026 ACM Internet Measurement Conference · 0 citations · 14 references

TL;DR

This work deploys a platform (RetryWatch) that simulates various types of bounce errors returned by email receiving servers, covering the DNS, TCP, and SMTP communication stages and finds that many providers and software exhibit defects in their email retry implementations.

Abstract

Reliable delivery is a basic expectation of email services. The increasingly complex ecosystem has led to more frequent email bounces. When encountering various errors, properly handling failures and retrying are crucial to ensuring email deliverability. Currently, the community lacks a comprehensive understanding of email retry strategies, and bridging this gap can enhance the reliability of email services. In this paper, we measure the email delivery retry strategies of 11 email service providers and 8 open-source email software. We deploy a platform (RetryWatch) that simulates various types of bounce errors returned by email receiving servers, covering the DNS, TCP, and SMTP communication stages. We find that many providers and software exhibit defects in their email retry implementations. For example, we find that six email systems exhibit poor load balancing across MX record entries in dual-stack environments. In addition, when encountering blocklist- and greylist-related SMTP error messages, most email providers fail to handle retries appropriately. We release RetryWatch to support further testing by the email community. We hope this work helps improve email reliability and inspires efforts to refine email retry strategies.

Read PDF

Similar papers

#artificial intelligence Preprint Sep 2026

EmailBench: A Benchmark for Evaluating LLM Agents on Enterprise Email and Productivity Tasks

Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and produc...

Mukul Singh, Mansi Uniyal, Devin Devlin et al. · 0 citations
#artificial intelligence Preprint Sep 2026

MCP Error Messages Written for Developers Hurt the Most Capable Agents Most

Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the...

Xiaonan Xu, Wen-Jing Wu · 0 citations
Preprint Open access Aug 2026

Retry Amplification in Distributed Systems: A Systematic Analysis of Retry Policies and Their Role in Cascading Failures

Retry mechanisms are a standard component of resilient distributed systems, but their collective behavior, when every tier in a call path retries concurrently, is less well understood than the per-client guidance that produced them. This paper introduces the retry amplification factor (RAF), a metric quantifying the ad...

Rishabh Mehan, Jasmit Saluja · 1 citation
Preprint Aug 2026

Outcome Monitors: Recovery Affordances for Silent Tool Failures

Outcome Monitors, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas, are introduced, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas.

Sugam Panthi, Rabab Abdelfattah · 2 citations
Open access Aug 2026

Evaluating Indirect Prompt Injection Defenses in Tool-Using LLM Agents: Security, Utility, and Replication

Large language model (LLM) agents that retrieve external content and use tools are vulnerable to indirect prompt injection, in which untrusted content contains instructions intended to influence agent behavior. We evaluated four defenses and an undefended control across GPT-5.4, GPT-5.4-mini, and Claude Sonnet 4.6 on t...

Adil Khan, Khaled AlKhanbashi, Azza Mohamed · 2 citations
Conference Open access Sep 2026

83% Faster, 10x Fewer Errors: Engineering a Self-Healing Payment Saga Tracker

This paper describes the re-architecture of a distributed payment system around an event-driven saga pattern, tracking transactions from initiation through settlement and automatically triggering compensations when a step fails. The results are given directly: transaction processing time fell by 83 per cent, payment er...

Tarun Kumar Chatterjee · 0 citations

Related blog posts

MIT News · Artificial Intelligence Oct 2, 2026

Documenting the tech worker movement

Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.

GPT-Lab Sep 23, 2026

Requirements Don’t Live in Isolation: What We’re Exploring with Req-Space

Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…

GPT-Lab Sep 17, 2026

Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering

AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.

We use cookies to run the site and, with your consent, for analytics and to show ads. See our Cookie Policy.