Oct 2026· Proceedings of the 2026 ACM Internet Measurement Conference· 0 citations· 14 references
TL;DR
This work deploys a platform (RetryWatch) that simulates various types of bounce errors returned by email receiving servers, covering the DNS, TCP, and SMTP communication stages and finds that many providers and software exhibit defects in their email retry implementations.
Abstract
Reliable delivery is a basic expectation of email services. The increasingly complex ecosystem has led to more frequent email bounces. When encountering various errors, properly handling failures and retrying are crucial to ensuring email deliverability. Currently, the community lacks a comprehensive understanding of email retry strategies, and bridging this gap can enhance the reliability of email services. In this paper, we measure the email delivery retry strategies of 11 email service providers and 8 open-source email software. We deploy a platform (RetryWatch) that simulates various types of bounce errors returned by email receiving servers, covering the DNS, TCP, and SMTP communication stages. We find that many providers and software exhibit defects in their email retry implementations. For example, we find that six email systems exhibit poor load balancing across MX record entries in dual-stack environments. In addition, when encountering blocklist- and greylist-related SMTP error messages, most email providers fail to handle retries appropriately. We release RetryWatch to support further testing by the email community. We hope this work helps improve email reliability and inspires efforts to refine email retry strategies.
Enterprise email agents must combine information retrieval, structured state changes, temporal reasoning, and multi-step coordination. Recent agent benchmarks include productivity tasks, but few center on typed email workflows in a self-contained environment. We introduce EmailBench, a benchmark of 206 email and produc...
Mukul Singh, Mansi Uniyal, Devin Devlin et al.· 0 citations
Many Model Context Protocol (MCP) servers wrap web APIs built for human developers, and their error messages tell the reader to run a command, edit a configuration, open a web page or wait. Many agents that read them can only call the server's tools. In 150 widely used MCP servers, 949 of 3,001 error messages tell the...
Retry mechanisms are a standard component of resilient distributed systems, but their collective behavior, when every tier in a call path retries concurrently, is less well understood than the per-client guidance that produced them. This paper introduces the retry amplification factor (RAF), a metric quantifying the ad...
Outcome Monitors, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas, are introduced, which detect violations of outcome contracts mined from task-disjoint traces or derived from public schemas.
Large language model (LLM) agents that retrieve external content and use tools are vulnerable to indirect prompt injection, in which untrusted content contains instructions intended to influence agent behavior. We evaluated four defenses and an undefended control across GPT-5.4, GPT-5.4-mini, and Claude Sonnet 4.6 on t...
Adil Khan, Khaled AlKhanbashi, Azza Mohamed· Computers· 2 citations
This paper describes the re-architecture of a distributed payment system around an event-driven saga pattern, tracking transactions from initiation through settlement and automatically triggering compensations when a step fails. The results are given directly: transaction processing time fell by 83 per cent, payment er...
Tarun Kumar Chatterjee· Proceedings of the Raptors C...· 0 citations
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 2, 2026
Writing as a participant and researcher, PhD student JS Tan SM ’22 has co-authored a new book about the rise of tech worker protests and the employer backlash that followed.
Requirements in large systems rarely exist in isolation. Their meaning depends on the wider project context - other requirements, policies, decisions, tests, and implementation details. That becomes especially important when AI is used for review, because spotting a possible conflict or gap is only the beginning. ReqSpace explores how AI, visualisation, and connected project context can help reviewers understand those findings, trace the relationships behind them, and focus on the questions that…
AI is making software generation faster, but speed does not remove the need for expertise. As more work is delegated to AI, tacit knowledge may become one of the most important human advantages in software engineering. The post Beyond Prompt Engineering: The Role of Tacit Knowledge in Software Engineering appeared first on GPT-Lab.