14 edge cases that quietly destroy transactional email
Fourteen production failure modes that never show up in testing: retry duplicates, late OTPs, alignment failures, the Gmail and Yahoo thresholds, bounce handling done backwards, and the it-works-in-testing trap.
Transactional email looks finished the day the happy path works: call the API, the message arrives, ship it. Every case below is something that worked exactly like that in testing and then failed in production, quietly, for real users. Fourteen of them, grouped by where they bite.
Sending
1. The retry that sends twice
Your send call times out, your retry logic fires, and the customer gets two receipts, or two different one-time codes racing each other. The timeout does not mean the first send failed; it means you do not know. Fix: pass an idempotency key with every send so a retry returns the original result instead of a sibling.
2. The OTP that arrives after it expired
A five-minute code that spends six minutes in a queue is indistinguishable from a broken login. Codes and magic links belong on your fastest sending path, ahead of every receipt and digest, and the expiry you print in the message should include realistic delivery time, not just cryptographic caution.
3. Rate-limited by yourself, mid-flow
A batch job shares an API key with your auth emails, the batch hits the rate limit, and password resets start failing behind it. Separate keys for separate traffic classes, and treat a 429 on the auth path as an incident, not a retry.
4. Bounce webhooks out of order
A delivery notification can arrive after the bounce for the same message, because webhooks are racing HTTP requests, not a transcript. Process them by comparing states, never by assuming sequence: a terminal state like bounced must not be overwritten by a late delivered event.
Authentication and reputation
5. DKIM passes, DMARC fails
Your provider signs the mail with their own domain, the signature verifies, and DMARC still fails because the signing domain does not match your From header. That mismatch is called alignment, and it is the single most common "but everything passes" confusion. The fix is custom-domain DKIM, which is exactly the record your provider generates when you verify a domain. The full picture is in our plain-English DNS guide.
6. The neighbour on your shared IP
On a shared sending IP, a spammer three customers over can get the range throttled at a big receiver, and your delivery times stretch from seconds to hours through no act of yours. You cannot prevent it; you can choose a provider that polices its network aggressively, and you can watch your own delivery latency so you notice before your users tell you.
7. The Gmail and Yahoo thresholds
Since February 2024, bulk senders to Gmail and Yahoo live under published limits: keep spam complaints under 0.3 percent and bounces under 5 percent, or watch the whole domain’s mail degrade. Those numbers are a cliff, not a gradient, and the only defence is measuring your own rates continuously instead of discovering them in a postmaster dashboard after the fall.
8. No one-click unsubscribe where it was required
The same rules require RFC 8058 one-click unsubscribe headers on promotional mail. The edge case is the email you believed was transactional and the receiver classified as marketing, a "we miss you" note, a feature announcement. When in doubt, carry the headers; an unnecessary unsubscribe link costs nothing, a missing required one costs placement.
Lifecycle
9. The soft bounce you treated as hard
A full mailbox is temporary; suppressing the address forever turns one over-quota week into a customer you can never email again. The reverse is worse: retrying a hard bounce forever burns reputation on an address that cannot exist. The two verdicts arrive in the same webhook and deserve opposite handling.
10. The hard bounce you never suppressed
Every send to a known-bad address after the first bounce is a signal to receivers that you do not read your own feedback. Suppression must be automatic and enforced at send time, not a cleanup script somebody runs before campaigns.
11. The reply-to black hole
Users reply to receipts and password resets with real problems: wrong charge, hacked account. A no-reply@ address converts those into silence, and some fraction of the silenced go to their bank instead, which arrives back to you as a chargeback. Route replies to a mailbox a human reads.
12. The subdomain that quietly lost verification
A DNS migration drops one TXT record, nothing fails immediately because verification is cached, and days later mail starts soft-failing everywhere at once. Verification is a state that decays, so it needs monitoring like any other dependency: re-check continuously and alert on the transition, not on the eventual delivery failures.
Rendering and the last trap
13. Renders in Gmail, collapses in Outlook
Desktop Outlook still renders HTML with Word’s engine: no flexbox, unreliable margins, background images gone. The emails where that matters most are invoices and receipts, the exact ones businesses open in Outlook. Test both engines before shipping a template, not after the first complaint.
14. It works in testing
Your test inbox has received your mail for months; your domain is warm to it; every link resolves to localhost decisions you have already made. None of that is true for a stranger’s inbox on day one. Treat the first week of production sending as its own rollout: small volumes, watched bounce and complaint rates, and a way to pause. The fourteen items above are, in the end, one item: production email is a feedback system, and shipping it means shipping the loop that listens.