How to test email in development without spamming real people
A four-level ladder from local capture servers to sandboxed providers and stranger-proof staging, plus the five email assertions that actually catch production bugs.
Somewhere right now a developer is testing a password reset flow against their own Gmail account, and a staging cron is quietly mailing five hundred real customers a message that begins with "Test test ignore". Both problems have the same cause: the development environment can reach real inboxes. Here is the ladder of ways to take that ability away, from quickest to most thorough.
Level 1: capture everything locally

A local capture server accepts SMTP like a real provider, delivers nothing, and shows every message in a web inbox. Mailpit is the current standard; MailHog is its unmaintained ancestor and still everywhere. One container:
docker run -d -p 1025:1025 -p 8025:8025 axllent/mailpitPoint your app’s SMTP settings at localhost:1025 and open localhost:8025. Every email your code sends appears instantly, with the raw source, an HTML preview, and a link check. Nothing can escape to a real inbox because nothing is ever delivered.
This is the level most teams should live at for day-to-day work. It also composes with CI: start the container in the pipeline, run the test suite, then assert against Mailpit’s API that exactly one message went to the right address with the right subject.
curl -s localhost:8025/api/v1/messages | jq '.messages[0].To'Level 2: plus addressing for realistic-looking data
Seed data full of user1@example.com hides bugs that real addresses trigger. Gmail and most providers ignore anything after a plus sign, so one inbox you own becomes unlimited distinct addresses:
yourname+customer1@gmail.com
yourname+bounce-test@gmail.com
yourname+unicode-é@gmail.comEvery one of these is a different user to your database and the same inbox to you. Use them when you genuinely need delivery, for example checking how Gmail renders your template, and keep them out of bulk seeds.
Level 3: a provider sandbox
Capture servers cannot answer the questions that involve the provider: does the API accept this payload, what does the webhook for a delivery look like, what happens to a malformed recipient. A sandbox mode runs the full provider pipeline but only delivers to addresses you have verified as your own, so the API behaves exactly like production while the blast radius stays at your team.
Sandbox is also the honest way to develop before DNS exists. Domain verification takes a registrar login and a review; a sandbox needs neither, so the integration work and the DNS work can happen in parallel instead of blocking each other.
Level 4: staging that cannot address strangers

The staging cron that mails real customers is not a testing failure, it is an authorization failure: staging held production data and production sending power at the same time. Pick at least one of these guards and enforce it in code, not in configuration someone can forget:
An allowlist in the mail layer: outside production, refuse any recipient not on your own domain. Fail loudly so the test that tried it breaks.
Scrubbed data: replace every customer address with a plus-addressed or hashed one at import time, so staging cannot leak even if the guard fails.
Separate credentials: staging uses a sandbox or test key that the provider itself refuses to deliver from. The guard lives outside your codebase entirely.
What to actually test
Capturing email is the means. These are the assertions that pay:
The retry path sends ONE email. Run the job twice, assert the count is still one. Duplicate sends are the most common transactional bug in production.
Links resolve to the right environment. A staging email with production links, or the reverse, is a classic capture-server catch.
The plain-text part exists and reads sensibly. Some corporate clients and watches show only that.
Rendering survives Gmail AND Outlook. Outlook still uses Word to render HTML; a layout that ignores that will collapse exactly where the invoices go.
Expiry behaviour: request a reset, let the token expire in a time-warped test, click, and assert the failure page is helpful.
The trap at the end
After all of this passes, the first real send still behaves differently, because your own inbox has been receiving mail from you for months and your customers’ inboxes have not. Deliverability to strangers is a reputation problem, not a correctness problem, and no amount of local testing touches it. Correctness is what this article gives you; for the other half, start with the DNS records that authenticate your mail.