← All articles

We Verified 1.3 Million Email Addresses. Fewer Than Half Were Deliverable.

We Verified 1.3 Million Email Addresses. Fewer Than Half Were Deliverable.

The number that matters is not the failure rate

Between 5 February and 11 August 2026 our verification engine processed 1,302,910 email addresses across 78,445 distinct domains. This is what they looked like.

The headline is that fewer than half were deliverable. The more useful finding is the third category most tools hide from you: 175,304 addresses — 13.5% — that no verifier can honestly call good or bad. They sit on catch-all domains, which accept mail for any address whether the mailbox exists or not.

  • 47.8%valid and deliverable
  • 38.7%definitively dead
  • 13.5%unverifiable either way
  • 2.0%role accounts

Key takeaways

  • 47.8% valid, 38.7% definitively dead, 13.5% unverifiable. A binary valid/invalid verdict forces that last group into "invalid" and quietly deletes real customers.
  • Most failures are missing mailboxes, not typos. 34.6% failed the SMTP check; syntax errors were 0.1%.
  • 4.1% of addresses sat on domains that no longer resolve at all — dead companies, expired domains.
  • Ten domains account for 86.5% of all addresses in this corpus. Email is far more concentrated than most list strategies assume.
  • We found a bug in our own disposable detector while preparing this, so disposable figures are deliberately excluded. Details below.

Method and sample

Every address submitted to our verification engine between 5 February and 11 August 2026 was checked through the same pipeline: syntax, domain resolution, MX records, disposable and role detection, then an SMTP probe that asks the receiving server whether the mailbox exists without delivering a message. Results are reported here in aggregate only. No individual address, domain-level customer data or personal information appears in this article or was used beyond counting.

This is a self-selected sample, and that matters. People send lists to a verification service precisely when they suspect those lists are dirty — after a bad campaign, before a big send, or when a purchased list arrives. A random sample of all email addresses in existence would look considerably healthier. Read these numbers as "what lists people are worried about look like", not "what email looks like".

Sample characteristicValue
Addresses verified1,302,910
Distinct domains78,445
Period5 Feb – 11 Aug 2026
Free webmail share86.5%
Largest single batch~1,000,000 addresses (February)
Geographic skewHeavily India / South Asia consumer

Two skews worth stating plainly. First, roughly one million of the 1.3 million arrived in a single February batch, so the corpus is dominated by one large consumer list rather than being evenly spread across many customers. Second, the domain mix is heavily consumer webmail with strong Indian representation. Neither invalidates the findings, but both mean you should not read this as a study of B2B corporate lists.

The three-way split

1,302,910 addresses, three honest outcomes 47.8% 13.5% 38.7% Valid — deliverable622,764 Unverifiable175,304 Definitively dead504,842 A binary valid/invalid tool reports the middle band as "invalid" — deleting 175,304 addresses it never actually tested.
The middle band is the one that costs money in both directions: send blindly and you risk reputation; delete blindly and you lose real customers.

That middle band is the entire argument for a three-verdict system. A catch-all domain accepts every address you offer it, so an SMTP probe cannot distinguish a real mailbox from a fictional one. A tool that returns only valid or invalid has to guess — and the commercially convenient guess is "invalid", because a deleted address never bounces and never shows up in your quality metrics.

Why addresses actually fail

Failure reasonAddressesShare of all
Mailbox does not exist (SMTP rejection)450,32534.6%
Domain dead — no DNS or no MX record53,2804.1%
Syntax invalid1,2210.1%
Total definitively dead504,84238.7%

The distribution contradicts a common assumption. Teams often justify verification as a typo-catching exercise, but syntax errors were 0.1% of the corpus. Form validation already catches those. The real losses are mailboxes that once existed and no longer do — people changing jobs, closing accounts, abandoning addresses.

The 4.1% with no DNS or MX record deserves attention because it is invisible to a human reviewer. The address looks perfectly reasonable; the company behind the domain simply no longer exists. No amount of careful list hygiene at capture time prevents this — only re-verification does.

Pro tip

If your list has been sitting for a year, the dead-domain share alone justifies re-verification before a large send. Those addresses cannot bounce softly — they fail hard, and hard bounces are what mailbox providers judge you on.

The unverifiable 13.5%

175,304 addresses sat on catch-all domains. This is the population that separates honest verification tools from confident ones.

What you can do with them, in order of preference:

  1. Use engagement history. A catch-all address that opened or clicked in the past is a real person. This beats any technical check.
  2. Send in small batches from a separate subdomain, watching bounce and complaint rates closely before committing the rest.
  3. Segment them permanently and never mix them into a large send with your verified population.
  4. Do not delete them by default. Corporate domains are disproportionately catch-all, which means this bucket contains a meaningful share of your most valuable contacts.

Role accounts

25,932 addresses (2.0%) were role-based — info@, sales@, support@ and similar. These are valid addresses and they are also a disproportionate source of spam complaints, because the person reading a shared inbox rarely feels they personally subscribed to anything.

They are fine for transactional mail. For marketing, segment them, watch their complaint rate separately, and suppress quickly. Under Google's bulk sender rules a complaint rate above 0.3% is enough to damage delivery for everything you send, and a shared inbox marking you as spam counts exactly as much as an individual doing so.

Domain concentration

Across 78,445 distinct domains, the top ten accounted for 86.5% of all addresses. The corpus is overwhelmingly consumer webmail, with Gmail alone representing roughly two thirds.

The practical consequence: your deliverability is largely a relationship with a very small number of mailbox providers. Optimising for "email in general" is the wrong frame. Optimising for how Gmail in particular judges you — authentication, complaint rate, engagement — is most of the job.

What we got wrong, and fixed

While preparing this study we found a defect in our own tooling, and it would be dishonest to publish the dataset without saying so.

Our disposable-address detector was running against a blocklist containing five domains. Five. It correctly flagged mailinator.com and missed yopmail.com entirely, along with thousands of other throwaway providers. Across 1.3 million addresses it flagged 11 as disposable — a number that should have been obviously wrong to us long before it was obviously wrong to a reader.

The blocklist now contains 8,200 domains, validated against the corpus so that legitimate providers are not caught by mistake. Two candidates were deliberately excluded during that check: a legitimate Indian ISP that appears on some public blocklists but runs genuine corporate mail, and a common misspelling of gmail.com which is better reported as a dead domain than as a disposable one, because that is the accurate reason.

Consequently, no disposable-address statistic appears in this study. The historical figures understate the true rate by an unknown margin, and publishing them would mean publishing our own bug as a finding. We will report disposable rates in a future update once enough volume has passed through the corrected detector.

Limitations

  • Self-selected sample. Lists submitted for verification are lists someone already distrusted.
  • One batch dominates. Roughly a million of the 1.3 million arrived in a single February upload.
  • Consumer and India-skewed. 86.5% free webmail. Do not generalise to B2B lists.
  • No industry or geography breakdown. Those fields were not populated for this corpus, so we make no claims by sector.
  • SMTP probing is provider-dependent. Some providers deliberately return uniform responses to frustrate address harvesting, which inflates apparent failure on those domains.
  • Point-in-time. An address valid in February may be dead by August. Decay is the phenomenon being measured, and it applies to the measurements too.

Frequently asked questions

Is a 52% failure rate normal for an email list?

Not for a healthy opt-in list, no. This corpus is self-selected — people verify lists they already suspect are bad, including purchased and scraped ones. A well-maintained permission-based list that is mailed regularly should perform dramatically better. The useful comparison is your own list over time rather than against this number: if your failure rate is climbing campaign to campaign, that trend matters far more than any published benchmark.

Why can't a verifier tell whether a catch-all address is real?

Because the receiving server is configured to accept mail for every address at that domain, whether or not a mailbox exists. An SMTP probe asks "will you accept mail for this address?" and a catch-all server answers yes to everything, including addresses nobody has ever used. No technical check can see past that. Any tool reporting high confidence on catch-all domains is guessing, and the only real signals available are engagement history and careful small-batch testing.

Should I delete every address a verifier marks invalid?

Delete the definitively dead ones — no mailbox, no DNS, no MX — permanently and without hesitation. Treat the unverifiable group differently: segment it, use engagement data, and test in small batches. In this corpus that distinction covered 175,304 addresses, 13.5% of the total, that a binary tool would have told you to delete despite never having tested them. Corporate domains are over-represented in that group, so blanket deletion tends to remove your most commercially valuable contacts first.

How often should a list be re-verified?

Before any significant send, and at minimum quarterly for lists you mail regularly. The 4.1% of addresses sitting on domains that no longer resolve is the clearest argument: those companies did not announce their closure to your database. Decay is continuous, so a list verified six months ago is measurably less accurate today regardless of how clean it was then.

Can I cite this data?

Yes — please do, with a link back to this page. Cite it as: The Code Smith, "We Verified 1.3 Million Email Addresses" (August 2026), n=1,302,910, verified 5 February to 11 August 2026. We would ask that you carry the sampling caveat with the figures, because the self-selected nature of the corpus genuinely changes how the headline number should be read.

Sources & further reading

  1. Primary data: NoMailBounce verification corpus, 1,302,910 addresses, 5 Feb – 11 Aug 2026. Aggregate figures computed for this study.
  2. Google, Email sender guidelines — the 0.3% spam complaint threshold and authentication requirements referenced above.
  3. IETF, RFC 5321 — Simple Mail Transfer Protocol — the specification behind the SMTP responses these verdicts are derived from.

Want this working
in your business?