The 98% Consensus: How Bulk Messaging Platforms Arrived at the Same Delivery Number and Why That Should Alarm You
Across virtually every major bulk messaging platform operating in the United States today, reported delivery rates cluster with remarkable consistency in the 98 to 99 percent range. If you have evaluated more than two or three vendors in this space, you have encountered this number. It appears in sales decks, benchmark reports, case studies, and platform documentation with such regularity that it has acquired the status of an industry standard.
It is not a standard. It is a convergence point — and understanding why every competitor arrives at roughly the same figure reveals far more about how these platforms measure success than it does about how well they actually deliver messages.
The Measurement Problem at the Foundation
Delivery rate, as reported by bulk messaging platforms, is not a single agreed-upon metric. It is a category of measurements, each platform defining its numerator and denominator differently, and each definition shaped — consciously or otherwise — by what produces the most favorable result.
The most common formulation counts messages that did not generate a hard bounce as delivered. Under this framework, a message that was accepted by a receiving mail server, immediately filtered into a spam folder, and never seen by a human being is counted as a successful delivery. A message that was accepted and silently discarded by a carrier's filtering system is counted as delivered. A message sent to a recycled phone number now assigned to a different subscriber is counted as delivered.
What is being measured, in most cases, is acceptance — not delivery in any sense that a business stakeholder would recognize as meaningful. The distinction matters enormously at scale. On a campaign of five million messages, a true delivery rate of 91 percent looks catastrophically different from a reported acceptance rate of 98.6 percent. Both numbers can be simultaneously accurate under their respective definitions.
Why the Numbers Converge
If every platform defines delivery differently, why do the reported numbers end up in the same range? The answer involves several reinforcing dynamics that collectively produce metric convergence without requiring any coordination among vendors.
The first dynamic is survivorship bias in the contact data. Platforms that have been operating for several years have, through attrition and list hygiene processes, filtered their most active customer campaigns toward cleaner contact sets. Customers who maintain high-quality lists see high delivery rates. Customers with degraded data churn out or are quietly suppressed from benchmark calculations. The reported average reflects the survivors, not the full distribution.
The second dynamic is definitional elasticity. Platforms adjust what they count in ways that are rarely disclosed in the headline metric. Unsubscribed contacts removed before sending are not counted in the denominator. Addresses flagged by internal suppression lists never enter the send queue. Messages that fail at the queue stage due to rate limiting may be categorized as unsent rather than failed. Each of these exclusions improves the reported rate without improving actual performance.
The third dynamic is competitive anchoring. When every major vendor reports 98 to 99 percent, any platform reporting 94 percent faces an immediate commercial disadvantage — regardless of whether the 94 percent figure is more honest. The market has established a floor below which reported numbers become a liability. This creates a soft incentive to define metrics in whatever way keeps the reported number above that floor.
Sampling Bias and the Problem of Representative Campaigns
Platform-level delivery statistics are typically aggregated across all customers and all campaigns. This aggregation obscures enormous variance at the campaign level, and the variance is not random — it is systematically correlated with the factors that platforms control least.
Transactional messaging campaigns — password resets, order confirmations, two-factor authentication codes — reach engaged users who have recently interacted with the sending brand. These campaigns generate high delivery rates, low complaint rates, and strong engagement signals. They are also the campaigns most likely to be featured in case studies and benchmark reports.
Bulk promotional campaigns targeting cold or aged lists, campaigns sent to purchased contact data, and re-engagement campaigns targeting dormant subscribers perform materially worse. They are also the campaigns that generate the most support tickets, the most deliverability investigations, and the most churn. Their performance data is present in the aggregate statistics but diluted by the volume of high-performing transactional traffic.
The 98 percent figure that appears in a vendor's marketing materials is, in most cases, the weighted average of a distribution that includes campaigns performing at 99.7 percent and campaigns performing at 81 percent. The average obscures the distribution, and the distribution is where your campaign actually lives.
The Perverse Incentive Structure
The commercial incentives in the bulk messaging industry are not aligned with accurate measurement. Platforms are compensated based on messages sent, not messages received. Customers evaluate platforms based on reported delivery rates, not independently verified outcomes. Sales cycles are won and lost on benchmark comparisons where every competitor is presenting the same number.
In this environment, the rational commercial behavior is to define metrics in ways that produce acceptable numbers, to exclude from reporting the categories of failure that would depress those numbers, and to invest in dashboard presentation that makes the reported figures feel authoritative. None of this requires bad faith. It emerges naturally from an incentive structure that rewards optimistic reporting over operational transparency.
The platforms that would benefit most from honest measurement — customers whose campaigns are underperforming and need intervention — are precisely the ones whose data is most aggressively smoothed by aggregation.
What Ground-Truth Measurement Actually Requires
Organizations serious about understanding their actual delivery performance cannot rely on platform-reported metrics as their primary data source. Ground-truth measurement requires independent instrumentation.
Seed list testing — embedding known, monitored addresses in every campaign send and tracking whether those addresses receive the message, where it lands, and how long delivery takes — provides a delivery signal that is independent of platform reporting. Comparing seed list results against reported delivery rates surfaces the gap between what the platform claims and what is verifiable.
Webhook and event log reconciliation, cross-referenced against CRM engagement data, provides a second verification layer. If a platform reports 98.4 percent delivery but only 61 percent of the supposedly delivered contacts show any subsequent engagement signal over the following 30 days, the discrepancy demands explanation.
The 98 percent number is not wrong, exactly. It is simply measuring something other than what most business stakeholders assume it measures. Knowing the difference is the beginning of building bulk communication infrastructure that performs as well as it reports.