Most messaging teams find out about a delivery problem the slow way. A sales lead mentions that customers "didn't get the code," a campaign report looks oddly flat a day later, or someone notices the bounce column turned red over the weekend. By then, thousands of messages have already gone into a hole, and the sender reputation you spent months building has taken a hit. Delivery anomaly alerts flip that timeline.
Instead of waiting for someone to notice, your messaging stack watches its own vital signs and raises a flag the moment something drifts outside its normal range. This guide walks through which signals to monitor, how to set baselines that adapt to your traffic, how to keep alerts useful instead of noisy, and what to do in the first fifteen minutes after one fires.
## Why Static Thresholds Fail The first instinct is to set simple rules: alert if the bounce rate goes above 5%, or if SMS delivery drops below 90%. These rules are better than nothing, but they break down quickly in real programs. ### Normal changes by channel, country, and time A 92% delivery rate might be excellent for one SMS route and a red flag for another.
Email bounce rates for a freshly imported list look nothing like those for a warm, engaged segment. Weekend traffic, holiday sends, and time-zone waves all shift what "normal" looks like hour to hour. ### Volume hides problems When volume is low, a single failed batch can swing percentages wildly and trigger false alarms.
When volume is high, a 2% drop can mean tens of thousands of missed messages while still sitting comfortably above a fixed threshold. ### Slow drifts slip through Static rules catch cliffs, not slopes. A route that loses half a point of delivery every day for two weeks never crosses a hard line on any single day, yet ends up far worse than where it started.
Anomaly detection solves these problems by comparing each metric against its own recent history, sliced the way your traffic actually behaves. ## The Signals Worth Watching You do not need to monitor everything. A focused set of signals covers the vast majority of real incidents. ### SMS signals - **Delivery rate by route and country.** The single most important SMS metric. Track it per provider route and per destination country, since problems are rarely global.
- **Failure codes mix.** A sudden rise in a specific carrier error, such as unknown subscriber or blocked content, tells you far more than a generic failure count. - **Delivery receipt latency.** If receipts start arriving minutes later than usual, a route is often congested or degrading before it fails outright. - **Missing receipts.** Some routes never return delivery receipts.
Know which ones, so a "zero delivered" reading on those routes does not trigger a false alarm. - **Opt-out rate.** A spike in STOP replies after a send usually points to a content, frequency, or targeting problem. ### Email signals - **Hard and soft bounce rates by receiving domain.** A jump in deferrals at one mailbox provider often signals throttling or a reputation issue before full blocking starts.
- **Complaint rate.** Even small increases matter, because complaint thresholds at major inbox providers are very low. - **Open and click rates by domain.** A sharp drop at one provider while others stay steady is a classic sign of messages landing in spam. - **Authentication failures.** Any rise in SPF, DKIM, or DMARC failures deserves immediate attention, especially after DNS or infrastructure changes.
### Pipeline signals - **Queue depth and age.** Messages waiting longer than usual mean something upstream is slow or stuck. - **Send throughput.** A campaign that normally sends 5,000 messages per minute and suddenly sends 500 is an incident, even if every message that does go out is delivered. - **Provider balance and API errors.** Running out of credit or hitting authentication errors at a provider stops traffic instantly. These deserve their own alerts.
## Building Baselines That Adapt The heart of anomaly detection is a good baseline: a reasonable expectation of what a metric should look like right now. ### Use rolling windows Compare the last hour against the same hour across the past two to four weeks, rather than against a single fixed number. This automatically accounts for daily and weekly rhythms.
A rolling median is often more robust than an average because it ignores the occasional extreme day. ### Slice by the dimensions that matter Build baselines per route, per country, per sending domain, and per receiving domain where you have enough volume. A global average will mask a localized failure, and localized failures are the most common kind.
### Measure deviation, not just distance Rather than asking "is the delivery rate below 90%," ask "how far is the current delivery rate from its typical range?" A common approach is to calculate how many standard deviations the current value sits from the rolling baseline. Values that land well outside the normal spread are anomalies; small wobbles are not.
### Respect minimum sample sizes Set a floor for how many messages must be in a window before an alert can fire. Fifty failed messages out of sixty sends is worth knowing about. Two failures out of three is noise. A simple rule like "at least 200 sends in the window" eliminates most false alarms on quiet routes.
### Catch slow drifts separately Pair your short-term detector with a slower one that compares this week to the previous month. This catches gradual degradation that hourly comparisons treat as normal because the baseline drifts along with it. ## Designing Alerts People Actually Read An alert that fires ten times a day gets muted. The goal is fewer, better alerts that always deserve action.
### Tier alerts by severity - **Critical:** Traffic has effectively stopped, a provider is rejecting all messages, or delivery has collapsed on a major route. These should page someone immediately. - **Warning:** A meaningful deviation that needs attention within the hour, such as rising deferrals at one inbox provider or a route sliding below its usual band. - **Informational:** Trends worth reviewing in a daily summary, like a slow drift in opt-out rates.
### Include context in every alert A good alert answers the obvious questions before anyone asks them: which metric, which route or domain, the current value, the normal range, how many messages are affected, and when it started. A link to the relevant dashboard or campaign view saves precious minutes. ### Deduplicate and group If one provider goes down, you do not need forty alerts for forty campaigns.
Group related anomalies into a single incident and update it as the situation changes. Suppress repeat alerts for the same issue for a set cooldown period unless the severity escalates. ### Route alerts to the right place Critical alerts should reach a channel people watch in real time, such as a team chat or mobile notification. Warnings can go to an operations channel. Informational items belong in a daily digest.
Matching the channel to the urgency keeps attention where it belongs. ## Pairing Alerts With Automatic Safeguards Detection is half the value. The other half is what happens automatically while a human is still reading the alert. ### Pause before you burn reputation When bounce or complaint rates spike far above baseline during a send, automatically pausing that campaign protects your domain and IP reputation.
It is almost always cheaper to resume a paused campaign than to repair a damaged sender score. ### Fail over between routes For SMS, if a route's delivery rate collapses or the provider runs out of balance, shifting traffic to the next best route keeps time-sensitive messages flowing. Do this in measured steps, watch the results, and roll back automatically if the backup performs worse.
### Do not retry what will never deliver Invalid or disconnected numbers and hard-bounced addresses should be marked as failed and suppressed, not retried on another route. Retrying them wastes money and can make a healthy route look bad. ### Log every automatic action Every pause, shift, or suppression should be recorded with the reason and the data behind it. This makes post-incident reviews faster and builds trust in the automation over time.
## A First-Fifteen-Minutes Playbook When an alert fires, a short, repeatable routine prevents panic and guesswork. 1. **Confirm the scope.** Is the anomaly limited to one route, country, domain, or campaign, or is it everywhere? Scope usually points straight at the cause. 2. **Check recent changes.** New templates, DNS edits, list imports, or provider configuration changes in the last day are the most common culprits. 3.
**Look at the error details.** Specific failure codes and bounce messages tell you whether you are dealing with blocking, throttling, invalid data, or a provider outage. 4. **Contain the damage.** Pause affected sends or shift traffic if automation has not already done so. 5. **Communicate
Related Articles
Privacy expectations keep changing. So do messaging rules. Customers want useful SMS and email without losing control of their data. Compliance protects...
Build a working email and SMS operating system: identity, triggers, channel roles, suppression, contact policy, testing, and a pragmatic 30‑day rollout.
Mailbox providers do not judge your brand by intent. They judge by signals: authentication, complaints, engagement, bounce quality, and sending patterns.
Transactional email keeps products working. Learn how to run Amazon SES transactional mail as infrastructure—auth, warm-up, and monitoring.
Discover the critical changes in email deliverability for 2026, from AI-powered filters to stricter authentication. Learn actionable strategies to ensure your emails consistently
Unlock peak email deliverability for your new domain or IP. This guide covers essential warm-up strategies, sender reputation, and 2026 best practices for optimal inbox placement.
A practical 2026 guide to GDPR and CAN-SPAM compliance for email marketers, covering consent, data handling, unsubscribe rules, and campaign-safe automation.
Design an email preference center that cuts unsubscribes and spam complaints—without dark patterns or preferences you never honor.
Explore SESender
SESender brings audience preparation, contact validation, sender and provider controls, scheduling, delivery tracking, and campaign reporting into one workspace. Review the current product and pricing information before deciding whether the platform fits your messaging workflow.
Explore the platform or review pricing.