A broken form, a delayed dashboard, and an incorrect consent sync are all incidents, but they do not carry the same urgency. Without shared severity definitions, teams either escalate everything or wait too long on a high-impact failure. Classify incidents by impact and response needs rather than by which department discovered them.

Choose severity dimensions
Use a few dimensions people can assess quickly: customer or prospect impact, data integrity, number of affected records or campaigns, duration, and whether the issue is still active. A system outage may be severe even if few records are involved when it blocks an important service or creates a privacy risk.
Define examples for each level using the organization’s own volume and operating context. Avoid thresholds that appear precise but have no connection to how the business works.
- Critical: active customer harm, consent failure, or widespread inability to submit.
- High: important routing or measurement failure with material scope.
- Moderate or low: contained issue with a workaround and no immediate harm.
Assign response roles and actions
Every level should name a response owner, a decision-maker for pausing campaigns or data flows, and a communication channel. The person who first sees the issue should know how to report it without needing to diagnose the underlying system.
For data incidents, preserve logs and affected time ranges before correcting records. For live campaigns, decide whether to pause spend, redirect traffic, or show an alternate path. Choose actions that reduce further impact while the investigation continues.
- Record when the issue began and when it was detected.
- Notify system owners and downstream teams early.
- Use a documented workaround only when it is safe and reversible.
Learn without inflating the process
After resolution, record cause, scope, customer impact, recovery, and follow-up owner. A short review is valuable even when the root cause is a vendor outage; it may reveal that monitoring or fallback behavior was missing.
Look for repeat incident patterns and adjust validation, alerts, or release checks. Severity should guide response, not blame. A consistent record helps teams explain performance changes and identify reliability work.
- Separate confirmed facts from working hypotheses.
- Track remediation to completion.
- Revisit severity examples after a significant incident.
How did this article land?
Choose one reaction. You can change it anytime.
