Repeated failures
Configure consecutive-failure thresholds so one noisy check does not immediately become a confirmed outage.
Incident management / Context before action
Confirm the failure, coordinate your team, and keep customers informed. Every check, decision, and update belongs to the same incident story.
10 monitors free · No credit card required
Virginia and Frankfurt returned HTTP 503 on two consecutive checks.
Regional evidence
2 of 3 failing
Team notification
Delivered
One record for evidence, delivery history, and team updates.
01 / Confirm
An investigation collects evidence. Your configured thresholds decide when it becomes a confirmed outage and notifications should follow.
Configure consecutive-failure thresholds so one noisy check does not immediately become a confirmed outage.
Require independent regions to agree before confirming the incident.
Keep expected downtime separate with maintenance windows and monitor mute controls.
02 / Coordinate
Acknowledge an incident, assign a responder, and set severity. Keep investigation notes beside the checks that explain the failure.
Example response ownership
Maya Chen
Assigned responder · Acknowledged
Severity: Major · Illustrative team activity
Virginia and Frankfurt return HTTP 503. One failed check starts an investigation; the team has not been alerted yet.
Both regions fail their next check. Two consecutive failures in two of three regions meet this monitor’s confirmation rules.
Email and Slack delivery succeed. The alert includes the affected monitor and regional evidence.
An operator publishes a public update: “We are investigating errors affecting checkout. Our team is working to restore service.”
03 / Communicate
Give responders operational detail and customers a clear public update. Keep both connected to the incident.
Inspect notification attempts and outcomes. A delivery failure stays distinct from the outage itself.
Illustrative delivery history · INC-1042
Explore notification integrationsPublish updates to your status page while keeping internal investigation notes and monitor targets private.
Acme Cloud · Example public update
Investigating checkout errors
We are investigating errors affecting checkout. Our team is working to restore service.
04 / Resolve
Resolve through configured recovery thresholds or a responder action. Preserve the timeline and resolution summary so the next review starts with evidence.
Walk through an example incidentBe ready for the next signal
Start monitoring for free, with incident context built in.