Incidents & alerts
Incidents & AI reports
Incidents open by themselves when something breaks, collect related failures, explain themselves with AI, and close when everything recovers. You or your agents can step in at any point.
Automatic incidents
An incident opens automatically when:
- a monitor goes down (after confirmation). Degraded alone doesn't open one, or
- a push- or manifest-fed component reaches
partial_outageormajor_outage.
It is public on every page where an affected component lives and settings.autoIncidents is on. Otherwise it stays internal: your team is still alerted, but nothing shows on the page. The first public update is a neutral template (“We're investigating an issue affecting Search API…”). An internal note records the raw check error, such as Automated check failed: Expected status 200, got 502.
Grouping: one incident, not five
If another failure hits the same status page while an automatic incident there is still open and started less than 30 minutes ago, it joins that incident instead of opening a new one. The new components are added, the impact is raised if needed, and a public update says “We're also seeing issues with Billing, Webhooks.” A shared upstream failure therefore produces one incident and one round of notifications.
Automatic recovery
When a monitor recovers, or a component drops back below partial_outage, its part of the incident is marked operational. If other parts are still failing, the incident moves to monitoring with an update like “Search API has recovered. We're still working on the remaining issues.” When nothing is failing anymore, the incident resolves on its own with a recovery note, written by AI when available (see below) or from a template (“…operating normally again after roughly 14 minutes of disruption”).
Statuses and impact
| Field | Values |
|---|---|
Incident status | investigating → identified → monitoring → resolved |
Maintenance status | scheduled → in_progress → completed |
impact | none · minor · major · critical · maintenance |
source | monitor · heartbeat · component · user · agent · api |
When you don't set impact, it comes from the worst affected component: major_outage → critical, partial_outage → major, degraded → minor. A manual incident with no components defaults to minor.
Report an incident
Humans use the dashboard. Agents and scripts call POST /incidents (MCP: incidents_create):
curl -X POST https://upbutler.com/api/v1/incidents \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"title": "Elevated API errors",
"message": "We are seeing 5xx errors on search and are investigating.",
"impact": "major",
"components": [{ "componentId": "cmp_0n3q8kz2p7wd4yx0s3a", "status": "partial_outage" }]
}'componentssets each listed component's status for the duration (as an override). Their pages are added topageIdsautomatically.publicdefaults totruewhen pages or components are given. An incident with neither is internal.notify: falseskips subscriber notifications. Your alert channels still getincident.created.statuscan start atinvestigating(default),identifiedormonitoring.
Updates
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/updates \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"status": "identified",
"message": "bad deploy of search-indexer v2.14, rolling back",
"polish": true,
"components": [{ "componentId": "cmp_0n3q8kz2p7wd4yx0s3a", "status": "degraded" }]
}'
# Internal note: not shown on the page, not sent to subscribers
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/updates \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"message": "rollback ETA 10 min, paged @dana", "public": false}'Each update can change status and component statuses. Setting a component to operational clears its override. Public updates emit incident.updated to channels and subscribers. Internal notes ("public": false) emit nothing. Posting status: "resolved" resolves the incident.
Resolve
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/resolve \
-H "Authorization: Bearer $UPBUTLER_API_KEY"Send a message to write the resolution yourself. Without one, UpButler writes it: the AI recovery summary when available and the page allows AI public updates, otherwise a template with the duration. Resolving sets every affected component back to operational, clears the incident's overrides and emits incident.resolved with durationMinutes and durationText.
AI reports
UpButler's AI reads the monitor's configuration, the last checks (status codes, latencies, errors, failed assertions), evidence from the failing check (filtered headers, body snippet, TLS and DNS), the monitor's previous incidents, and any deploy markers from the hour before. From that it writes:
- An opening report, when a monitor opens an automatic incident: a status-page title, a customer-safe
publicSummary(no hostnames, IPs or stack traces), a technicalanalysisfor engineers, rankedlikelyCauses, concretesuggestedActions, and aseverity. - A recovery report, when the incident resolves: a post-incident summary covering duration, failure mode and how recovery happened, plus follow-ups to prevent a repeat.
"ai": {
"opening": {
"model": "deepseek/deepseek-v4.1-flash",
"at": "2026-10-09T08:12:31.402Z",
"title": "Elevated errors on the Search API",
"publicSummary": "Some search requests are failing. We're investigating and will post an update shortly.",
"analysis": "Since 08:10:44 every check returns HTTP 502 from Cloudflare within ~300ms while TLS and DNS are healthy; the origin is refusing connections rather than timing out. The pattern is hard-down, not intermittent.",
"likelyCauses": ["Origin process crashed or was not restarted after deploy", "Load balancer health checks removed all backends"],
"suggestedActions": ["Check origin process status and recent deploys", "Inspect load balancer target health", "Roll back the last release if it coincides"],
"severity": "critical"
},
"recovery": { "...": "same shape, written when the incident resolves" }
}Where each part goes:
analysis,likelyCausesandsuggestedActionsgo to your team: they are attached tomonitor.downalerts (email, Slack, Discord, Telegram and webhook payloads underdata.ai) and added to the timeline as an internal note.titleandpublicSummaryreplace the template on the status page, but only if every page the incident is on hasaiPublicUpdateson, and the text respects the pages' AI policy (guidelines and avoided terms). Otherwise the neutral template stays.- When the timing fits, the analysis names the deploy, such as “errors began 2 minutes after deploy 1.4.2 (commit abc1234)”. Deploys are never mentioned in public text.
Polish rough notes
Pass polish: true to POST /incidents, /incidents/:id/updates or /incidents/:id/resolve, and your draft (“bad deploy of indexer, rolling back”) is rewritten into a calm 2–4 sentence public update with no internal details. Polished updates are flagged ai: true. If polishing fails, your original text is posted.
Analyze on demand
POST /incidents/:id/analyze generates or regenerates the report: the opening report for an open incident, the recovery report for a resolved one. It stores the report under ai.opening or ai.recovery and returns it. Incidents without a monitor (declared by people or agents, or opened by pushed statuses) get a report summarizing their timeline instead. If AI is unavailable or the quota is used up, it returns 503 unavailable.
Quota
Every opening report, recovery report, polish and analysis counts as one AI report. Usage resets monthly and is shown in GET /workspace under usage.aiReports.
| Plan | AI reports / month |
|---|---|
| Free | 10 |
| Starter | 100 |
| Pro | 1,000 |
| Business | 5,000 |
Acknowledging
POST /incidents/:id/ack (every plan) marks an incident as owned. It stops “still down” reminders and escalation, adds an internal timeline note and sends incident.acknowledged to the alerted channels. Alerts also carry one-click Acknowledge links. See On-call & escalation.
Postmortems and templates
After an incident, POST /incidents/:id/postmortem drafts a blameless postmortem with AI that you can edit and publish to your status page. For the next incident, incident templates give you ready-made titles and messages with placeholders. See Insights.
Maintenance
curl -X POST https://upbutler.com/api/v1/maintenance \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"title": "Database upgrade",
"message": "We are upgrading our primary database to PostgreSQL 18. Search may be read-only for up to 30 minutes.",
"scheduledStart": "2026-11-02T22:00:00Z",
"scheduledEnd": "2026-11-02T23:00:00Z",
"componentIds": ["cmp_0n3q8kz2p7wd4yx0s3a"]
}'Maintenance is an incident with kind: "maintenance" and id mnt_…. It is created as scheduled and emits maintenance.scheduled. At scheduledStart, it moves to in_progress by itself, its components show Under maintenance, and maintenance.started goes out. At scheduledEnd, it completes, the overrides are cleared and maintenance.completed goes out. Pass notify: false to keep subscribers out of the loop for the whole window (scheduled, started and completed). Your team's channels still get all three. Progress updates posted during the window emit maintenance.updated. Post updates during the window with /incidents/:id/updates, the same as incidents.
Finding incidents
GET /incidents filters by open=true, kind, pageId and monitorId, up to limit=200. GET /incidents/:id returns the full timeline, including internal notes and AI reports. The public view (/public/pages/:slug/incidents) only ever contains public updates.