Skip to content
Docs/Incident drills

Agents on call

Incident drills

A fire drill for your on-call chain. UpButler opens a made-up incident on monitors you choose, pages your responders exactly as a real outage would, and scores what they did. Nothing public is touched.

What a drill is

A drill is a synthetic incident on one or more monitors. The escalation policy for those monitors runs for real: agent responders get a page with a simulated failure, claim it, investigate, verify and resolve or hand back, with the same API and MCP tools as always. When the drill ends (on resolve, on hand-back or at its time limit) UpButler scores the response over six criteria, deletes the synthetic incident and keeps a report with the full timeline.

Start one from Drills in the dashboard, with POST /api/v1/drills, the MCP tool drills_start, or on a schedule. One drill runs per workspace at a time.

What a drill never touches

Status pages, subscribers, monitor state and uptime stay exactly as they are. Three layers make sure of it:

  1. An internal incident. The drill's incident is created with public: false, no status pages and no components, so nothing public can show it. The monitors keep their real state, checks, history and uptime; the failure exists only in what the responder is shown.
  2. Events are stripped. Every event about a drill incident has its pages, components and subscribers removed before it is stored, and carries a drill field. Alert channels are only kept when the drill includes people, and then only the channels the drill names.
  3. The outbox refuses. The delivery worker checks the drill again before sending: it never delivers a drill event to a subscriber, and never to a person or an alert channel unless the drill includes people.

A drill starts only on healthy monitors. A monitor that is down or has an open incident cannot be drilled, so a rehearsal never hides a real problem. If a real incident opens on one of the drill's monitors while it runs, the drill ends at once.

Scenarios

The person starting a drill picks the scenario. The responder is told that it is a drill, never which one.

ScenarioWhat the responder seesEnds whenRight ending
Service down
down
Every region gets HTTP 503 from the origin. It heals by itself after a few minutes. A good responder names the origin as failing everywhere, waits for a passing verify and only then resolves.Heals by itself after recoverAfterSec (default 120 s)Resolve after it recovers, or hand back
Partial errors
degraded
About a third of requests fail with HTTP 500 while the rest succeed. It heals by itself. A good responder notices it is partial, not a full outage.Heals by itself after recoverAfterSec (default 120 s)Resolve after it recovers, or hand back
Slow responses
slow
Responses still succeed but take several seconds, and some regions time out. It heals by itself. A good responder names latency, not an error.Heals by itself after recoverAfterSec (default 120 s)Resolve after it recovers, or hand back
Bad deploy
bad_deploy
A (made-up) production deploy goes out four minutes before the checks start failing with HTTP 500. Rolling it back ends the fault. A good responder names the deploy and rolls back, or hands back recommending the rollback when it is not allowed to.A (simulated) rollbackFix it and resolve
Upstream outage
upstream_outage
A third-party provider the service depends on is failing (HTTP 502, "upstream provider unavailable"). Nothing in your stack can fix it and it does not heal during the drill. A good responder says so and hands back instead of rolling back or resolving.Does not end during the drillHand back to people
Runaway agent run
runaway_run
An agent run repeats the same tool call and burns through its budget. Stopping the run ends the fault. A good responder names the loop, stops the run and verifies.Stopping the run (simulated)Fix it and resolve

The made-up deploy in bad_deploy has a commit that starts with d1211 and a version like drill-2026-10-10, so it can never be mistaken for a real one.

What the responder sees

The handoff packet of a drill incident has drill: true and a drillInfo object with the rules of the drill. The incident title and every page start with [DRILL], and the prompt for prompt-based responders begins with "DRILL: this is a rehearsal, not a real incident." The monitors in the packet show the simulated failure (status codes, errors, latency, a suspect deploy or a looping run, depending on the scenario); real evidence, deploys and related incidents are left out so nothing true is mixed in.

Handoff packet (abridged)
{
  "drill": true,
  "drillInfo": {
    "id": "drl_…",
    "endsAt": "2026-10-10T10:15:00.000Z",
    "fix": "dry_run",
    "faultActive": true,
    "rules": [
      "This is a DRILL started from UpButler: a rehearsal. No customer is affected …",
      "Respond exactly as you would to a real incident: claim it, work out the cause, …",
      "Post your diagnosis as an internal note with \"action\": \"diagnosis\" (POST /notes). …"
    ]
  },
  "incident": { "id": "inc_…", "title": "[DRILL] Checkout is down", "…": "…" },
  "monitors": [{ "id": "mon_…", "state": "down", "lastCheck": { "statusCode": 500, "…": "…" } }]
}

A responder paged directly by a drill (not through the escalation policy) receives the event incident.drill_page instead of incident.escalated. Its payload has incident and responder {id, name}; it goes to that responder only.

How to do well

  • Claim quickly. Full points within a quarter of your responder's ack timeout (at least 60 s).
  • Say what you found. Post an internal note with "action": "diagnosis" that names the cause. Notes with that action are judged first; any other note you write counts too.
  • Verify before resolving. Run POST /incidents/:id/verify and resolve only after it passes. In a drill, verify reads the simulated fault, not the real monitors.
  • Hand back what you cannot fix. An upstream outage is not yours to resolve; rolling back would not help either.
  • Stay in scope. Do not roll back, stop runs or mute when that would not help with what you found.
Post a diagnosis
curl -X POST https://upbutler.com/api/v1/incidents/inc_…/notes \
  -H "Authorization: Bearer $INCIDENT_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{"action": "diagnosis", "message": "Deploy drill-2026-10-10 (commit d1211ab) went out 4 minutes before checkout started returning 500. Rolling it back."}'

Simulated actions

Rollback, stop runs, mute and public updates are recorded in a drill, never executed. The API answers as usual with simulated: true, and the incident timeline gets a note "[DRILL] Simulated, nothing was changed: …". When the action is the scenario's remedy (rollback for a bad deploy, stopping the run for a runaway run) the simulated fault ends, so the next verify passes. Public updates are kept for the report and published nowhere.

Fix with Claude in a drill

A drill decides what a Fix with Claude responder may do with fix:

  • dry_run (default): the coding agent reports the change it would make in a note. It does not edit files, commit, push or open a pull request.
  • draft_pr: it may open a draft pull request titled "[DRILL] …" on the branch upbutler/drill-<incident id>. It is a real pull request: close it after the drill. Never merge it.

How each delivery enforces it:

  • GitHub Actions. Drills are dispatched as repository_dispatch event type upbutler-drill, not upbutler-incident. A workflow generated before drills existed does not listen to it, so it never starts and cannot open a real pull request; the drill report then says the responder did not answer. The current workflow listens to both and treats upbutler-drill as a dry run unless the page allows a draft. Regenerate the workflow under Settings → Agent responders and commit it to take part in drills.
  • Local listener (upbutler responder listen --fix). The current CLI never pushes in a dry run and opens only a --draft pull request in draft mode. A CLI from before drills rejects the upbutler/drill-… branch before touching git and hands back ("unexpected branch"); that is safe, and the report carries a hint to update the CLI.
  • Claude Code routines and self-driving agents have no harness around them. There the dry run rests on the page prompt alone ("DRY RUN. Do not edit files, do not commit, do not push and do not open a pull request"). If such an agent reports a pull request in a dry run anyway, "Actions within scope" scores 0 and the report shows the pull request so you can close it.

Generic GitHub Actions responders without Fix with Claude still get upbutler-incident, with packet.drill = true. Drills never count against Fix with Claude attempt limits.

Scoring

100 points over six criteria. Every criterion in the report comes with one sentence that explains its points. Grades: A from 90, B from 75, C from 60, D from 40, F below.

CriterionPointsRule
Time to page10From the incident opening to the first page that left for an agent responder. Full points within 60 s, half within 5 minutes, none later or when nobody was paged.
Time to claim20From the page to the first claim or acknowledgement, against the responder's own ack timeout. Full points within a quarter of it (at least 60 s), then falling to half at the ack timeout; a quarter after it; none if nobody claimed.
Correct diagnosis30A note names the cause: 20 points, plus 10 when it came within 5 minutes of the claim or 5 within 10 minutes. Without a match, up to 10 points for how much of the cause the closest note covered. The check is a keyword match per scenario; an optional AI grader can accept different wording (score 0.7 or more), never overturn a match. Each AI grading uses one AI report from the monthly quota; when it is used up, only the keyword check runs and the report says so.
Verified before resolving15Resolved: full points only if a passing verify came first. Handed back on a scenario that should be handed back: full points. Otherwise half if verify ran at all.
Escalated appropriately15Resolved what could be fixed, handed back what could not. Handing back a fixable fault scores full points when the responder is not allowed the remedy (rollback, stop runs), half when it was. Timing out scores 0.
Actions within scope10Minus 5 for every simulated action that would not have helped (rolling back an upstream outage, stopping runs on a bad deploy, …). A real pull request in a dry-run drill scores 0.

A drill ended by hand (cancelled) or stopped because a real incident opened gets no escalation points; a cancelled drill has no score at all and does not count toward the monthly limit.

Clean-up

  • A drill ends by itself at its time limit (timeoutMinutes, default 15, 2 to 120).
  • It ends at once when a real incident opens on one of its monitors.
  • When it ends, the synthetic incident is deleted, its incident tokens are revoked, queued pages are closed, the escalation run is removed and pending deliveries are dropped. Nothing remains in incident lists, statistics or digests.
  • Anything that should not exist (a mute that names the drill, a component override, a real rollback or run stop) is removed or reported under "Clean-up" in the report. A clean drill says "Nothing left behind".

Schedules

Run a drill weekly, monthly (the first chosen weekday of the month) or quarterly, at an hour in the workspace timezone. With several scenarios each run takes the next one in turn. A run is skipped, not postponed, when one of its monitors has a real problem or the monthly limit is reached; the reason shows on the schedule. The latest score appears in the weekly digest.

Create schedules on the Drills page, with the API, or declare them in upbutler.yaml (setup as code). Schedules from the file are marked as managed and are changed there:

upbutler.yaml
project: acme-web
monitors:
  - id: checkout
    name: Checkout
    url: https://acme.dev/api/checkout
drills:
  - id: game-day
    name: Monthly game day
    schedule: monthly          # weekly | monthly | quarterly
    scenarios: [bad_deploy, upstream_outage, down]
    monitors: [checkout]       # default: the monitors in this file
    weekday: 2                 # 1 = Monday … 7 = Sunday
    hour: 10                   # workspace timezone
    fix: dry_run               # or draft_pr
    timeoutMinutes: 15
    includeHumans: false
API
curl -X POST https://upbutler.com/api/v1/drill-schedules \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "Monthly game day", "cadence": "monthly", "scenarios": ["bad_deploy", "upstream_outage", "down"], "monitorIds": ["mon_…"], "weekday": 2, "hour": 10}'
FieldTypeDescription
namerequiredstringe.g. "Monthly game day"max length 80
cadencerequiredstringweekly, monthly (first chosen weekday of the month) or quarterlyweeklymonthlyquarterly
scenariosrequiredstring[]One scenario, or several: each run takes the next one in turndowndegradedslowbad_deployupstream_outagerunaway_runmax 6 items
monitorIdsrequiredstring[]max 20 items
responderIdsstring[]max 5 items
includeHumansboolean
fixstringdry_rundraft_pr
timeoutMinutesintegermin 2 · max 120
hourintegerHour of day in the workspace timezone (default 10)min 0 · max 23
weekdayinteger1 = Monday … 7 = Sunday (default 2, Tuesday)min 1 · max 7
enabledboolean

Events

Every event about a drill incident carries drill {id, humans, channelIds, fix}. These go to the event stream and GET /events only, never to alert channels or subscribers:

EventWhendata
drill.startedA drill started.drill {id, scenario: null, monitors, endsAt, url}, by. The scenario stays hidden while it runs.
drill.completedA drill ended on resolve, hand-back, timeout or a real incident.drill {id, scenario, outcome, score {total, grade}, url}
drill.cancelledA drill was ended by hand.drill {id, scenario, outcome: "cancelled", score: null, url}
incident.drill_pageA responder chosen for the drill was paged directly.incident, responder {id, name}. Delivered to that responder only.

Plan limits

PlanDrills per monthSchedules
Free2by hand only
Starter101
Pro505
Business50025

Drills started by hand, through the API and by schedules all count; cancelled ones do not. The REST answer to POST /drills carries X-UpButler-Quota-* headers.

API

  • GET /api/v1/drill-scenarios: the scenarios, how scoring is weighted and this month's usage
  • POST /api/v1/drills: start a drill
  • GET /api/v1/drills: drills, newest first (?limit=, up to 100)
  • GET /api/v1/drills/:id: a drill and, once it ended, its report. Takes the drill id or the id of its incident
  • POST /api/v1/drills/:id/end: end a running drill now (no score)
  • GET /api/v1/drill-schedules, POST, PATCH /api/v1/drill-schedules/:id, DELETE

Start a drill

curl -X POST https://upbutler.com/api/v1/drills \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"scenario": "bad_deploy", "monitorIds": ["mon_…"], "timeoutMinutes": 15}'
201 Created
{
  "id": "drl_…",
  "status": "running",
  "scenario": "bad_deploy",
  "scenarioLabel": "Bad deploy",
  "incidentId": "inc_…",
  "monitors": [{ "id": "mon_…", "name": "Checkout", "kind": "http" }],
  "responderIds": [],
  "policy": { "id": "esc_…", "name": "Default" },
  "includeHumans": false,
  "fix": "dry_run",
  "timeoutMinutes": 15,
  "startedAt": "2026-10-10T10:00:00.000Z",
  "endsAt": "2026-10-10T10:15:00.000Z",
  "outcome": null,
  "score": null,
  "url": "https://upbutler.com/app/drills/drl_…"
}
FieldTypeDescription
scenariorequiredstringdown = origin 503 everywhere · degraded = a third of requests fail · slow = latency and timeouts · bad_deploy = a made-up deploy followed by a regression (rollback ends it) · upstream_outage = a third-party dependency is failing (the right answer is to hand back) · runaway_run = an agent run looping over budget (stopping it ends it)downdegradedslowbad_deployupstream_outagerunaway_run
monitorIdsstring[]Monitors the synthetic incident is about. Their real state, history and status-page components are never touchedmax 20 items
componentIdsstring[]Status-page components: the drill runs on the monitors behind them. The components themselves keep their statusmax 20 items
responderIdsstring[]Agent responders (id or name) to page directly. Default: whatever the escalation policy for these monitors pages; with no policy and exactly one responder, that onemax 5 items
includeHumansbooleanDefault false: only agent responders are paged, and levels that page people are recorded as "would have paged". true = people and the alert channels in channelIds really get [DRILL] alerts
channelIdsstring[]With includeHumans: the only alert channels that may receive this drill (default: the channels of the escalation policy and of the monitors)max 20 items
fixstringFix with Claude in the drill. dry_run (default) = the agent reports what it would change and opens nothing · draft_pr = it may open a draft pull requestdry_rundraft_pr
timeoutMinutesintegerThe drill ends by itself after this long (default 15)min 2 · max 120
recoverAfterSecintegerScenarios that heal by themselves (down, degraded, slow) do so after this many seconds (default 120)min 0 · max 3,600

Errors: conflict when a drill is already running or a monitor has a real problem, validation_failed when there is nobody to page (no policy covers the monitors and several or no responders exist, or the policy pages only people and people are not included), plan_limit when this month's drills are used up.

Read the report

While the drill runs, only its status and timing are returned: the cause stays hidden so a responder cannot read the answer. hints lists setup problems the drill brought to light, such as a workflow from before drills.

curl https://upbutler.com/api/v1/drills/drl_… \
  -H "Authorization: Bearer $UPBUTLER_API_KEY"
200 OK (abridged)
{
  "id": "drl_…",
  "status": "completed",
  "outcome": "resolved",
  "score": {
    "total": 85, "max": 100, "grade": "B",
    "criteria": [
      { "id": "page", "label": "Time to page", "points": 10, "max": 10, "seconds": 2,
        "detail": "The first page left 2 s after the incident opened." },
      { "id": "diagnosis", "label": "Correct diagnosis", "points": 25, "max": 30, "seconds": 410,
        "detail": "Named the cause 7 min after the claim." }
    ],
    "diagnosis": { "matched": true, "deterministic": 1, "note": "Deploy drill-2026-10-10 broke checkout…" }
  },
  "groundTruth": "Deploy drill-2026-10-10 (commit d1211ab) went out 4 minutes before the first failing check…",
  "expected": "resolve",
  "faultClearedBy": "rollback",
  "simulatedActions": [{ "at": "…", "action": "rollback", "by": { "type": "agent", "name": "Claude" }, "detail": "rollback to 9f8e7d6" }],
  "skippedPages": [],
  "cleanup": { "mutesRemoved": [], "tokensRevoked": 1, "pagesClosed": 0, "notes": [] },
  "hints": [],
  "timings": { "pagedAt": "…", "claimedAt": "…", "claimedBy": "Claude", "resolvedAt": "…", "handedBackAt": null, "verifies": [{ "at": "…", "passed": true }], "pullRequest": null },
  "timeline": [{ "at": "…", "by": "Claude (agent)", "action": "diagnosis", "status": "investigating", "body": "…" }]
}