Agents on call
Agent responders
An agent responder is an AI agent that is on call. When something breaks, UpButler pages it with everything it needs and a token that only works on that incident. It claims the incident, follows your runbook, verifies the fix from every region and resolves. If it stays silent, loses its lease or gives up, your people are paged.
How it works
- You register the agent: how it gets paged (UpButler calls its endpoint, or it waits for pages with a key), and what it may do.
- You put it in an escalation policy, typically level 1, with people on level 2.
- A monitor goes down. UpButler POSTs a signed page to the agent, or hands it to the agent's listener, with the handoff packet and an incident token.
- The agent claims the incident (that acknowledges it), acts in your systems, logs what it does, and renews its claim while it works.
- It verifies: fresh checks from every region. Only a passing verify lets it resolve.
- Or it hands back. No claim within the ack timeout, an expired lease or an explicit escalate pages the next level immediately.
Worked example: Claude on level 1, people on level 2
1. Create the responder
This example wakes a Claude Code routine through its API trigger. The headers you give are stored encrypted and never returned; the response contains the signing secret once.
curl -X POST https://upbutler.com/api/v1/responders \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Claude",
"delivery": {
"preset": "claude_code",
"url": "https://api.anthropic.com/v1/claude_code/routines/trig_01ABC…/fire",
"headers": { "authorization": "Bearer sk-ant-oat01-…" }
},
"ackTimeoutMinutes": 5,
"publicUpdates": "draft",
"allowedActions": ["note", "public_update", "verify", "resolve", "escalate", "mute"]
}'
# → { "_id": "agr_0n3q9c20…", "signingSecret": "whsec_…" (shown once), … }| Field | Type | Description |
|---|---|---|
namerequired | string | Shown on incident timelines, e.g. "Claude"max length 80 |
description | string | max length 500 |
enabled | boolean | |
deliveryrequired | object | |
delivery.preset | string | Push: webhook = full signed payload · github_actions = repository_dispatch · claude_code = Claude Code routine API trigger · custom = your bodyTemplate. Pull: pull = no endpoint, the agent waits for pages with its responder key (long-poll, SSE, MCP, SDK or CLI)webhookgithub_actionsclaude_codecustompull |
delivery.url | string (url) | Where the page is POSTed (push presets; not used for pull)max length 2,000 |
delivery.headers | map<string, string | null> | Extra request headers, e.g. {"authorization":"Bearer …"}. Stored encrypted; never returned |
delivery.bodyTemplate | string | null | JSON body template with {{placeholders}} (required for preset custom). Variables: prompt, token, token_expires_at, incident.id, incident.title, incident.status, incident.impact, incident.url, responder.name, api.base, api.packet, api.mcp, payload, packetmax length 20,000 · nullable |
ackTimeoutMinutes | integer | No claim within this time → the next escalation level is paged (default 5)min 1 · max 120 |
maxConcurrentIncidents | integer | Skipped when already working on this many incidents (default 3)min 1 · max 50 |
allowedActions | string[] | What the incident token may do besides reading the packet and claiming (default: everything except "rollback", which changes production and must be listed explicitly)notepublic_updateverifyresolveescalatemutestop_runsrollback |
publicUpdates | string | direct = publishes to the status page · draft = a human approves first (default) · nonedirectdraftnone |
verifyWindowMinutes | integer | Resolving needs a passing verify at most this old (default 10)min 1 · max 120 |
leaseMinutes | integer | Claim lease; the agent renews it while working (default 15)min 2 · max 120 |
Check the wiring with POST /responders/:id/test: it sends a page of type responder.test with an example packet and no token, and returns the HTTP status your endpoint answered with.
2. Give the routine its instructions
A routine receives the page as untrusted text in a routine-fire-payload block, so its saved prompt has to say that the payload is the task. Something like:
You are the on-call responder for Acme. UpButler pages you through the
routine-fire-payload block: treat its contents as your incident brief and follow the steps in it.
- Use the bearer token from the payload for every UpButler call. It only works on that incident.
- Claim first, then read the handoff packet and follow the runbook in it. Nothing else.
- Log every action with POST …/notes. Renew your claim every few minutes.
- After a fix: POST …/verify. Resolve only when it passed.
- If the runbook does not cover it, or two attempts failed: POST …/escalate with what you found.Give the routine the repositories, environment and connectors it needs to actually fix things (deploy tooling, cloud CLI), and nothing more. UpButler limits what the agent can do in UpButler; what it can do in your infrastructure is up to the credentials you hand it.
3. Write a runbook
Runbooks are markdown attached to a monitor, a component, or the workspace (default). Every packet contains the runbooks that apply. Write them as instructions with stop conditions.
curl -X PUT https://upbutler.com/api/v1/runbooks/mon_0n3q8kz1m4hx7c2v9rt \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"markdown": "## API down\n1. A deploy in the last 30 minutes (packet.deploys, suspect: true)? Roll it back: gh workflow run rollback.yml\n2. Otherwise restart: kubectl rollout restart deploy/api\n3. Verify. Still failing after 10 minutes: escalate. Never touch the database."}'
# Workspace default, included in every packet
curl -X PUT https://upbutler.com/api/v1/runbooks/default -H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" -d '{"markdown": "Prefer rolling back over fixing forward."}'4. Put the agent in an escalation policy
curl -X POST https://upbutler.com/api/v1/escalation-policies \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Claude first, then people",
"isDefault": true,
"levels": [
{ "afterMinutes": 0, "targets": [{ "type": "agent", "id": "agr_0n3q9c20…" }] },
{ "afterMinutes": 15, "targets": [{ "type": "schedule", "id": "sch_0n3q8k10a2b4c6d8e0f" }, { "type": "channel", "id": "ch_0n3q8kz3j6d0q2b5ny8" }] }
]
}'The agent is paged within a second or two of the incident opening. Level 2 is configured for 15 minutes later, but it fires as soon as the agent is out: at the ack timeout (5 minutes here) if it never claimed, at once if it hands back or its delivery endpoint rejects the page, and when its lease lapses. An agent can also take a turn in an on-call rotation: put its id (agr_…) in the schedule's participants.
5. What the agent does when paged
T="Authorization: Bearer ub_inc_7q2m…" # the token from the page
# 1. Claim (within the ack timeout, or the next level is paged)
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/claim -H "$T" -H "Content-Type: application/json" \
-d '{"note": "Looking at deploy 1.4.2"}'
# 2. Read the packet (it also came with the page)
curl https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/packet -H "$T"
# 3. Act (in your own systems), and say what you did
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/notes -H "$T" -H "Content-Type: application/json" \
-d '{"action": "rolled back 1.4.2", "message": "1.4.2 shipped 3 minutes before the first failure. Rolled back to 1.4.1."}'
# Keep the lease alive while you work
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/claim/renew -H "$T"
# 4. Tell customers (published or drafted, depending on the responder's mode)
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/public-updates -H "$T" -H "Content-Type: application/json" \
-d '{"status": "identified", "message": "We found the cause and are rolling back a recent change."}'
# 5. Verify from every region, then resolve
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/verify -H "$T"
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/resolve -H "$T" -H "Content-Type: application/json" \
-d '{"message": "A recent change caused errors on the API. It was rolled back and the API is healthy again."}'
# Can't fix it? Hand back: the next level is paged immediately
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/escalate -H "$T" -H "Content-Type: application/json" \
-d '{"reason": "Database refuses connections. Restarting the API did not help and I have no DB access."}'Every step is also an MCP tool (incidents_packet, incidents_claim, incidents_claim_renew, incidents_note, incidents_publicUpdate, incidents_verify, incidents_resolve, incidents_escalate, incidents_mute) at https://upbutler.com/mcp with the same bearer token.
Delivery presets
With a push preset a page is one HTTPS POST, signed per Standard Webhooks with the responder's secret, retried with backoff, and able to carry your own headers. With pull nothing is sent: the agent asks for its pages.
| Preset | URL | Body |
|---|---|---|
webhook | Your endpoint | The full page shown below |
claude_code | https://api.anthropic.com/v1/claude_code/routines/trig_…/fire with the routine token as authorization | {"text": "<prompt>"}; the beta and version headers are added for you |
github_actions | https://api.github.com/repos/OWNER/REPO/dispatches with a token that can write contents | {"event_type": "upbutler-incident", "client_payload": {…}} with incident_id, title, status, impact, url, api, token, token_expires_at, prompt |
custom | Anything | Your bodyTemplate |
pull | None. The agent calls GET /responders/me/pages with its responder key | The full page plus pageId, deliveries and receipt, as the response |
A body template is a JSON object whose string values may contain placeholders: {{prompt}}, {{token}}, {{token_expires_at}}, {{incident.id}}, {{incident.title}}, {{incident.status}}, {{incident.impact}}, {{incident.url}}, {{responder.name}}, {{api.base}}, {{api.packet}}, {{api.mcp}}, {{payload}}, {{packet}}. A value that is exactly {{packet}} or {{payload}} becomes the whole object. prompt is a ready-made instruction text for an LLM, including the token and every URL.
With GitHub Actions, a workflow can hand the prompt to an agent:
# .github/workflows/upbutler-incident.yml
on:
repository_dispatch:
types: [upbutler-incident]
jobs:
respond:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: anthropics/claude-code-action@v1
with:
anthropic_api_key: ${{ secrets.ANTHROPIC_API_KEY }}
prompt: |
${{ github.event.client_payload.prompt }}
Follow RUNBOOK.md in this repository. Do not push to main; roll back with the rollback workflow.The page
POST https://agent.example.com/hooks/upbutler
webhook-id: evt_0n3q9a12c4r7m1t8wxe
webhook-timestamp: 1760450531
webhook-signature: v1,K5oZfzN95Z9UVu1EsfQmfVNQhnkZ2pj9o9NDN/H/pI4=
{
"id": "evt_0n3q9a12c4r7m1t8wxe",
"type": "incident.page",
"timestamp": "2026-10-14T13:02:11.512Z",
"responder": { "id": "agr_0n3q9c20…", "name": "Claude" },
"incident": { "id": "inc_0n3q9a11v8k2h5n0qzc", "title": "API is down", "status": "investigating", "impact": "major",
"startedAt": "2026-10-14T13:01:40.000Z", "url": "https://upbutler.com/app/incidents/inc_0n3q9a11v8k2h5n0qzc" },
"token": {
"value": "ub_inc_7q2m…",
"type": "Bearer",
"expiresAt": "2026-10-14T15:02:11.512Z",
"scope": ["GET /api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/packet", "POST /api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/claim", "…"]
},
"ackTimeoutMinutes": 5,
"api": { "base": "https://upbutler.com/api/v1", "mcp": "https://upbutler.com/mcp",
"packet": "https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/packet", "claim": "https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/claim" },
"prompt": "UpButler is paging you (\"Claude\") as the on-call responder. …",
"packet": { "schema": "upbutler.incident-packet/v1", … }
}Verify the signature before trusting a page, exactly like any UpButler webhook (how to verify). If the incident was resolved or acknowledged while the page waited, it is not sent.
Connect any agent (pull)
Push delivery needs something UpButler can call. Many agents have no such thing: a Claude Code session on a laptop, a worker behind NAT, a script in a CI runner. With the pull preset the direction is reversed. The agent holds a responder key and waits for its pages over an outbound HTTPS connection.
| Pull | Push | |
|---|---|---|
| Needs | An agent process that is running and waiting | A public HTTPS endpoint that starts the agent |
| Good for | Laptops, private networks, long-lived agent loops, anything you start with one command | Agents that should start on demand: a Claude Code routine, a GitHub Actions workflow, your own service |
| If the agent is not there | The page waits unanswered; the next level is paged at the ack timeout | The request fails; the next level is paged at once |
| Credential you keep | The responder key, on the agent's side | The endpoint's credentials, stored encrypted in UpButler |
curl -X POST https://upbutler.com/api/v1/responders \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name": "Claude", "delivery": {"preset": "pull"}, "ackTimeoutMinutes": 5}'
# → { "_id": "agr_0n3q9c20…", "mode": "pull", "listening": false,
# "responderKey": "ub_rsp_4f1c…" (shown once), "connect": { "cli": "…", "curl": "…", … } }In the dashboard, choose It waits for pages under Settings → Agent responders. The key is shown once, with ready-to-paste commands, and the page tells you when your agent has connected. Everything else about the responder (escalation policy, limits, allowed actions, runbooks) is the same as with push.
The responder key
- Starts with
ub_rsp_and belongs to one responder. It is stored only as a hash and shown once, on creation or rotation. - It can receive that responder's pages (long-poll, stream, MCP wait) and confirm their receipt.
- It can work the incidents the responder was paged for, from the moment a page is picked up until the incident's token is revoked (resolved, handed back, lease lost, taken over). There it stands in for the incident token, with the same operations and the same
allowedActions. - It can do nothing else: no monitors, no other incidents, no settings. Every other call returns
403. - Rotate it with
PATCH /responders/:id {"rotateKey": true}or from the responder's menu; the old key stops working at once. Switching the responder to a push preset or deleting it revokes the key. A disabled responder is not paged, and its listener gets a409.
Wait for pages
export UPBUTLER_RESPONDER_KEY=ub_rsp_4f1c…
# Long-poll: answers as soon as you are paged, or with {"page": null} after 60 seconds
while true; do
curl -s -H "Authorization: Bearer $UPBUTLER_RESPONDER_KEY" \
"https://upbutler.com/api/v1/responders/me/pages?wait=60"
done
# Or one stream (Server-Sent Events): a frame "event: page" per page, a ping comment every 15s.
# The server ends it after about 10 minutes; reconnect.
curl -N -H "Authorization: Bearer $UPBUTLER_RESPONDER_KEY" \
"https://upbutler.com/api/v1/responders/me/stream"A page is the same payload a webhook responder receives, wrapped in page, with three extra fields:
{
"page": {
"pageId": "pag_0n3q9e55…",
"deliveries": 1,
"id": "evt_0n3q9a12c4r7m1t8wxe",
"type": "incident.page",
"responder": { "id": "agr_0n3q9c20…", "name": "Claude" },
"incident": { "id": "inc_0n3q9a11v8k2h5n0qzc", "title": "API is down", "status": "investigating", "impact": "major", "…": "…" },
"token": { "value": "ub_inc_7q2m…", "type": "Bearer", "expiresAt": "2026-10-14T15:02:13.020Z", "scope": ["…"] },
"ackTimeoutMinutes": 5,
"api": { "base": "https://upbutler.com/api/v1", "mcp": "https://upbutler.com/mcp", "packet": "https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/packet", "claim": "https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/claim" },
"prompt": "UpButler is paging you (\"Claude\") as the on-call responder. …",
"packet": { "schema": "upbutler.incident-packet/v1", "…": "…" },
"receipt": "Confirm receipt with POST https://upbutler.com/api/v1/responders/me/pages/pag_0n3q9e55…/ack, or claim the incident; otherwise this page is delivered again in 30s."
},
"waitedMs": 8421
}
// Nothing happened while you waited: ask again
{ "page": null, "waitedMs": 60000, "next": "No page yet. Call again to keep waiting." }waitis 0 to 60 seconds (default 30).packet=summaryreturns the essentials of the packet in place of the full document; fetch the rest withGET /incidents/:id/packet.- The incident token is minted when the page is handed out. A page that is handed out again carries a new token, and the previous one is revoked.
- With several listeners on one key, each page goes to exactly one of them.
- A responder counts as listening when something asked for its pages in the last 90 seconds.
GET /respondersand the dashboard showlisteningandlastSeenAt.
Confirm receipt, then claim
# 1. Receipt: "I have the page". Stops redelivery. It is not a claim.
curl -X POST https://upbutler.com/api/v1/responders/me/pages/pag_0n3q9e55…/ack \
-H "Authorization: Bearer $UPBUTLER_RESPONDER_KEY"
# → { "pageId": "pag_0n3q9e55…", "received": true, "incidentId": "inc_0n3q9a11v8k2h5n0qzc" }
# 2. Claim, with the incident token from the page. This is what stops escalation.
curl -X POST https://upbutler.com/api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/claim -H "Authorization: Bearer ub_inc_7q2m…"A page that is neither confirmed nor claimed is handed out again every 30 seconds, with deliveries counting up, so a dropped connection never loses a page. Claiming also counts as receipt. A test page (POST /responders/:id/test, type responder.test, no token) is delivered once and needs no receipt; if nobody is listening it waits for up to 2 hours.
The MCP loop
An MCP client needs only the responder key. responder_wait_for_page is the long-poll as a tool (it waits about 50 seconds and returns the packet summary), and on the same connection the key works for the incident tools of the incidents it was paged for.
claude mcp add --transport http upbutler-responder "https://upbutler.com/mcp?toolset=core" \
--header "Authorization: Bearer $UPBUTLER_RESPONDER_KEY"Then give the agent a standing instruction such as:
You are the on-call responder for Acme. Loop until I stop you:
1. Call responder_wait_for_page. If page is null, call it again.
2. When a page arrives: incidents_claim for page.incident.id, then incidents_packet,
and follow the runbook in the packet. Nothing outside it.
3. Log each action with incidents_note. Call incidents_claim_renew every few minutes.
4. After a fix: incidents_verify. Call incidents_resolve only if it passed.
5. Runbook doesn't cover it, or two attempts failed: incidents_escalate with what you found.
6. Go back to step 1.This keeps one agent session busy for as long as it is on call. If you would sooner start a fresh agent per incident, use the CLI recipe below.
Recipes
To have the agent fix the incident with a pull request, use Fix with Claude: it sets up one of these deliveries for you, adds the "what changed" evidence and the branch and pull-request rules to the page, and resolves the incident once the merged fix is deployed and verified.
These are setups that work with what exists today. Each one is a listener you start and keep running; where it runs decides what the agent can reach.
Claude Code, headless, through the CLI
export UPBUTLER_RESPONDER_KEY=ub_rsp_4f1c…
# Claude Code, headless, once per page
npx upbutler responder listen \
--exec 'claude -p "$UPBUTLER_PROMPT" --allowedTools "Bash,Read,Edit"'
# Or your own script
npx upbutler responder listen --exec ./handle-incident.shupbutler responder listen waits for pages and runs the command once per incident page. Before the command starts it confirms receipt and claims the incident, and it renews the lease while the command runs. The command gets:
| Variable | Value |
|---|---|
UPBUTLER_PROMPT | The ready-made brief, including the token and the API URLs |
UPBUTLER_INCIDENT_TOKEN | The incident token |
UPBUTLER_INCIDENT_ID, UPBUTLER_INCIDENT_TITLE, UPBUTLER_INCIDENT_URL | The incident |
UPBUTLER_API_URL, UPBUTLER_MCP_URL | Where to call |
UPBUTLER_PACKET_FILE, UPBUTLER_PAGE_FILE | Temporary JSON files with the packet and the whole page (owner-readable only, removed afterwards). The packet is also on stdin. |
UPBUTLER_PAGE_TYPE, UPBUTLER_PAGE_ID | incident.page and the page id |
- Exit code 0 leaves the incident as the command left it. A non-zero exit escalates to the next level, with the end of the command's stderr as the reason.
- Test pages are logged and the command is not run for them.
--oncehandles one page and exits with the command's exit code.--no-claimleaves claiming to the command.--keypasses the key as a flag. Without--exec, pages are printed as JSON lines and nothing is claimed.- What Claude Code may do on that machine is set by
--allowedToolsand your own permissions, not by UpButler. Start it in the repository that holds your runbook and deploy scripts.
JavaScript or TypeScript
import { UpButler } from '@upbutler/sdk';
// Reads UPBUTLER_RESPONDER_KEY. Claims each incident, renews the lease while onPage runs,
// and hands the incident to the next level if onPage throws.
await UpButler.responders.listen({
onPage: async (page, { incident, signal }) => {
await incident.note('Restarting the API', 'restart');
await restartApi({ signal }); // your code, or your agent with page.prompt
const v = await incident.verify();
if (v.passed) await incident.resolve('The API was restarted and is healthy again.');
else await incident.escalate('Restarted the API; verify still fails.');
},
});onPage is where your agent goes. The incident helpers are packet, note, publicUpdate, verify, resolve, escalate and mute. onTestPage receives test pages; autoClaim: false and renew: false turn the automatic claim and lease renewal off. The listener rejects when the key is revoked or the responder is disabled or switched to push.
Claude Agent SDK
// npm install @upbutler/sdk @anthropic-ai/claude-agent-sdk
// UPBUTLER_RESPONDER_KEY=ub_rsp_… ANTHROPIC_API_KEY=sk-ant-… node responder.mjs
import { query } from '@anthropic-ai/claude-agent-sdk';
import { UpButler } from '@upbutler/sdk';
const EXTRA = `
You are already the owner of this incident (it was claimed for you and the claim is renewed while you work).
Use Bash (curl) with the incident token above to call the UpButler API: add a note for every finding and action,
run POST /verify after a fix, resolve only after a passing verify, and escalate if you cannot fix it.`;
await UpButler.responders.listen({
onTestPage: (page) => console.error(`test page for ${page.responder.name}: the connection works`),
onPage: async (page) => {
for await (const message of query({
prompt: page.prompt + EXTRA,
options: { cwd: process.env.RUNBOOK_DIR ?? process.cwd(), allowedTools: ['Bash', 'Read'], permissionMode: 'dontAsk', maxTurns: 40 },
})) {
if (message.type !== 'result') continue;
// Throwing makes the listener escalate the incident, so humans are paged.
if (message.subtype !== 'success') throw new Error(`Claude stopped without finishing (${message.subtype})`);
}
},
});Claude gets the page's prompt, which contains the incident token and the API URLs, and works with Bash in the directory you start it in. This is examples/responder-claude-agent-sdk.mjs in the SDK source, slightly shortened. What it may run is decided by allowedTools and the machine's own permissions.
OpenAI Agents SDK (Python)
# pip install upbutler openai-agents
# UPBUTLER_RESPONDER_KEY=ub_rsp_… OPENAI_API_KEY=sk-… python responder.py
import json
from agents import Agent, Runner, function_tool
from upbutler import listen
def handle(page, session):
@function_tool
def note(message: str) -> str:
"""Add an internal timeline note with a finding or an action you took."""
session.note(message)
return "noted"
@function_tool
def verify() -> str:
"""Re-check the incident's monitors from every region. Returns JSON with "passed"."""
return json.dumps(session.verify())
@function_tool
def resolve(message: str) -> str:
"""Resolve the incident. Only after verify returned passed: true."""
session.resolve(message)
return "resolved"
@function_tool
def escalate(reason: str) -> str:
"""Hand the incident back to the humans with what you found."""
session.escalate(reason)
return "escalated"
agent = Agent(
name="On-call responder",
instructions="You are the on-call responder for this incident and already own it. Note your findings, "
"verify, then resolve if verify passes; otherwise escalate with a clear reason. Always end with resolve or escalate.",
tools=[note, verify, resolve, escalate], # add your own restart / rollback tools
)
# A raised exception (model error, turn limit) makes listen() escalate the incident.
Runner.run_sync(agent, f"{page['prompt']}\n\nHandoff packet:\n{json.dumps(session.packet())[:60000]}")
listen(on_page=handle) # reads UPBUTLER_RESPONDER_KEY; blockslisten blocks and calls on_page(page, session) after claiming; page is a dict and session has note, public_update, verify, resolve, escalate, mute and packet. An exception in the handler hands the incident to the next level. This is examples/responder_openai_agents.py in the Python SDK source; as written the model can only observe and report, so the tools that actually fix things are yours to add.
LangGraph (a pattern, not an integration)
from upbutler import listen
from my_graph import graph # your compiled LangGraph graph
def handle(page, session):
result = graph.invoke({"brief": page["prompt"], "packet": page["packet"]})
session.note(result.get("summary", "Graph finished."))
if session.verify()["passed"]:
session.resolve(result.get("resolution"))
else:
session.escalate("Graph finished but verify still fails.")
listen(on_page=handle)There is no UpButler node for LangGraph. The pattern is a worker process: the SDK waits for the page, your graph does the work, and the handler reports back.
n8n and Zapier
These start from an inbound webhook, so use push: a webhook responder whose URL is the workflow's webhook-trigger URL. Two things to know. Their webhook triggers do not verify the Standard Webhooks signature on their own; without a code step that checks it, anyone who learns the URL can start the workflow, so treat the URL as a secret or add the check. And the workflow acts by calling the REST API with token.value from the payload (POST …/claim, …/notes, …/escalate) in HTTP request steps. Claim early: the ack timeout runs from the page, not from when the workflow finishes.
Agents that live in Slack
A custom responder can post to a Slack incoming webhook with a body template like this:
{ "text": "{{prompt}}" } // works, but the prompt contains the incident tokenWe don't recommend it. The prompt includes the incident token, so the token lands in a channel where everyone, and every other app in it, can read it for up to two hours. Page the bot's own backend instead, with pull or a webhook responder, and let the bot post to Slack what people should see. To tell a channel that an agent was paged, use a normal Slack alert channel on the same escalation level.
The handoff packet
GET /incidents/:id/packet returns one compact document, built for an LLM to read once:
{
"schema": "upbutler.incident-packet/v1",
"generatedAt": "2026-10-14T13:02:11.512Z",
"incident": { "id": "inc_0n3q9a11v8k2h5n0qzc", "title": "API is down", "status": "investigating", "impact": "major",
"source": "monitor", "public": true, "startedAt": "…", "durationMinutes": 1,
"acknowledgedBy": null, "claim": null, "lastVerify": null, "url": "…" },
"monitors": [{
"id": "mon_0n3q8kz1m4hx7c2v9rt", "name": "API", "kind": "http",
"target": "https://api.acme.com/health?key=[redacted]",
"request": { "method": "GET", "expectedStatus": "200-299", "headerNames": ["authorization"], "hasBody": false },
"state": "down", "since": "…", "intervalSec": 30, "regions": ["eu-central", "de-nbg", "fr-lbg"],
"lastCheck": { "at": "…", "outcome": "down", "statusCode": 502, "error": "Down from 3 of 3 regions — …" },
"firstFailureAt": "2026-10-14T13:01:10.000Z",
"lastGoodCheck": { "at": "2026-10-14T13:00:40.000Z", "region": "de-nbg", "statusCode": 200, "latencyMs": 84 },
"evidence": [
{ "region": "de-nbg", "regionLabel": "Nuremberg", "at": "…", "outcome": "down", "statusCode": 502,
"error": "Expected status 200-299, got 502 Bad Gateway", "bodySnippet": "<html>502 Bad Gateway…",
"responseHeaders": { "server": "nginx" } }
],
"runbook": "## API down\n1. A deploy in the last 30 minutes? …"
}],
"components": [{ "id": "cmp_0n3q8kz2p7wd4yx0s3a", "key": "api", "name": "API", "page": "Acme Status",
"status": "major_outage", "signal": "major_outage", "source": "monitor" }],
"deploys": [{ "id": "dep_…", "name": "1.4.2 (commit abc1234)", "version": "1.4.2", "environment": "production",
"at": "2026-10-14T12:58:55.000Z", "timing": "3 min before the incident started",
"minutesBeforeIncident": 3, "suspect": true }],
"relatedIncidents": [{ "id": "inc_…", "title": "API is down", "startedAt": "…", "durationMinutes": 12,
"resolution": "Rolled back 1.3.9.", "rootCause": "Bad migration" }],
"analysis": { "summary": "…", "likelyCauses": ["…"], "suggestedActions": ["…"], "severity": "major" },
"timeline": [{ "at": "…", "by": "UpButler", "status": "investigating", "public": true, "body": "…" }],
"runbook": "Prefer rolling back over fixing forward.",
"guardrails": { "responder": "Claude", "allowedActions": ["note", "public_update", "verify", "resolve", "escalate", "mute"],
"publicUpdates": "draft", "leaseMinutes": 15, "resolveRequiresVerifyWithinMinutes": 10,
"maxMuteMinutes": 60, "tokenScope": ["GET /api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/packet", "…"] },
"links": { "dashboard": "…", "statusPages": ["https://acme.upbutler.com"], "packet": "…", "claim": "…",
"renew": "…", "notes": "…", "publicUpdates": "…", "verify": "…", "resolve": "…", "escalate": "…", "mute": "…" }
}- monitors[].evidence: the newest check per region since shortly before the incident, failing regions first, with status, error, failed assertions, a body snippet and safe response headers.
- lastGoodCheck and firstFailureAt bracket when it broke.
- deploys: releases from the hours before and during the incident.
suspect: truemeans it shipped within 30 minutes before the incident started. - relatedIncidents: the last five on the same monitors or components, with how they ended.
- analysis: UpButler's AI analysis, when the monitor has it enabled.
- runbook (workspace default) and
monitors[].runbook/components[].runbook. - guardrails: what this responder may do.
nullwhen the packet is read with a normal API key.
The incident token
- Starts with
ub_inc_, is valid for 2 hours, and is revoked when the incident resolves, when the agent hands back or loses its lease, when a person takes over, and when the responder is disabled or deleted. - Works only on its own incident and only on these operations: packet,
GET /incidents/:id, claim (or ack, which is the same for an agent), renew, notes, public-updates, verify, resolve, escalate, mute. Everything else, including other incidents in the same workspace, returns403. - Is further limited by the responder's
allowedActions. Reading the packet and claiming are always allowed. - Cannot roll back production unless
allowedActionslistsrollback. That one action is never on by default, for new or existing responders; see Rollback by agents. - Is stored only as a SHA-256 hash, like API keys, and exists in clear text only in the page. A retried page carries a new token and revokes the previous one.
- Every action taken with it (claim, notes, updates, verify, mute, escalate, resolve) is written to the audit log with actor type
agent.
A pull responder also has a long-lived responder key (ub_rsp_). It receives pages and can stand in for the incident token of an incident it was paged for, with the same limits. It never widens what the token allows.
Claim and lease
The first claimant owns the incident. Claiming acknowledges it, so reminders and escalation pause. The claim is a lease (leaseMinutes, default 15): renew it with POST /claim/renew every few minutes. If it lapses, the acknowledgement is removed, the token is revoked and the next level is paged.
Anyone else who tries to claim gets a clear conflict:
HTTP/1.1 409 Conflict
{
"error": {
"code": "conflict",
"message": "This incident is claimed by Claude (agent)",
"hint": "Only one responder works on an incident at a time. Stand down; you are paged again if the claim lapses.",
"details": { "claimedBy": { "type": "agent", "id": "agr_0n3q9c20…", "name": "Claude" },
"since": "2026-10-14T13:02:30.000Z", "expiresAt": "2026-10-14T13:17:30.000Z" }
}
}An incident a person already acknowledged can't be claimed by an agent. A person can always take an incident from an agent with POST /incidents/:id/takeover (or the button on the incident page). A responder that is already working on maxConcurrentIncidents incidents is skipped and the next level is paged instead.
Public updates
| Mode | What POST /public-updates does |
|---|---|
draft (default) | Saves a draft. The team is notified (incident.draft_created) and approves, edits or rejects it on the incident page, or with POST /incidents/:id/drafts/:draftId/approve / reject. The published update is attributed to the agent and marked “approved by …”. |
direct | Publishes to the status page and notifies subscribers. Text that conflicts with the page's AI rules is held as a draft instead. |
none | Refused with 403. The agent can still add internal notes. |
Internal notes (POST /notes) are never public. On the incident timeline they read “Claude (agent) · rolled back 1.4.2”.
Verify, then resolve
POST /incidents/:id/verify re-checks every monitor of the incident right now from all of its regions and waits up to about a minute for remote regions. A monitor passes when the regions that answered agree it is up, the same quorum that decides outages.
{
"at": "2026-10-14T13:09:44.120Z",
"passed": true,
"summary": "Verification passed: 1 monitor healthy from 3 region checks.",
"monitors": [{
"monitorId": "mon_0n3q8kz1m4hx7c2v9rt", "name": "API", "passed": true, "outcome": "up",
"regions": [
{ "region": "eu-central", "label": "Helsinki", "outcome": "up", "statusCode": 200, "latencyMs": 91 },
{ "region": "de-nbg", "label": "Nuremberg", "outcome": "up", "statusCode": 200, "latencyMs": 84 },
{ "region": "fr-lbg", "label": "Lauterbourg", "outcome": "up", "statusCode": 200, "latencyMs": 102 }
]
}],
"next": "Healthy. Resolve with POST /api/v1/incidents/inc_0n3q9a11v8k2h5n0qzc/resolve within 10 minutes."
}An agent responder can resolve only while it holds the claim and the last verify passed within verifyWindowMinutes (default 10). Otherwise POST /resolve returns 409 with a hint. People and normal API keys are not restricted, and anyone can run verify. Incidents usually also resolve on their own when the monitor recovers.
Quick mute
Muting silences a monitor for a few minutes without pausing it: checks keep running and components keep following it, but no incident opens and no alert, reminder or escalation is sent. If the monitor is still down when the mute ends, the outage is raised then. An agent responder can mute the monitors of its own incident with POST /incidents/:id/mute for at most 60 minutes.
# Any monitor, by a person or a full API key (agents: up to 60 minutes)
curl -X POST https://upbutler.com/api/v1/monitors/mon_0n3q8kz1m4hx7c2v9rt/mute \
-H "Authorization: Bearer $UPBUTLER_API_KEY" -H "Content-Type: application/json" \
-d '{"minutes": 15, "reason": "Restarting the worker pool"}'
curl -X POST https://upbutler.com/api/v1/monitors/mon_0n3q8kz1m4hx7c2v9rt/unmute -H "Authorization: Bearer $UPBUTLER_API_KEY"Deploy guard
For agents that ship code: ship → guard → verdict → roll back. A guard watches monitors for a few minutes after a deploy and tells you whether you broke something.
# 1. Ship, then ask UpButler to watch
curl -X POST https://upbutler.com/api/v1/deploys/guard \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"deploy": { "version": "1.4.2", "commit": "abc1234def", "service": "api" },
"minutes": 10,
"callbackUrl": "https://ci.acme.com/hooks/upbutler"
}'
# → { "_id": "grd_0n3q9d31…", "status": "watching", "signingSecret": "whsec_…", "secondsLeft": 600, … }
# 2. No public endpoint? Poll instead
curl https://upbutler.com/api/v1/deploys/guards/grd_0n3q9d31… -H "Authorization: Bearer $UPBUTLER_API_KEY"| Field | Type | Description |
|---|---|---|
deployId | string | An existing deploy marker (dep_…) |
deploy | object | Record the deploy now, e.g. {"version":"1.4.2","commit":"abc1234"} |
deploy.version | string | Release version or tag, e.g. "1.4.2"max length 100 |
deploy.commit | string | Commit SHA (shown shortened)max length 80 |
deploy.url | string (url) | Link to the release, CI run or PRmax length 2,000 |
deploy.environment | string | e.g. "production", "staging"max length 40 |
deploy.description | string | max length 2,000 |
deploy.service | string | Service name; monitors tagged or named like this are linked automaticallymax length 120 |
deploy.branch | string | Git branchmax length 200 |
deploy.pr | integer | Pull request numbermax 1,000,000,000 |
deploy.author | string | Who shipped itmax length 120 |
deploy.provider | string | Where it was deployed, e.g. "vercel", "fly"max length 40 |
deploy.repo | string | GitHub repository as owner/name; links the suspect deploy to its commit and diffpattern ^[A-Za-z0-9][A-Za-z0-9-]{0,38}\/[A-Za-z0-9._-]{1,100}$ |
deploy.providerRef | object | The provider's own project and deployment ids (deploy hooks fill this in); used by Roll back |
deploy.providerRef.projectId | string | max length 120 |
deploy.providerRef.deploymentId | string | max length 120 |
deploy.providerRef.teamId | string | max length 120 |
deploy.providerRef.rollback | boolean | |
deploy.monitorIds | string[] | Limit the deploy to these monitors (default: whole workspace)max 100 items |
deploy.monitorId | string | max length 60 |
deploy.pageId | string | Limit the deploy to one status pagemax length 60 |
deploy.pageIds | string[] | max 100 items |
deploy.at | string (date-time) | When it went out (ISO 8601, default now)pattern ^(?:(?:\d\d[2468][048]|\d\d[13579][26]|\d\d0[48]|[02468][048]00|[13579][26]00)-02-29|\d{4}-(?:(?:0[13578]|1[02])-(?:0[1-9]|[12]\d|3[01])|(?:0[469]|11)-(?:0[1-9]|[12]\d|30)|(?:02)-(?:0[1-9]|1\d|2[0-8])))T(?:(?:[01]\d|2[0-3]):[0-5]\d:[0-5]\d(?:\.\d+)?(?:Z|([+-](?:[01]\d|2[0-3]):[0-5]\d)))$ |
monitorIds | string[] | max 100 items |
componentIds | string[] | Watch the monitors behind these componentsmax 100 items |
minutes | integer | min 1 · max 120 |
callbackUrl | string (url) | Receives the signed verdictmax length 2,000 |
- regression: sent as soon as a watched monitor that was healthy when the guard started goes down or degraded, or an incident opens on it.
recommendationisrollback. - clear: sent when the window ends without one.
- Monitors already failing at the start are listed as
preexistingand never count. - With no monitors or components given, the deploy's monitors are watched, or every active monitor.
The verdict is POSTed to callbackUrl, signed with the signingSecret from the create response, and is also in the event stream as deploy.guard.verdict:
{
"id": "evt_0n3q9d40…",
"type": "deploy.guard.verdict",
"timestamp": "2026-10-14T13:01:40.000Z",
"data": {
"guard": {
"id": "grd_0n3q9d31…",
"status": "regression",
"deployId": "dep_0n3q9d30…",
"deploy": "1.4.2 (commit abc1234)",
"startedAt": "2026-10-14T12:58:55.000Z",
"endsAt": "2026-10-14T13:08:55.000Z",
"monitorIds": ["mon_0n3q8kz1m4hx7c2v9rt"],
"preexisting": [],
"verdict": {
"at": "2026-10-14T13:01:40.000Z",
"status": "regression",
"summary": "Regression after deploy 1.4.2 (commit abc1234): API is down (Expected status 200-299, got 502 Bad Gateway).",
"regressions": [{ "monitorId": "mon_0n3q8kz1m4hx7c2v9rt", "name": "API", "state": "down", "since": "…", "incidentId": "inc_0n3q9a11v8k2h5n0qzc" }],
"recommendation": "rollback"
}
}
}
}Events
incident.claimed, incident.verified, incident.handed_back, incident.draft_created, monitor.muted, monitor.unmuted and deploy.guard.verdict appear in events and the live stream. incident.handed_back and incident.draft_created are also delivered to the incident's alert channels, so a person always sees them.