Skip to content
Docs/Alert actions & rollback

Incidents & alerts

Alert actions & rollback

An alert you can end an outage from, without opening a laptop: one sentence about what your users see, then Roll back, Let Claude fix and Mute 1h.

One plain sentence

Every team alert about something failing or recovering leads with a sentence built from four facts:

PartComes fromExample
WhatThe public name of the status page component the monitor feeds, else the monitor's name (a URL becomes its host)Checkout
EffectThe monitor kind and its stateis failing for users · is slow or partly failing for users · has not been running · can't be reached · is not answering
SinceWhen the incident began, in the workspace timezone (labelled UTC when none is set)since 14:02
DeployThe deploy a deploy guard blamed, else the newest deploy marker from the hour before(4 minutes after deploy abc1234 by Vercel)

The sentence is a template, not AI output, so it is the same every time and never waits for a model. It is the email subject and preheader, the headline in Slack, Discord and Telegram, the notification text on a phone, and summaryPlain in webhooks. Reminders read “Checkout is still failing for users, 25 minutes now (since 14:02, …)” and recoveries “Checkout is working again after 6 minutes.” The target, the error and the AI analysis still follow underneath. Incidents a person wrote by hand keep their own title.

The buttons

Alerts about an open incident (monitor.down, monitor.reminder, incident.created, incident.escalated, incident.handed_back) carry these, in this order:

ButtonShown whenDoesPlan
Roll backA deploy from Vercel or Netlify (or one whose project a connected account knows) went out in the hour before the incident beganPuts the previous production deployment back, then watches for 10 minutesStarter, Pro, Business
Let Claude fixThe workspace has a Fix with Claude responder that can take the incident. Without one the button reads “Set up Claude fixes” and leads to set-upHands the incident and its evidence to your coding agent, which opens a pull requestEvery plan (daily attempts vary)
AcknowledgeNobody acknowledged yetStops reminders and escalation (Alert channels)Every plan
Stop runThe monitor is an agent monitorTells the flagged runs to stopEvery plan
Mute 1hThe incident has monitorsNo alerts, reminders or escalation for those monitors for an hour. Checks keep running; if it is still failing when the hour ends, alerts resumeEvery plan

Email shows the first button large and the rest underneath. Slack uses URL buttons (no Slack app needed), Telegram inline URL buttons, Discord links in an “Act on it” field, and webhooks an actions object. On the Free plan Roll back appears on the first 3 alerts that follow a deploy: its page explains what it would do and how to get it. After that it is left out. “Set up Claude fixes” follows the same rule, and its page has a “Don't show this button again” link. Turn any button off for the whole workspace in Settings → Integrations.

What happens when you tap one

  1. The link opens a confirm page. It shows the plain sentence, what will happen in a few short lines, and one large button. Opening the link does nothing, however often: mail scanners, chat previews and prefetchers open links, and none of them can trigger an action.
  2. You press the button. That sends a POST with a one-time value the page was rendered with. Only then does the action run.
  3. The page shows the result and keeps itself up to date: “Rolled back. Watching whether it recovers…”, then “It worked” or “Still failing after the rollback.”
  4. It is recorded. An internal note on the incident timeline and an audit log entry name who did it: the signed-in member, else the recipient of the alert (“Someone on #ops (slack)”, “Dana (via alert link)”).
  • Signed and scoped. Each link carries an HMAC-signed token bound to one incident, one action and the recipient it was sent to. Changing any of them invalidates it.
  • Short-lived. Links expire after 24 hours and stop working when the incident is resolved or the action is turned off.
  • Single use. A Roll back or Let Claude fix link works once. A second visit shows what happened. If the action failed before anything changed, the link can be tried again.
  • Confirm with POST. The confirm button posts a value tied to a cookie set on that page, from the same origin. A POST without it is refused.
  • Webhook links need a sign-in to roll back. Webhook payloads pass through other systems, so a Roll back link delivered to a webhook channel only works for a signed-in member of the workspace. Links in email, Slack, Discord and Telegram work for whoever received them.
  • Rate limited, per client and per incident.
  • Plan and settings are checked on the server when the button is pressed, not only when the alert is sent.

Roll back

Roll back asks your hosting provider to serve the previous production deployment again. Your repository and the broken deployment are untouched; only where production traffic goes changes. UpButler picks the newest finished production deployment that is older than the suspect one, and the confirm page names it before you press.

ProviderWhat UpButler callsGood to know
VercelPOST /v1/projects/{projectId}/rollback/{deploymentId} (Instant Rollback)On Vercel's Hobby plan only the immediately previous deployment can be restored, which is the one UpButler picks. After a rollback Vercel stops sending new pushes to production: promote the fixed deployment (Undo Rollback in the dashboard, or vercel promote) to switch that back on. Environment variable changes are not rolled back.
NetlifyPOST /api/v1/sites/{site_id}/deploys/{deploy_id}/restoreThe restored deploy stays published until the next production deploy finishes.

Railway and Render are next. Deploys from other hosts get no Roll back button.

After the rollback

  • A deploy marker “Rollback to 9f8e7d6” is recorded, so charts and later incidents show it.
  • A deploy guard starts on it for 10 minutes, to catch anything the rollback itself breaks.
  • As soon as the incident's monitors are healthy again, or after 10 minutes if they are not, the timeline says “The rollback to 9f8e7d6 worked” or “Still failing 10 minutes after the rollback … The deploy was probably not the cause.”
  • The channels that were alerted get incident.rollback twice: when it was rolled back, and with the result.
  • For a public incident, a draft update (“We rolled back a recent change and are watching the results.”) waits for you to approve. Nothing is posted to customers on its own.

Connect Vercel or Netlify

In Settings → Integrations, paste an API token. It is tested against the provider before it is saved, stored encrypted, and never shown again (reads return its last four characters).

  • Vercel: create an access token scoped to the team, or to the single project, you deploy from. A full-account token also needs the team id (team_…) or slug. The token's user must be allowed to roll back: an Owner, Member or Developer, or a Project Administrator.
  • Netlify: a personal access token of a member who can publish deploys.

Deploys that arrive through the provider's own deploy hook carry the provider's project id, so nothing has to be mapped. For markers from CI or the GitHub Action, UpButler matches the marker's service to a project of the same name; when the names differ, map them on the same page.

curl -X POST https://upbutler.com/api/v1/integrations \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"provider": "vercel", "token": "vcp_…", "teamId": "team_…"}'
201 Created
{
  "provider": "vercel",
  "label": "Vercel",
  "account": "dana",
  "tokenHint": "x9Qf",          // the token itself is never returned
  "teamId": "team_…",
  "projects": [{ "id": "prj_12HKQaOm…", "name": "acme-web" }],
  "projectMap": [],
  "hasToken": true
}
FieldTypeDescription
providerrequiredstringHosting providervercelnetlify
tokenrequiredstringAPI token. Stored encrypted and never returned by any endpointmin length 8 · max length 500
teamIdstringVercel only: team id (team_…) or slug the projects live in. Needed for full-account tokensmax length 120

Webhook payload

Team webhooks gain three fields; everything that was there before is unchanged (links.acknowledge stays too).

{
  "id": "evt_0n3qc1a2b3c4d5e6f7g",
  "type": "monitor.down",
  "timestamp": "2026-10-10T11:02:31.000Z",
  "workspaceId": "ws_…",
  "summaryPlain": "Checkout is failing for users since 14:02 (4 minutes after deploy abc1234 by Vercel).",
  "summaryShort": "Checkout is failing for users since 14:02",
  "actions": {
    "rollback": "https://upbutler.com/act/eyJpIjoi…",
    "fix": "https://upbutler.com/act/eyJpIjoi…",
    "mute": "https://upbutler.com/act/eyJpIjoi…",
    "ack": "https://upbutler.com/ack/eyJpIjoi…"
  },
  "data": { "monitor": { "name": "Checkout", "…": "…" }, "check": { "…": "…" }, "incident": { "id": "inc_…" } }
}

API

What can be done about an incident

curl "https://upbutler.com/api/v1/incidents/inc_…/actions?preview=true" \
  -H "Authorization: Bearer $UPBUTLER_API_KEY"
200 OK
{
  "summaryPlain": "Checkout is failing for users since 14:02 (4 minutes after deploy abc1234 by Vercel).",
  "suspectDeploy": { "_id": "dep_…", "ref": "abc1234", "provider": "vercel", "service": "acme-web", "at": "2026-10-10T10:58:02.000Z" },
  "rollback": {
    "state": "ready",                      // ready | connect | map | none
    "provider": "vercel",
    "target": { "deploymentId": "dpl_…", "label": "9f8e7d6", "title": "Add coupon field", "createdAt": "2026-10-10T08:01:44.000Z" },
    "plan": { "rollback": true },
    "history": []
  },
  "fix": { "available": true, "responder": { "id": "agr_…", "name": "Claude Code" } },
  "mute": { "available": true, "minutes": 60 }
}

preview=true asks the provider which deployment a rollback would restore. rollback.state is connect when the provider is not connected, map when the project is unknown, none when no deploy went out shortly before.

Roll back

curl -X POST https://upbutler.com/api/v1/incidents/inc_…/rollback \
  -H "Authorization: Bearer $UPBUTLER_API_KEY"
201 Created
{
  "_id": "rbk_…",
  "status": "requested",                   // then recovered | still_failing (or failed)
  "provider": "vercel",
  "projectName": "acme-web",
  "fromDeploymentId": "dpl_bad…",
  "toDeploymentId": "dpl_good…",
  "toLabel": "9f8e7d6",
  "deployId": "dep_…",                     // the "Rollback to 9f8e7d6" marker
  "guardId": "grd_…",
  "checkBy": "2026-10-10T11:17:40.000Z",
  "restored": { "deploymentId": "dpl_good…", "label": "9f8e7d6", "commit": "9f8e7d6c5b4a…" }
}

Needs the write scope and the Starter plan or higher (402 plan_limit with details.upgrade otherwise). Returns 409 conflict with a plain message when there is nothing to roll back, a rollback is already under way, or the provider refused.

Rollback by agents

Agents can roll back in two ways, both through the same operation, plan check, deploy marker and guard, with the agent as the actor on the timeline and in the audit log:

  • An agent responder with its incident token, only when its allowed actions include rollback. That action is off for every responder, new and existing, until an admin ticks “Roll back production” on the responder (or lists it in allowedActions). The token works for its own incident only and the agent must claim it first. With the action on, the token can also read GET /incidents/:id/actions.
  • A full API key or an OAuth connection with the write scope, over REST or as the MCP tool incidents_rollback. The tool is not part of the small default toolset (?toolset=core). An agent with a full key should roll back only when a person asked for it.

Connecting a hosting account and changing these settings stay with people: incident tokens cannot call them, and integrations.connect is not an MCP tool.

Settings

GET /api/v1/integrations returns the connected accounts, the settings and whether the plan includes rollback. Change the settings with PATCH /api/v1/alert-actions:

FieldTypeDescription
rollbackbooleanShow "Roll back" on alerts that follow a deploy
fixbooleanShow "Let Claude fix" on alerts
mutebooleanShow "Mute 1h" on alerts
rollbackRequiresSignInbooleanOnly signed-in members can confirm a rollback from any alert link. Links delivered through webhook channels always need a sign-in
fixSetupHintbooleanShow "Set up Claude fixes" on the first alerts while no Fix-with-Claude responder exists

Other endpoints: POST /integrations/:provider/test, PUT /integrations/:provider/projects, DELETE /integrations/:provider. See the API reference.