Skip to content
Docs/Status for agents

Status pages

Status for agents

An agent that gets a 503 from your API has one question: what now? Every UpButler status page answers it in one small request: proceed, retry, fallback or pause.

Status pages are written for people. The agent contract is the same information reduced to one verdict. It needs no key, allows any origin, is under 1,000 bytes and can be cached for 15 seconds. It works on every plan and on every host a page has: https://<slug>.upbutler.com, your custom domain, and https://upbutler.com/s/<slug>.

The verdict endpoint

Request
curl https://status.acme.com/status.agent.json
{
  "schema": "upbutler.agent-status/v1",
  "verdict": "proceed",
  "status": "operational",
  "reason": "All systems operational",
  "retryAfterSec": null,
  "eta": null,
  "components": [],
  "incident": null,
  "page": "https://status.acme.com",
  "updatedAt": "2026-10-09T11:42:07.000Z"
}
FieldTypeMeaning
schemastringupbutler.agent-status/v1. New optional fields can appear within v1; a change in meaning gets a new version.
verdictstringproceed, retry, fallback or pause. See what each one asks of you.
statusstringThe worst component status: operational, degraded, partial_outage, major_outage, maintenance or unknown.
reasonstringOne short public sentence: the title of the open incident or maintenance, otherwise a generated summary. At most 140 characters.
retryAfterSecnumber | nullFor retry: wait this long before the next attempt. For fallback and pause: read the verdict again after this long. null with proceed.
etastring | nullWhen the service is expected back (ISO 8601). etaSource says how we know: scheduled is the end of a maintenance window, history is an estimate from how long this page's recent incidents lasted. null when there is nothing to base it on.
caveatstringOnly with proceed. degraded: expect slow or flaky responses, keep your retries on. no_data: the page has no fresh signal, so it cannot vouch for the service.
componentsarrayComponents that are not operational, worst first, each with key, status, verdict and an optional hint written by the page owner. A component that is not listed is operational.
omittednumberPresent when more components are affected than fit in 1 KB. Ask about the one you need with ?component=.
incidentobject | nullid, url (the human-readable incident page) and startedAt.
nextMaintenanceobjectThe next scheduled maintenance window that starts within 7 days: startsAt, endsAt, title and the affected components (keys; absent when the window covers the whole page). It does not change the verdict; use it to plan ahead. Left out when nothing is scheduled, or when an ongoing problem needs the space.
pagestringThe status page for people.
updatedAtstringWhen the underlying state last changed.

One component

Most callers depend on one part of a service. Add ?component=<key> and the verdict, reason, ETA and incident are about that component alone. The document then carries component and always lists that one component, also when it is operational. An unknown key answers 404, so a typo cannot look like good news.

curl "https://status.acme.com/status.agent.json?component=payments-api"
# HEAD is enough when you only need the word:
curl -sI https://status.acme.com/status.agent.json | grep -i x-status-verdict
# x-status-verdict: fallback

Caching

Responses carry Cache-Control: public, max-age=15, stale-while-revalidate=15 and an ETag that only changes when the verdict changes. Send If-None-Match and an unchanged verdict costs a 304 with no body. retryAfterSec is never shorter than the cache lifetime. GET, HEAD and OPTIONS are allowed from any origin; the verdict is also in the X-Status-Verdict response header.

What each verdict asks of you

VerdictDo this
proceedCall the service. If your call just failed, the page has not seen a problem yet: retry with your normal backoff.
retryThe problem is recent or small. Wait retryAfterSec, try again, and back off further on each failure.
fallbackIt has lasted long enough that waiting is the wrong plan. Use your alternative (another provider, a cache, a queue) and read the verdict again after retryAfterSec.
pausePlanned work. Stop jobs that need the service until eta, or read the verdict again after retryAfterSec.

Mapping rules

The verdict is computed, not written by hand, and the rules are fixed. Each component is mapped on its own:

Component statusConditionVerdict
operationalproceed
degradedNo incident, or an incident with impact none or minorproceed, caveat degraded
degradedAn open incident on it with impact major or criticalretry
partial_outageFor less than 15 minutesretry
partial_outageFor 15 minutes or morefallback
major_outageFor less than 5 minutesretry
major_outageFor 5 minutes or morefallback
maintenancepause until the window ends
unknownNo signal, or a pushed status older than its TTLproceed, caveat no_data
  • How long is counted from the start of the incident on that component, or from the moment its status changed when there is no incident.
  • The page verdict is the most restrictive component verdict: fallback, then pause, then retry, then proceed. An open incident with impact major or critical that names no component makes the page retry.
  • Hidden components, private incidents and private pages never appear. An incident or maintenance window that only concerns hidden components is ignored as well, here and in the MCP tools. A private page answers 404.
  • retryAfterSec for retry and fallback grows with how long the trouble has lasted:under 2 min: 15, 2 min to 10 min: 30, 10 min to 30 min: 60, 30 min to 2 h: 120, 2 h or more: 300 seconds. For pause it is the time left in the maintenance window, between 60 and 3,600 seconds, or 300 when no end is known.
  • eta is the scheduled end of the maintenance window, or the incident start plus the median duration of the page's last 20 resolved incidents (at least 3 are needed, and the estimate must still be in the future).

Your own policy

You know your service better than a default table. In the page settings, agentPolicy replaces the verdict for a status, for the whole page (default) or per component key (components), and adds one sentence of guidance that is shown as hint while that component is not operational. A pinned verdict does not change with time; timing fields keep working as above.

Checkout outage ⇒ fallback
curl -X PATCH https://upbutler.com/api/v1/pages/acme \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "settings": {
      "agentPolicy": {
        "default": { "degraded": "retry" },
        "components": {
          "checkout":     { "verdicts": { "major_outage": "fallback", "partial_outage": "fallback" }, "guidance": "Use the EU endpoint." },
          "batch-export": { "verdicts": { "major_outage": "pause" }, "guidance": "Queue exports; they run when we are back." }
        }
      }
    }
  }'

The same settings are in the dashboard under Status pages → your page → Agents & API → What agents are told. Over the API the policy is replaced as a whole on each update; send {"agentPolicy": {}} to go back to the defaults. Guidance is public: write it for your customers' agents.

Discovery: tell agents where to look

RFC 8631 registers the link relation status for exactly this: a link from an API to a resource that describes its health. Add it to your API's error responses and an agent holding a failed response finds the verdict in one hop, with no prior knowledge of your status page.

What the agent sees
HTTP/1.1 503 Service Unavailable
Retry-After: 30
Link: <https://status.acme.com/status.agent.json?component=payments-api>; rel="status"

Send it at least on 429 and 5xx. Point it at the page verdict, or at the component that serves the request. Your status page URL and the copyable header are in the dashboard under Status pages → your page → Agents & API.

const STATUS = '<https://status.acme.com/status.agent.json>; rel="status"';

// After your routes: every error response says where the status is.
app.use((err, req, res, next) => {
  res.set('Link', STATUS).status(err.status ?? 500).json({ error: err.message });
});

The same endpoint is advertised in three more places, so crawlers and agents that start from the status page find it as well:

  • HTML: every status page has <link rel="status" type="application/json" href="/status.agent.json"> in its head.
  • /llms.txt on the status page host lists the verdict, the MCP server, the JSON status and the feeds.
  • /.well-known/status on the status page host serves the same document. This path is a convention of ours, not a registered well-known URI. If you want it on your API host too, redirect it:
# Optional, on your API host: one hop from a hostname to the verdict.
location = /.well-known/status {
    return 302 https://status.acme.com/status.agent.json;
}

A read-only MCP server per status page

Every status page is also an MCP server at /mcp on its own host, for example https://status.acme.com/mcp or https://acme.upbutler.com/mcp. It needs no key and no OAuth, speaks Streamable HTTP, and has 4 tools, so it costs a few hundred tokens in an agent's context. Your customers add one URL.

claude mcp add --transport http acme-status https://status.acme.com/mcp
ToolArgumentsReturns
get_verdictcomponentproceed, retry, fallback or pause, with retryAfterSec, eta and the incident link. Pass component for one component.
get_statusnoneOverall status, every component with its key, open incidents and maintenance.
list_incidentslimitPast and open incidents and maintenance, newest first.
subscribe_webhookurl, components, agentGet signed incident, maintenance and component events at url. We first POST {"type":"subscription.verify","challenge"}; reply with the challenge.

subscribe_webhook is the only tool that writes anything: it creates a webhook subscription, and it is absent when the page has webhook subscriptions turned off. This server is separate from the workspace MCP server at upbutler.com/mcp, which manages monitors and incidents and needs credentials.

SDK helpers

Both SDKs ship a small client for the contract. It caches the verdict for the response's max-age, revalidates with the ETag, times out after 3 seconds and never throws for network trouble: when the status page cannot be reached you get the last known verdict (source: "stale"), or proceed with caveat: "no_data" (source: "unavailable"). A status page must never be the reason your call did not happen.

import { checkStatus, retryDelayMs, statusUrlFromResponse } from '@upbutler/sdk';

async function charge(body: unknown, attempt = 0): Promise<Response> {
  const res = await fetch('https://api.acme.com/v1/charges', { method: 'POST', body: JSON.stringify(body) });
  if (res.status !== 429 && res.status < 500) return res;

  // The failed response says where the status is. Fall back to the page you know.
  const s = await checkStatus(statusUrlFromResponse(res) ?? 'https://status.acme.com', { component: 'payments-api' });
  switch (s.verdict) {
    case 'fallback':
      return chargeWithBackupProvider(body);
    case 'pause':
      throw new PausedUntil(s.eta ?? new Date(Date.now() + s.retryAfterMs));
    default: // retry, or proceed (the page has not noticed yet): back off and try again
      if (attempt >= 4) throw new Error(`payments-api: ${s.reason}`);
      await new Promise((r) => setTimeout(r, retryDelayMs(s, attempt) || 1000 * 2 ** attempt));
      return charge(body, attempt + 1);
  }
}

checkStatus accepts a status page URL, a verdict URL or a bare UpButler page slug. retryDelayMs(status, attempt) turns retryAfterSec into a delay for a retry loop: doubled per attempt, capped at 5 minutes, with jitter.

See also