Status pages
Status for agents
An agent that gets a 503 from your API has one question: what now? Every UpButler status page answers it in one small request: proceed, retry, fallback or pause.
Status pages are written for people. The agent contract is the same information reduced to one verdict. It needs no key, allows any origin, is under 1,000 bytes and can be cached for 15 seconds. It works on every plan and on every host a page has: https://<slug>.upbutler.com, your custom domain, and https://upbutler.com/s/<slug>.
The verdict endpoint
curl https://status.acme.com/status.agent.json{
"schema": "upbutler.agent-status/v1",
"verdict": "proceed",
"status": "operational",
"reason": "All systems operational",
"retryAfterSec": null,
"eta": null,
"components": [],
"incident": null,
"page": "https://status.acme.com",
"updatedAt": "2026-10-09T11:42:07.000Z"
}{
"schema": "upbutler.agent-status/v1",
"verdict": "fallback",
"status": "major_outage",
"reason": "Card payments are failing",
"retryAfterSec": 60,
"eta": "2026-10-09T12:40:00.000Z",
"etaSource": "history",
"components": [
{ "key": "payments-api", "status": "major_outage", "verdict": "fallback", "hint": "Use the EU endpoint." },
{ "key": "dashboard", "status": "degraded", "verdict": "proceed" }
],
"incident": {
"id": "inc_1k4h5i0bqr4gb8nbezp",
"url": "https://status.acme.com/incidents/inc_1k4h5i0bqr4gb8nbezp",
"startedAt": "2026-10-09T12:04:11.000Z"
},
"page": "https://status.acme.com",
"updatedAt": "2026-10-09T12:15:30.000Z"
}| Field | Type | Meaning |
|---|---|---|
schema | string | upbutler.agent-status/v1. New optional fields can appear within v1; a change in meaning gets a new version. |
verdict | string | proceed, retry, fallback or pause. See what each one asks of you. |
status | string | The worst component status: operational, degraded, partial_outage, major_outage, maintenance or unknown. |
reason | string | One short public sentence: the title of the open incident or maintenance, otherwise a generated summary. At most 140 characters. |
retryAfterSec | number | null | For retry: wait this long before the next attempt. For fallback and pause: read the verdict again after this long. null with proceed. |
eta | string | null | When the service is expected back (ISO 8601). etaSource says how we know: scheduled is the end of a maintenance window, history is an estimate from how long this page's recent incidents lasted. null when there is nothing to base it on. |
caveat | string | Only with proceed. degraded: expect slow or flaky responses, keep your retries on. no_data: the page has no fresh signal, so it cannot vouch for the service. |
components | array | Components that are not operational, worst first, each with key, status, verdict and an optional hint written by the page owner. A component that is not listed is operational. |
omitted | number | Present when more components are affected than fit in 1 KB. Ask about the one you need with ?component=. |
incident | object | null | id, url (the human-readable incident page) and startedAt. |
nextMaintenance | object | The next scheduled maintenance window that starts within 7 days: startsAt, endsAt, title and the affected components (keys; absent when the window covers the whole page). It does not change the verdict; use it to plan ahead. Left out when nothing is scheduled, or when an ongoing problem needs the space. |
page | string | The status page for people. |
updatedAt | string | When the underlying state last changed. |
One component
Most callers depend on one part of a service. Add ?component=<key> and the verdict, reason, ETA and incident are about that component alone. The document then carries component and always lists that one component, also when it is operational. An unknown key answers 404, so a typo cannot look like good news.
curl "https://status.acme.com/status.agent.json?component=payments-api"
# HEAD is enough when you only need the word:
curl -sI https://status.acme.com/status.agent.json | grep -i x-status-verdict
# x-status-verdict: fallbackCaching
Responses carry Cache-Control: public, max-age=15, stale-while-revalidate=15 and an ETag that only changes when the verdict changes. Send If-None-Match and an unchanged verdict costs a 304 with no body. retryAfterSec is never shorter than the cache lifetime. GET, HEAD and OPTIONS are allowed from any origin; the verdict is also in the X-Status-Verdict response header.
What each verdict asks of you
| Verdict | Do this |
|---|---|
proceed | Call the service. If your call just failed, the page has not seen a problem yet: retry with your normal backoff. |
retry | The problem is recent or small. Wait retryAfterSec, try again, and back off further on each failure. |
fallback | It has lasted long enough that waiting is the wrong plan. Use your alternative (another provider, a cache, a queue) and read the verdict again after retryAfterSec. |
pause | Planned work. Stop jobs that need the service until eta, or read the verdict again after retryAfterSec. |
Mapping rules
The verdict is computed, not written by hand, and the rules are fixed. Each component is mapped on its own:
| Component status | Condition | Verdict |
|---|---|---|
operational | proceed | |
degraded | No incident, or an incident with impact none or minor | proceed, caveat degraded |
degraded | An open incident on it with impact major or critical | retry |
partial_outage | For less than 15 minutes | retry |
partial_outage | For 15 minutes or more | fallback |
major_outage | For less than 5 minutes | retry |
major_outage | For 5 minutes or more | fallback |
maintenance | pause until the window ends | |
unknown | No signal, or a pushed status older than its TTL | proceed, caveat no_data |
- How long is counted from the start of the incident on that component, or from the moment its status changed when there is no incident.
- The page verdict is the most restrictive component verdict:
fallback, thenpause, thenretry, thenproceed. An open incident with impactmajororcriticalthat names no component makes the pageretry. - Hidden components, private incidents and private pages never appear. An incident or maintenance window that only concerns hidden components is ignored as well, here and in the MCP tools. A private page answers
404. retryAfterSecforretryandfallbackgrows with how long the trouble has lasted:under 2 min:15, 2 min to 10 min:30, 10 min to 30 min:60, 30 min to 2 h:120, 2 h or more:300seconds. Forpauseit is the time left in the maintenance window, between 60 and 3,600 seconds, or 300 when no end is known.etais the scheduled end of the maintenance window, or the incident start plus the median duration of the page's last 20 resolved incidents (at least 3 are needed, and the estimate must still be in the future).
Your own policy
You know your service better than a default table. In the page settings, agentPolicy replaces the verdict for a status, for the whole page (default) or per component key (components), and adds one sentence of guidance that is shown as hint while that component is not operational. A pinned verdict does not change with time; timing fields keep working as above.
curl -X PATCH https://upbutler.com/api/v1/pages/acme \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"settings": {
"agentPolicy": {
"default": { "degraded": "retry" },
"components": {
"checkout": { "verdicts": { "major_outage": "fallback", "partial_outage": "fallback" }, "guidance": "Use the EU endpoint." },
"batch-export": { "verdicts": { "major_outage": "pause" }, "guidance": "Queue exports; they run when we are back." }
}
}
}
}'The same settings are in the dashboard under Status pages → your page → Agents & API → What agents are told. Over the API the policy is replaced as a whole on each update; send {"agentPolicy": {}} to go back to the defaults. Guidance is public: write it for your customers' agents.
Discovery: tell agents where to look
RFC 8631 registers the link relation status for exactly this: a link from an API to a resource that describes its health. Add it to your API's error responses and an agent holding a failed response finds the verdict in one hop, with no prior knowledge of your status page.
HTTP/1.1 503 Service Unavailable
Retry-After: 30
Link: <https://status.acme.com/status.agent.json?component=payments-api>; rel="status"Send it at least on 429 and 5xx. Point it at the page verdict, or at the component that serves the request. Your status page URL and the copyable header are in the dashboard under Status pages → your page → Agents & API.
const STATUS = '<https://status.acme.com/status.agent.json>; rel="status"';
// After your routes: every error response says where the status is.
app.use((err, req, res, next) => {
res.set('Link', STATUS).status(err.status ?? 500).json({ error: err.message });
});const STATUS = '<https://status.acme.com/status.agent.json>; rel="status"';
app.use(async (c, next) => {
await next();
if (c.res.status === 429 || c.res.status >= 500) c.res.headers.append('Link', STATUS);
});
app.onError((err, c) => c.json({ error: err.message }, 500, { Link: STATUS }));// middleware.ts: add the header to every API response (it is small and harmless on a 200).
import { NextResponse } from 'next/server';
export function middleware() {
const res = NextResponse.next();
res.headers.append('Link', '<https://status.acme.com/status.agent.json>; rel="status"');
return res;
}
export const config = { matcher: '/api/:path*' };const STATUS = '<https://status.acme.com/status.agent.json>; rel="status"';
Bun.serve({
async fetch(req) {
const res = await handle(req).catch(() => new Response('Internal error', { status: 500 }));
if (res.status === 429 || res.status >= 500) res.headers.append('Link', STATUS);
return res;
},
});STATUS = '<https://status.acme.com/status.agent.json>; rel="status"'
@app.middleware("http")
async def status_link(request, call_next):
try:
response = await call_next(request)
except Exception:
return JSONResponse({"error": "internal"}, status_code=500, headers={"Link": STATUS})
if response.status_code == 429 or response.status_code >= 500:
response.headers.append("Link", STATUS)
return response# "always" makes nginx add the header to error responses too, including the ones it
# generates itself (502, 504) when your app is the thing that is down.
location /api/ {
add_header Link '<https://status.acme.com/status.agent.json>; rel="status"' always;
proxy_pass http://app;
}The same endpoint is advertised in three more places, so crawlers and agents that start from the status page find it as well:
- HTML: every status page has
<link rel="status" type="application/json" href="/status.agent.json">in its head. /llms.txton the status page host lists the verdict, the MCP server, the JSON status and the feeds./.well-known/statuson the status page host serves the same document. This path is a convention of ours, not a registered well-known URI. If you want it on your API host too, redirect it:
# Optional, on your API host: one hop from a hostname to the verdict.
location = /.well-known/status {
return 302 https://status.acme.com/status.agent.json;
}A read-only MCP server per status page
Every status page is also an MCP server at /mcp on its own host, for example https://status.acme.com/mcp or https://acme.upbutler.com/mcp. It needs no key and no OAuth, speaks Streamable HTTP, and has 4 tools, so it costs a few hundred tokens in an agent's context. Your customers add one URL.
claude mcp add --transport http acme-status https://status.acme.com/mcp{
"mcpServers": {
"acme-status": { "type": "http", "url": "https://status.acme.com/mcp" }
}
}| Tool | Arguments | Returns |
|---|---|---|
get_verdict | component | proceed, retry, fallback or pause, with retryAfterSec, eta and the incident link. Pass component for one component. |
get_status | none | Overall status, every component with its key, open incidents and maintenance. |
list_incidents | limit | Past and open incidents and maintenance, newest first. |
subscribe_webhook | url, components, agent | Get signed incident, maintenance and component events at url. We first POST {"type":"subscription.verify","challenge"}; reply with the challenge. |
subscribe_webhook is the only tool that writes anything: it creates a webhook subscription, and it is absent when the page has webhook subscriptions turned off. This server is separate from the workspace MCP server at upbutler.com/mcp, which manages monitors and incidents and needs credentials.
SDK helpers
Both SDKs ship a small client for the contract. It caches the verdict for the response's max-age, revalidates with the ETag, times out after 3 seconds and never throws for network trouble: when the status page cannot be reached you get the last known verdict (source: "stale"), or proceed with caveat: "no_data" (source: "unavailable"). A status page must never be the reason your call did not happen.
import { checkStatus, retryDelayMs, statusUrlFromResponse } from '@upbutler/sdk';
async function charge(body: unknown, attempt = 0): Promise<Response> {
const res = await fetch('https://api.acme.com/v1/charges', { method: 'POST', body: JSON.stringify(body) });
if (res.status !== 429 && res.status < 500) return res;
// The failed response says where the status is. Fall back to the page you know.
const s = await checkStatus(statusUrlFromResponse(res) ?? 'https://status.acme.com', { component: 'payments-api' });
switch (s.verdict) {
case 'fallback':
return chargeWithBackupProvider(body);
case 'pause':
throw new PausedUntil(s.eta ?? new Date(Date.now() + s.retryAfterMs));
default: // retry, or proceed (the page has not noticed yet): back off and try again
if (attempt >= 4) throw new Error(`payments-api: ${s.reason}`);
await new Promise((r) => setTimeout(r, retryDelayMs(s, attempt) || 1000 * 2 ** attempt));
return charge(body, attempt + 1);
}
}import { shouldProceed } from '@upbutler/sdk';
// Before starting a long job that needs the service:
if (!(await shouldProceed('https://status.acme.com', { component: 'payments-api' }))) return;from upbutler import check_status, retry_delay, should_proceed, status_url_from_headers
s = check_status("https://status.acme.com", component="payments-api")
if s["verdict"] == "fallback":
use_backup_provider()
elif s["verdict"] == "pause":
sleep_until(s["eta"] or time.time() + s["retryAfterSec"])
elif s["verdict"] == "retry":
time.sleep(retry_delay(s, attempt))
# From a failed response (requests / httpx):
url = status_url_from_headers(response.headers, str(response.url))
# As a gate:
if not should_proceed("https://status.acme.com"):
returnupbutler verdict https://status.acme.com -c payments-api
# fallback: Card payments are failing
# check again in 60s, expected back 2026-10-09T12:40:00.000Z
# payments-api: major_outage → fallback (Use the EU endpoint.)
# https://status.acme.com/incidents/inc_1k4h5i0bqr4gb8nbezp
# The exit code is the verdict: 0 proceed, 10 retry, 11 fallback, 12 pause.
upbutler verdict https://status.acme.com && ./run-batch.shcheckStatus accepts a status page URL, a verdict URL or a bare UpButler page slug. retryDelayMs(status, attempt) turns retryAfterSec into a delay for a retry loop: doubled per attempt, capped at 5 minutes, with jitter.