Skip to content
Docs/Introduction

Get started

Introduction

UpButler is uptime monitoring, status pages and incident communication in one place. It is built API-first, so an AI agent can set up and run everything a human can do from the dashboard.

You point UpButler at your services. It checks them, opens incidents when they fail, writes the status page update, alerts your team and notifies your customers. When the service recovers, it does all of that again in reverse. There is a dashboard, a REST API, an MCP server and signed webhooks. All four run the same operations, so anything you can click, an agent can call.

One call: monitor a URL and publish it on a status page
curl -X POST https://upbutler.com/api/v1/services \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Search API",
    "url": "https://api.example.com/health",
    "statusPage": "acme",
    "group": "APIs"
  }'

Core concepts

Workspace

A workspace owns everything else: monitors, status pages, channels, API keys and the plan. A person can belong to several workspaces with the role owner, admin or member. API keys (ub_live_…) belong to one workspace and carry the read and/or write scope.

An agent can create an unclaimed workspace with no human involved (POST /api/v1/agent/bootstrap). It gets a claim link to pass to its human. Unclaimed workspaces are deleted after 7 days, and email features only unlock after a human claims the workspace.

Monitors and checks

A monitor checks one thing on a schedule. There are seven kinds: http, tcp, dns, heartbeat (your job pings us), manifest (one JSON endpoint that reports many component statuses), script and browser. Each run stores a check with outcome up, degraded or down, plus latency and evidence (status code, headers, body snippet, TLS and DNS details).

A monitor's state is pending, up, degraded, down or paused. The state only changes after failureThreshold consecutive bad checks (2 by default). UpButler re-checks sooner to confirm a suspected failure, so a single network blip never pages anyone.

Status pages, groups and components

A status page has ordered groups (such as “APIs” or “Dashboard”), and each group holds components, the rows your customers see. A component's status is one of operational, degraded, partial_outage, major_outage, maintenance or unknown. It gets that status from its source:

SourceStatus comes from
monitorOne or more monitors, combined with worst or majority
pushYour system POSTs a status to the component's push URL. See Push & manifest
manifestA key inside a manifest monitor's JSON
manualNothing automatic. You set it (default)

Incidents, maintenance windows and manual changes set an override. An override wins over the automatic status until it is cleared.

Incidents and maintenance

Incidents open automatically when a monitor goes down or a pushed component reports an outage. Failures on the same status page within 30 minutes join one incident instead of opening five. You or your agents can also open incidents by hand. Every incident has a timeline of public or internal updates, and optional AI reports: an opening analysis with likely causes and suggested actions, and a recovery summary. Maintenance is an incident with kind: "maintenance" that starts and completes on its own schedule.

Events and deliveries

Every state change is stored as an event (evt_…), such as monitor.down, incident.created or component.status_changed. Each event fans out into deliveries, which are queued and retried:

  • to your team's alert channels (email, Slack, Discord, Telegram, webhook), and
  • to a status page's subscribers (email with double opt-in, webhooks, Slack, Discord).

Webhooks follow the Standard Webhooks spec. Agents without a public endpoint can poll GET /api/v1/events?after=… instead.

Vocabulary at a glance

ThingValues
Monitor statepending up degraded down paused
Check outcomeup degraded down
Component statusoperational degraded partial_outage major_outage maintenance unknown
Incident statusinvestigating identified monitoring resolved
Maintenance statusscheduled in_progress completed
Impactnone minor major critical maintenance
ID prefixesws_ mon_ chk_ pg_ cmp_ inc_ mnt_ evt_ ch_ sub_ key_