Monitoring
Monitors
A monitor checks one thing on a schedule and decides whether it is up, degraded or down. This page covers every kind and field, and exactly when a monitor changes state.
Kinds
| Kind | Checks | Needs | Plan |
|---|---|---|---|
http | An HTTP(S) request: status code, body, JSON, headers, latency, TLS expiry | url | All |
tcp | A TCP connection opens within the timeout | host + port | All |
dns | A DNS record resolves (optionally to expected values) | host (+ recordType) | All |
heartbeat | Your job pings us at least every periodSec. See Heartbeats | periodSec | All |
manifest | One JSON endpoint that reports many component statuses. See Manifest | kind + url | All |
script | Your code runs a multi-step API check in an isolated sandbox | code | Pro, Business |
browser | A real-browser (Playwright) check in an isolated sandbox | kind + code | Business |
Kind inference
The input is flat and forgiving: {"name": "API", "url": "https://…"} is a complete monitor. When you leave out kind, it is inferred in this order:
periodSecpresent →heartbeatcodepresent →scripturlpresent →httphost+port→tcphost+recordType→dns
manifest and browser are never inferred, so set kind explicitly. If nothing matches, you get a 422 validation_failed error with a hint listing the options.
Field reference
Fields for POST /monitors, POST /services and PATCH /monitors/:id (where every field is optional). This table is generated from the server's input schema.
| Field | Type | Description |
|---|---|---|
namerequired | string | Human-friendly name, e.g. "Search API"max length 120 |
description | string | null | null or "" clears itmax length 1,000 · nullable |
kind | string | Inferred when omitted. See Kind inference below.httptcpdnsheartbeatmanifestscriptbrowser |
url | string (url) | http and manifest: the URL to check.max length 2,000 |
method | string | http: request method (default GET).GETHEADPOSTPUTPATCHDELETEOPTIONS |
headers | map<string, string> | http and manifest: extra request headers. |
body | string | null | http: request body (ignored for GET and HEAD).max length 100,000 · nullable |
followRedirects | boolean | http: follow redirects (default true). |
expectedStatus | string | http: accepted status codes. Default "200-299". See Expected status.max length 100 |
keyword | string | http: response body must contain this textmax length 500 |
keywordAbsent | string | http: response body must NOT contain this textmax length 500 |
degradedAfterMs | integer | null | Mark degraded when slower than this (null clears)min 50 · max 120,000 · nullable |
sslExpiryDays | integer | null | Mark degraded when the TLS cert expires within N days (null clears)min 1 · max 365 · nullable |
host | string | tcp/dns hostmax length 255 |
port | integer | min 1 · max 65,535 |
recordType | string | AAAAACNAMEMXTXTNS |
expected | string[] | dns: values that must be presentmax 20 items |
periodSec | integer | heartbeat: expected ping intervalmin 30 · max 2,678,400 |
graceSec | integer | heartbeat: extra time before alertingmin 0 · max 86,400 |
code | string | script/browser: test codemax length 50,000 |
viewport | object | |
viewport.widthrequired | integer | browser: viewport width (default 1280).min 320 · max 3,840 |
viewport.heightrequired | integer | browser: viewport height (default 720).min 240 · max 2,160 |
assertions | object[] | max 30 items |
assertions[].sourcerequired | string | statuslatencyheaderbodyjsonssl_days |
assertions[].path | string | max length 300 |
assertions[].oprequired | string | eqneqltltegtgtecontainsnot_containsmatchesexistsnot_exists |
assertions[].value | string | number | boolean | max length 2,000 |
assertions[].severity | string | downdegraded |
intervalSec | integer | Seconds between checks. Default 60, or the plan minimum if that is higher (300 on Free). Must be at least the plan minimum.min 10 · max 86,400 |
timeoutMs | integer | Per-check timeout. Default 15000.min 1,000 · max 120,000 |
failureThreshold | integer | Consecutive failures before DOWN (default 2)min 1 · max 10 |
recoveryThreshold | integer | Consecutive good checks before a failing monitor is up again (default 1).min 1 · max 10 |
channelIds | string[] | Alert channels. Default: channels marked as defaultmax 50 items |
reminderMinutes | integer[] | “Still down” reminders, in minutes after the incident started. Default [30, 120, 480].max 10 items |
ai | boolean | AI incident analysis (default true) |
tags | string[] | Free-form labels. Filter with GET /monitors?tag=…max 20 items |
paused | boolean | Pause (true) or resume (false) checking. |
regions | string[] | Check regions (GET /api/v1/regions). With several, a round only counts as DOWN when a majority of regions agree. Default: main region onlymax 10 items |
HTTP checks
curl -X POST https://upbutler.com/api/v1/monitors \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Search API",
"url": "https://api.example.com/v1/search?q=ping",
"headers": { "x-api-key": "monitoring-only-key" },
"expectedStatus": "200",
"degradedAfterMs": 1500,
"sslExpiryDays": 14,
"assertions": [
{ "source": "json", "path": "$.status", "op": "eq", "value": "ok" },
{ "source": "json", "path": "$.results.length", "op": "gt", "value": 0 },
{ "source": "header", "path": "content-type", "op": "contains", "value": "application/json" },
{ "source": "latency", "op": "lt", "value": 800, "severity": "degraded" }
]
}'Requests go out from the main region eu-central (Helsinki), or from several regions, with the user agent UpButler/2.0 (+https://upbutler.com/bot). Redirects are followed by default. Up to 512 KB of the response body is read. Private and internal addresses are blocked. A result is decided in this order:
- The status code doesn't match
expectedStatus→ down. - An assertion with severity
down(the default) fails → down. - An assertion with severity
degradedfails → degraded. - Latency is above
degradedAfterMs→ degraded. - The TLS certificate expires within
sslExpiryDaysdays → degraded. - Otherwise → up.
Network errors (timeouts, DNS failures, refused or reset connections, TLS errors) are always down, with a readable error message.
Expected status
A comma-separated list of codes, ranges and classes. Spaces are ignored.
| Value | Accepts |
|---|---|
200-299 | Any 2xx (the default) |
200,204 | Exactly 200 or 204 |
2xx,301 | Any 2xx, or 301 |
200-299,301,401 | 2xx, 301 or 401. Useful for endpoints that require auth |
keyword and keywordAbsent
Shortcuts for body assertions. keyword becomes {source: "body", op: "contains"} and keywordAbsent becomes {source: "body", op: "not_contains"}. Setting either one on an update replaces any existing body contains/not_contains assertions.
Assertions
Up to 30 per monitor (HTTP only). Each one reads a value from the response (source), compares it (op) with value, and on failure marks the check down or, with severity: "degraded", degraded.
source | Reads | path |
|---|---|---|
status | HTTP status code | — |
latency | Response time in ms | — |
header | A response header value | Header name (case-insensitive) |
body | The raw response body (up to 512 KB) | — |
json | A value inside the parsed JSON body | JSON path (default $) |
ssl_days | Days until the TLS certificate expires (https only) | — |
op | Passes when |
|---|---|
eq / neq | Value equals / differs, compared as strings (200 equals "200", true equals "true") |
lt lte gt gte | Numeric comparison |
contains | The string includes value, or the array has an element equal to value |
not_contains | The opposite of contains. Also passes when the value is neither a string nor an array |
matches | A JavaScript regular expression matches (an invalid pattern fails) |
exists / not_exists | The value is present (not null or missing) / absent. No value needed |
JSON path syntax
A small, predictable subset. A leading $. is optional, array indexes work with brackets or dots, and .length gives the length of an array or string.
| Path | On {"data":{"items":[{"status":"ok"}]}} |
|---|---|
$.data.items[0].status | "ok" |
data.items.0.status | "ok" |
$.data.items.length | 1 |
$.data.missing | missing (fails everything except not_exists and not_contains) |
If the body isn't valid JSON, every json assertion sees a missing value.
TCP and DNS
{ "name": "Postgres primary", "host": "db.example.com", "port": 5432 }{ "name": "MX records", "host": "example.com", "recordType": "MX", "expected": ["aspmx.l.google.com"] }TCP is up when a connection to host:port opens within timeoutMs. DNS resolves recordType (A by default, or AAAA, CNAME, MX, TXT, NS). It is down when there are no records, or when any expected value doesn't appear in the answers. That comparison is case-insensitive substring matching that ignores trailing dots, and MX answers look like "10 aspmx.l.google.com".
Script and browser checks
script (Pro, Business) and browser (Business) monitors run your code (up to 50,000 characters) inside locked-down, short-lived containers: no host environment, a read-only filesystem, dropped capabilities, and memory and CPU limits. Browser checks take a viewport (default 1280×720). Creating either kind on a plan without it returns 402 plan_limit. A browser run costs about 100 times an HTTP check, so the Business plan includes 20 browser monitors, checked at most every 5 minutes (the default for browser monitors). A faster browser interval returns 422 validation_failed.
Your code is the body of an async function. If it finishes, the check is up. If it throws or a promise rejects, the check is down with that error. Call degraded("reason") to mark it degraded instead. Code can only reach the public internet: requests to private, loopback or link-local addresses (including cloud metadata) are blocked, and redirects are re-checked on every hop.
| Global | Script checks | Browser checks |
|---|---|---|
fetch | Standard fetch, 10 s default timeout per request | Use page instead |
page | — | A real Playwright page (Chromium) |
expect | toBe, toEqual, toContain, toMatch, toBeGreaterThan/LessThan, toBeTruthy/Falsy, toHaveProperty, toHaveLength, and .not | Playwright's expect, including locator matchers like toBeVisible |
assert(cond, msg), sleep(ms), degraded(reason) | Yes | degraded only |
console | Captured into the check's logs (200 lines max) | Captured; a screenshot is attached when the check fails |
const login = await fetch('https://api.example.com/login', {
method: 'POST',
headers: { 'content-type': 'application/json' },
body: JSON.stringify({ user: 'probe', password: 'probe-password' }),
});
expect(login.status).toBe(200);
const { token } = await login.json();
const t0 = Date.now();
const res = await fetch('https://api.example.com/search?q=test', {
headers: { authorization: 'Bearer ' + token },
});
expect((await res.json()).results).toHaveProperty('length');
if (Date.now() - t0 > 1500) degraded('Search slower than 1.5s');Multi-region checks
A single vantage point can't tell “your service is down” from “the network between us and you is down”. Set regions on a monitor and every check runs as a round from each region at once. The results are combined by quorum before the monitor state can change.
| Region id | Location |
|---|---|
eu-central | Helsinki, Finland (main region) |
de-nbg | Nuremberg, Germany |
de-fsn | Falkenstein, Germany |
fr-lbg | Lauterbourg, France |
GET /regions (MCP regions_list) returns the live list with each region's health, the main region and how many regions your plan allows per monitor.
{ "name": "Search API", "url": "https://api.example.com/health", "regions": ["eu-central", "de-nbg", "fr-lbg"] }Quorum
Only regions that actually reported in a round vote. A region whose probe is unhealthy or missed the round is left out. It never counts as a failure, so our own network trouble can't page you.
| Reporting regions say | Round outcome |
|---|---|
| A strict majority down (2 of 3, 2 of 2, 1 of 1) | down |
| A strict majority failing (down or degraded), or exactly half failing (1 of 2) | degraded: a warning, not a page |
| A minority failing (1 of 3) | up. The regional failure stays visible in that region's checks |
A round closes when every expected region has answered, or at timeoutMs + 25 seconds. The round's latency is the median across regions, and its error names the failing regions, such as “Down from 2 of 3 regions — Helsinki: Connection refused; Nuremberg: Connection refused”. Confirmation (failureThreshold) then applies round by round, exactly as for single-region monitors. If none of a monitor's regions are healthy, the main region checks it alone.
- Multi-region works for
http,tcp,dnsandmanifestmonitors. Heartbeats, script and browser checks run from the main region only. - Plans: Free 1 region (main only), Starter up to 2 regions, Pro up to 3 regions, Business every region. Asking for more returns
402 plan_limit. Omittingregions(or sending["eu-central"]) checks from the main region only. GET /monitors/:id/regions?hours=24(up to 168) returns per-region checks, failures, uptime, average and p95 latency, and each region's last result. Every regional check also carriesregion,roundIdandroundOutcome.
When a monitor changes state
Monitors are deliberately hard to fool. One failed check never alerts.
- A new or resumed monitor starts as
pending. Its first good check makes itupsilently. Failures still go through confirmation, and they do alert. - From
up, a failed check (for multi-region monitors, a failed quorum round) increments a counter, and UpButler re-checks after 15 seconds (or the interval, if that is shorter) to confirm. Only whenfailureThresholdconsecutive checks (default 2) are down does the monitor flip todown, opening an incident and alerting. - Degraded results work the same way with their own counter. A
downmonitor that starts responding but isn't healthy moves todegraded. - Recovery needs
recoveryThresholdconsecutive good checks (default 1), again with a quick confirmation re-check when the threshold is higher. - Heartbeats skip confirmation in both directions. An explicit failure ping (
/fail,status=downordegraded) flips the monitor immediately, and a success ping recovers a down heartbeat immediately.
Check intervals by plan
| Plan | Minimum interval | Monitors | Regions | Scripted | Browser |
|---|---|---|---|---|---|
| Free | 5 min | 10 | 1 | — | — |
| Starter | 1 min | 25 | 2 | — | — |
| Pro | 30 s | 75 | 3 | Yes | — |
| Business | 30 s | 300 | All | Yes | 20 (every 5 min+) |
Heartbeat monitors are exempt from the interval minimum, because their schedule comes from periodSec. Values outside intervalSec 10–86,400 are rejected by validation, and values below your plan minimum return 402 plan_limit.
Alerts, reminders and AI
channelIds: which alert channels get this monitor's events. When omitted, all channels marked as default are used. Unknown ids return 422.reminderMinutes: while the monitor stays down with an open incident, amonitor.remindergoes out once each listed duration has passed since the incident started. Up to 10 values from 1 to 10,080 minutes. Pass[]to disable reminders.ai(defaulttrue): attach an AI opening analysis and recovery report to this monitor's incidents and alerts. Each report counts toward your monthly AI quota. See AI reports.
Pause, resume and delete
PATCH /monitors/:id with {"paused": true} stops checks and sets the state to paused. Paused monitors don't count toward component outages. {"paused": false} resets the counters and checks right away. Changing url, host, code, assertions, keyword or expectedStatus also triggers an immediate check.
DELETE /monitors/:id removes the monitor and its check history. Components that followed only this monitor (or read its manifest) switch to the manual source.
Run a check now
POST /monitors/:id/check runs the monitor immediately (from the main region) and returns the full result with evidence. By default it is a dry run: the check is stored as manual, doesn't change state and doesn't count toward uptime. Pass {"apply": true} to feed the result into the state machine, which can open or resolve incidents.
curl -X POST https://upbutler.com/api/v1/monitors/mon_0n3q8kz1m4hx7c2v9rt/check \
-H "Authorization: Bearer $UPBUTLER_API_KEY"{
"check": {
"_id": "chk_0n3q9a10e9n3p5s7hz2",
"monitorId": "mon_0n3q8kz1m4hx7c2v9rt",
"outcome": "down",
"statusCode": 502,
"latencyMs": 312,
"error": "Expected status 200, got 502 Bad Gateway",
"region": "eu-central",
"manual": true,
"evidence": {
"responseHeaders": { "server": "cloudflare", "content-type": "text/html" },
"bodySnippet": "<html><head><title>502 Bad Gateway</title>…",
"resolved": ["104.18.12.33"]
}
},
"state": "up",
"transition": null
}History: GET /monitors/:id returns 24-hour, 30-day and 90-day uptime, daily rollups and the last checks. GET /monitors/:id/checks lists raw results (filter with outcome and before, and add evidence=true for full evidence). GET /checks/:id returns one check in full.