Monitoring
Error-rate monitors
Synthetic checks tell you a URL answers. This tells you what your users are getting: the share of real requests that end in a 5xx, from the logs your host already has.
How it works
- Create a monitor of kind
error_rate(dashboard: New monitor → Error rate, or the API below). You get a drain URL likehttps://upbutler.com/hooks/drain/ld_…. - Add that URL as a log drain at your hosting provider. The provider posts its request logs to it, in batches.
- UpButler keeps counts only: requests and errors per minute and per normalized path (
/api/orders/8812becomes/api/orders/:id). - Every minute the monitor reads the last window (5 minutes by default). At or above
downAtPct(5%) it is down, at or abovedegradedAtPct(1%) it is degraded. From there it behaves like any monitor: incident, alerts, status page component, AI report, agent responders. The alert names the routes that fail most.
The rules
| Situation | Monitor |
|---|---|
Rate ≥ downAtPct, with at least minRequests in the window | down |
Rate ≥ degradedAtPct, with at least minRequests | degraded |
Fewer than minRequests (default 20) in the window | up, noted as low_traffic. Three failures at 4 a.m. are not an outage. A monitor that is already down stays down while the little traffic that arrives still fails |
| No requests at all (the drain is paused, deleted, or there is no traffic) | up, noted as no_data. A silent drain never opens an incident. The dashboard shows when logs last arrived |
| No requests at all while the monitor is down or degraded | holds its state for 15 minutes after the last logs (holding), because a hard outage often stops the logs too; then no_data |
An error is a response with status statusFrom or higher (default 500, so 4xx never count). On Vercel a function that crashed without a response (statusCode: -1) counts as an error, build logs are skipped, and preview deployments are left out unless you turn on includePreviews. Narrow what counts with hosts, include and exclude path prefixes. Pair this monitor with an HTTP check on the same site: the check catches "nothing answers", this catches "it answers, with errors".
Set up your provider
Vercel (first-class)
- Team Settings → Drains → Add Drain, data type Logs. Drains need a Pro or Enterprise team and are billed by Vercel per GB.
- Choose the projects, the sources
static,lambda,edgeandexternal, and the production environment. On a busy site add a sampling rule: a 10% sample gives the same rate. - Destination Custom Endpoint: paste the drain URL. Format NDJSON or JSON, both work.
- Optional: copy the drain's Signature Verification Secret into the monitor (
errorRate.secret). From then on every delivery must carry a validx-vercel-signature.
Read from each log line: proxy.statusCode (else statusCode), proxy.path (else path), proxy.host, proxy.timestamp, requestId, environment, source. A function that logs several lines for one request is counted once per delivery.
Netlify
- Site → Logs & metrics → Log Drains → Enable a log drain (an Enterprise feature at Netlify).
- Service General HTTP endpoint, log type Traffic logs, and tick Exclude personally identifiable information.
- Full URL: the drain URL. Format NDJSON or JSON. If the monitor has a secret, enter
Bearer <secret>as the Authorization header.
Read: status_code, url, timestamp, request_id.
Cloudflare
- Analytics & Logs → Logpush → Create a Logpush job, destination HTTP, dataset HTTP requests.
- Fields:
EdgeResponseStatus,ClientRequestPath,ClientRequestHost,EdgeStartTimestamp. LeaveClientIPandClientRequestUserAgentout. - Destination: the drain URL, plus
?header_Authorization=Bearer%20<secret>if the monitor has a secret. Bodies arrive gzipped; that is handled.
Workers Trace Events Logpush is read as well (Event.Response.Status, Event.Request.URL, EventTimestampMs; an uncaught exception without a response counts as an error).
Anything else
Post newline-delimited JSON or a JSON array, one object per request. A status is required (status, statusCode or status_code); path or url and timestamp (RFC 3339, or Unix seconds / milliseconds) are optional.
curl -fsS -X POST "https://upbutler.com/hooks/drain/<token>" \
-H "Content-Type: application/x-ndjson" \
--data-binary $'{"status":200,"path":"/api/orders/17","timestamp":"2026-10-10T10:00:00Z"}\n{"status":502,"path":"/api/checkout","timestamp":"2026-10-10T10:00:01Z"}'{ "ok": true, "accepted": 2, "errors": 1, "filtered": 0, "late": 0, "lines": 2, "skipped": 0, "duplicates": 0, "malformed": 0 }Signatures and secrets
The token in the URL is enough to authenticate a delivery. With a secret stored on the monitor, deliveries must also prove it, and anything else is rejected with 403.
| Provider | Checked |
|---|---|
| Vercel | x-vercel-signature: hex HMAC-SHA1 of the raw body with the drain's signature secret |
| Netlify, Cloudflare, generic | Authorization: Bearer <secret> (these providers do not sign log deliveries) |
curl -X PATCH https://upbutler.com/api/v1/monitors/mon_0n3qc1a2b3c4d5e6f7g \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"errorRate": {"secret": "the-drain-signature-secret"}}'The secret is write-only and stored encrypted; the API only says hasSecret. Send null to remove it. An unsigned request that contains no request logs (a provider's endpoint test) is answered with 200 and stores nothing. If your provider asks the endpoint to answer an x-vercel-verify header, set that value as errorRate.verify.
What the alert says
Alerts name the drain, the paths the monitor counts and the rate that tripped it, instead of a URL:
Checkout API is returning errors to users (6.2% of /api/* requests) since 14:02.
Checkout API is DOWN
Target Vercel drain · /api/* error rate 6.2%
Error 5xx rate 6.2% (62 of 1000 requests in the last 5 min; down at 5%). Failing most: /api/checkout/:id (62 of 1000)The scope is the include prefixes (/api/*), or "all paths". The first line is the plain sentence every channel leads with; reminders and recoveries carry the target without a rate. The same target string is in the API as errorRate.target.
In upbutler.yaml
monitors:
- id: shop-errors
name: Shop errors
kind: error_rate
errorRate:
provider: vercel
include: ["/api/"]
exclude: ["/api/health"]
downAtPct: 5
degradedAtPct: 1
minRequests: 20
secret: ${env:VERCEL_DRAIN_SECRET} # optionalupbutler apply creates, updates and adopts error-rate monitors like any other entry; a changed threshold or secret is an update, and the drain URL stays the same. The URL is never written to the file: the apply that creates the monitor returns it once (resources.monitors.<id>.drainUrl, printed by the CLI and repeated under next). After that, POST /monitors/:id/error-rate/rotate-token is the only way to a URL. upbutler export writes thresholds and filters only, never the token, the secret or the verify value. A dry run reports the plan limit as a blocker.
Limits
- Bodies up to 5 MB (24 MB after gzip), at most 50,000 lines each, and 1,200 deliveries a minute per monitor. Above that, sample the drain at the provider.
- Up to 50 paths per minute get their own row; the rest are counted together as "other paths". Failing paths are given a row first.
- Logs older than 2 days are ignored. Providers batch their logs, so expect the rate to trail real time by up to a minute or two.
- A paused monitor, or one on a plan without error-rate monitors, answers
200and stores nothing.
| Plan | Error-rate monitors |
|---|---|
| Free | Not included |
| Starter | Not included |
| Pro | 3 |
| Business | 20 |
API
curl -X POST https://upbutler.com/api/v1/monitors \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name": "Shop errors", "kind": "error_rate", "errorRate": {"provider": "vercel", "exclude": ["/_next/"]}}'{
"_id": "mon_0n3qc1a2b3c4d5e6f7g",
"kind": "error_rate",
"drainUrl": "https://upbutler.com/hooks/drain/ld_Zk3…", // shown once
"errorRate": {
"provider": "vercel",
"windowSec": 300, "downAtPct": 5, "degradedAtPct": 1, "minRequests": 20, "statusFrom": 500,
"exclude": ["/_next/"],
"hasSecret": false,
"drainUrlHint": "https://upbutler.com/hooks/drain/ld_Zk3x1Q…",
"target": "Vercel drain · all paths"
}
}| Field | Type | Description |
|---|---|---|
provider | string | Where the logs come from: vercel (default), netlify, cloudflare or generic. Only changes the setup steps shown; every supported format is readvercelnetlifycloudflaregeneric |
secret | string | null | Write-only.min length 4 · max length 500 · nullable |
verify | string | null | Value to answer in the x-vercel-verify response header, when the provider asks for one. null removes itmax length 200 · nullable |
windowSec | integer | Window the rate is computed over (default 300)min 60 · max 3,600 |
downAtPct | number | Error rate in percent at or above which the monitor is down (default 5)min 0.1 · max 100 |
degradedAtPct | number | null | Error rate in percent at or above which the monitor is degraded (default 1; null = no degraded step)min 0.1 · max 100 · nullable |
minRequests | integer | A window with fewer requests cannot take the monitor down (default 20)min 1 · max 1,000,000 |
statusFrom | integer | Responses with this status or higher are errors (default 500)min 400 · max 599 |
include | string[] | Only count requests whose path starts with one of these, e.g. ["/api/"]max 20 items |
exclude | string[] | Leave out requests whose path starts with one of these, e.g. ["/_next/", "/healthz"]max 20 items |
hosts | string[] | Only count requests to these hostnamesmax 20 items |
includePreviews | boolean | Also count preview deployments (default false) |
| Endpoint | MCP tool | Does |
|---|---|---|
POST /monitors | monitors_create | Create one with "kind": "error_rate"; returns drainUrl once |
PATCH /monitors/:id | monitors_update | Thresholds, filters, secret (errorRate object; fields left out keep their value) |
GET /monitors/:id/error-rate | errorRate_get | The current window, the rate per minute and the failing paths (?minutes=60, up to 2880) |
POST /monitors/:id/error-rate/rotate-token | REST only | New drain URL; the old one stops working at once |
POST /hooks/drain/<token> | The drain itself. GET answers whether the URL is live |
{
"thresholds": { "windowSec": 300, "downAtPct": 5, "degradedAtPct": 1, "minRequests": 20, "statusFrom": 500 },
"lastIngestAt": "2026-10-10T10:04:58.120Z",
"window": { "total": 1204, "errors": 96, "ratePct": 7.97 },
"minutes": 60, "total": 14120, "errors": 131, "ratePct": 0.93,
"series": [{ "at": "2026-10-10T09:05:00.000Z", "total": 231, "errors": 0 }],
"failingPaths": [{ "path": "/api/checkout/:id", "total": 310, "errors": 92, "ratePct": 29.68 }],
"busiestPaths": [{ "path": "/", "total": 5120, "errors": 0, "ratePct": 0 }]
}