Skip to content
Docs/Error-rate monitors

Monitoring

Error-rate monitors

Synthetic checks tell you a URL answers. This tells you what your users are getting: the share of real requests that end in a 5xx, from the logs your host already has.

How it works

  1. Create a monitor of kind error_rate (dashboard: New monitor → Error rate, or the API below). You get a drain URL like https://upbutler.com/hooks/drain/ld_….
  2. Add that URL as a log drain at your hosting provider. The provider posts its request logs to it, in batches.
  3. UpButler keeps counts only: requests and errors per minute and per normalized path (/api/orders/8812 becomes /api/orders/:id).
  4. Every minute the monitor reads the last window (5 minutes by default). At or above downAtPct (5%) it is down, at or above degradedAtPct (1%) it is degraded. From there it behaves like any monitor: incident, alerts, status page component, AI report, agent responders. The alert names the routes that fail most.

The rules

SituationMonitor
Rate ≥ downAtPct, with at least minRequests in the windowdown
Rate ≥ degradedAtPct, with at least minRequestsdegraded
Fewer than minRequests (default 20) in the windowup, noted as low_traffic. Three failures at 4 a.m. are not an outage. A monitor that is already down stays down while the little traffic that arrives still fails
No requests at all (the drain is paused, deleted, or there is no traffic)up, noted as no_data. A silent drain never opens an incident. The dashboard shows when logs last arrived
No requests at all while the monitor is down or degradedholds its state for 15 minutes after the last logs (holding), because a hard outage often stops the logs too; then no_data

An error is a response with status statusFrom or higher (default 500, so 4xx never count). On Vercel a function that crashed without a response (statusCode: -1) counts as an error, build logs are skipped, and preview deployments are left out unless you turn on includePreviews. Narrow what counts with hosts, include and exclude path prefixes. Pair this monitor with an HTTP check on the same site: the check catches "nothing answers", this catches "it answers, with errors".

Set up your provider

Vercel (first-class)

  1. Team Settings → Drains → Add Drain, data type Logs. Drains need a Pro or Enterprise team and are billed by Vercel per GB.
  2. Choose the projects, the sources static, lambda, edge and external, and the production environment. On a busy site add a sampling rule: a 10% sample gives the same rate.
  3. Destination Custom Endpoint: paste the drain URL. Format NDJSON or JSON, both work.
  4. Optional: copy the drain's Signature Verification Secret into the monitor (errorRate.secret). From then on every delivery must carry a valid x-vercel-signature.

Read from each log line: proxy.statusCode (else statusCode), proxy.path (else path), proxy.host, proxy.timestamp, requestId, environment, source. A function that logs several lines for one request is counted once per delivery.

Netlify

  1. Site → Logs & metrics → Log Drains → Enable a log drain (an Enterprise feature at Netlify).
  2. Service General HTTP endpoint, log type Traffic logs, and tick Exclude personally identifiable information.
  3. Full URL: the drain URL. Format NDJSON or JSON. If the monitor has a secret, enter Bearer <secret> as the Authorization header.

Read: status_code, url, timestamp, request_id.

Cloudflare

  1. Analytics & Logs → Logpush → Create a Logpush job, destination HTTP, dataset HTTP requests.
  2. Fields: EdgeResponseStatus, ClientRequestPath, ClientRequestHost, EdgeStartTimestamp. Leave ClientIP and ClientRequestUserAgent out.
  3. Destination: the drain URL, plus ?header_Authorization=Bearer%20<secret> if the monitor has a secret. Bodies arrive gzipped; that is handled.

Workers Trace Events Logpush is read as well (Event.Response.Status, Event.Request.URL, EventTimestampMs; an uncaught exception without a response counts as an error).

Anything else

Post newline-delimited JSON or a JSON array, one object per request. A status is required (status, statusCode or status_code); path or url and timestamp (RFC 3339, or Unix seconds / milliseconds) are optional.

curl -fsS -X POST "https://upbutler.com/hooks/drain/<token>" \
  -H "Content-Type: application/x-ndjson" \
  --data-binary $'{"status":200,"path":"/api/orders/17","timestamp":"2026-10-10T10:00:00Z"}\n{"status":502,"path":"/api/checkout","timestamp":"2026-10-10T10:00:01Z"}'
200
{ "ok": true, "accepted": 2, "errors": 1, "filtered": 0, "late": 0, "lines": 2, "skipped": 0, "duplicates": 0, "malformed": 0 }

Signatures and secrets

The token in the URL is enough to authenticate a delivery. With a secret stored on the monitor, deliveries must also prove it, and anything else is rejected with 403.

ProviderChecked
Vercelx-vercel-signature: hex HMAC-SHA1 of the raw body with the drain's signature secret
Netlify, Cloudflare, genericAuthorization: Bearer <secret> (these providers do not sign log deliveries)
curl -X PATCH https://upbutler.com/api/v1/monitors/mon_0n3qc1a2b3c4d5e6f7g \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"errorRate": {"secret": "the-drain-signature-secret"}}'

The secret is write-only and stored encrypted; the API only says hasSecret. Send null to remove it. An unsigned request that contains no request logs (a provider's endpoint test) is answered with 200 and stores nothing. If your provider asks the endpoint to answer an x-vercel-verify header, set that value as errorRate.verify.

What the alert says

Alerts name the drain, the paths the monitor counts and the rate that tripped it, instead of a URL:

Checkout API is returning errors to users (6.2% of /api/* requests) since 14:02.

Checkout API is DOWN
Target   Vercel drain · /api/* error rate 6.2%
Error    5xx rate 6.2% (62 of 1000 requests in the last 5 min; down at 5%). Failing most: /api/checkout/:id (62 of 1000)

The scope is the include prefixes (/api/*), or "all paths". The first line is the plain sentence every channel leads with; reminders and recoveries carry the target without a rate. The same target string is in the API as errorRate.target.

In upbutler.yaml

upbutler.yaml
monitors:
  - id: shop-errors
    name: Shop errors
    kind: error_rate
    errorRate:
      provider: vercel
      include: ["/api/"]
      exclude: ["/api/health"]
      downAtPct: 5
      degradedAtPct: 1
      minRequests: 20
      secret: ${env:VERCEL_DRAIN_SECRET}   # optional

upbutler apply creates, updates and adopts error-rate monitors like any other entry; a changed threshold or secret is an update, and the drain URL stays the same. The URL is never written to the file: the apply that creates the monitor returns it once (resources.monitors.<id>.drainUrl, printed by the CLI and repeated under next). After that, POST /monitors/:id/error-rate/rotate-token is the only way to a URL. upbutler export writes thresholds and filters only, never the token, the secret or the verify value. A dry run reports the plan limit as a blocker.

Limits

  • Bodies up to 5 MB (24 MB after gzip), at most 50,000 lines each, and 1,200 deliveries a minute per monitor. Above that, sample the drain at the provider.
  • Up to 50 paths per minute get their own row; the rest are counted together as "other paths". Failing paths are given a row first.
  • Logs older than 2 days are ignored. Providers batch their logs, so expect the rate to trail real time by up to a minute or two.
  • A paused monitor, or one on a plan without error-rate monitors, answers 200 and stores nothing.
PlanError-rate monitors
FreeNot included
StarterNot included
Pro3
Business20

API

curl -X POST https://upbutler.com/api/v1/monitors \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"name": "Shop errors", "kind": "error_rate", "errorRate": {"provider": "vercel", "exclude": ["/_next/"]}}'
201 Created
{
  "_id": "mon_0n3qc1a2b3c4d5e6f7g",
  "kind": "error_rate",
  "drainUrl": "https://upbutler.com/hooks/drain/ld_Zk3…",   // shown once
  "errorRate": {
    "provider": "vercel",
    "windowSec": 300, "downAtPct": 5, "degradedAtPct": 1, "minRequests": 20, "statusFrom": 500,
    "exclude": ["/_next/"],
    "hasSecret": false,
    "drainUrlHint": "https://upbutler.com/hooks/drain/ld_Zk3x1Q…",
    "target": "Vercel drain · all paths"
  }
}
FieldTypeDescription
providerstringWhere the logs come from: vercel (default), netlify, cloudflare or generic. Only changes the setup steps shown; every supported format is readvercelnetlifycloudflaregeneric
secretstring | nullWrite-only.min length 4 · max length 500 · nullable
verifystring | nullValue to answer in the x-vercel-verify response header, when the provider asks for one. null removes itmax length 200 · nullable
windowSecintegerWindow the rate is computed over (default 300)min 60 · max 3,600
downAtPctnumberError rate in percent at or above which the monitor is down (default 5)min 0.1 · max 100
degradedAtPctnumber | nullError rate in percent at or above which the monitor is degraded (default 1; null = no degraded step)min 0.1 · max 100 · nullable
minRequestsintegerA window with fewer requests cannot take the monitor down (default 20)min 1 · max 1,000,000
statusFromintegerResponses with this status or higher are errors (default 500)min 400 · max 599
includestring[]Only count requests whose path starts with one of these, e.g. ["/api/"]max 20 items
excludestring[]Leave out requests whose path starts with one of these, e.g. ["/_next/", "/healthz"]max 20 items
hostsstring[]Only count requests to these hostnamesmax 20 items
includePreviewsbooleanAlso count preview deployments (default false)
EndpointMCP toolDoes
POST /monitorsmonitors_createCreate one with "kind": "error_rate"; returns drainUrl once
PATCH /monitors/:idmonitors_updateThresholds, filters, secret (errorRate object; fields left out keep their value)
GET /monitors/:id/error-rateerrorRate_getThe current window, the rate per minute and the failing paths (?minutes=60, up to 2880)
POST /monitors/:id/error-rate/rotate-tokenREST onlyNew drain URL; the old one stops working at once
POST /hooks/drain/<token>The drain itself. GET answers whether the URL is live
GET /monitors/:id/error-rate
{
  "thresholds": { "windowSec": 300, "downAtPct": 5, "degradedAtPct": 1, "minRequests": 20, "statusFrom": 500 },
  "lastIngestAt": "2026-10-10T10:04:58.120Z",
  "window": { "total": 1204, "errors": 96, "ratePct": 7.97 },
  "minutes": 60, "total": 14120, "errors": 131, "ratePct": 0.93,
  "series": [{ "at": "2026-10-10T09:05:00.000Z", "total": 231, "errors": 0 }],
  "failingPaths": [{ "path": "/api/checkout/:id", "total": 310, "errors": 92, "ratePct": 29.68 }],
  "busiestPaths": [{ "path": "/", "total": 5120, "errors": 0, "ratePct": 0 }]
}