Skip to content
Docs/MCP, LLM & credit monitors

Monitoring

AI-stack monitors

Three monitor kinds for what agents depend on. They fail in ways an HTTP 200 check doesn't see: a tool loses a parameter, a stream starts and dies, a model is retired, the credits run out. They use the same confirmation, incidents, status pages, alerts and AI reports as every other monitor.

KindChecksDefault intervalRegions
mcpHandshake, tool list, optional probe call, schema drift5 min (plan minimum applies)Multi-region
llmOne tiny streamed completion: first-token time, stream completes, error class10 min, never under 5Multi-region (each region is another completion)
balanceRemaining credits and projected run-out date1 hour, never under 5 minMain region only

Create them like any monitor, with POST /monitors, the MCP tool monitors_create, or in the dashboard under Monitors → New monitor. The kind is inferred when you send an mcp, llm or balance object.

MCP server

curl -X POST https://upbutler.com/api/v1/monitors \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "Acme MCP",
    "kind": "mcp",
    "url": "https://mcp.acme.com/mcp",
    "headers": { "Authorization": "Bearer <token for the MCP server>" },
    "mcp": {
      "probe": { "tool": "ping", "args": {}, "expectContains": "ok" },
      "driftStatus": "degraded",
      "driftHoldHours": 24
    }
  }'

Each check speaks MCP over Streamable HTTP, and falls back to the older HTTP+SSE transport when the server only offers that:

  1. initialize, then notifications/initialized.
  2. tools/list, following nextCursor until the list is complete. With listPrompts / listResources, also prompts/list and resources/list when the server advertises them.
  3. If mcp.probe is set: one tools/call. The check fails when the tool is missing, returns isError, or its text doesn't satisfy expectContains / expectRegex.

The time of every step is stored with each check (initMs, listMs, callMs), and GET /monitors/:id/metrics returns their p50 and p95.

What a failure means

evidence.mcp.failureMeaningStatus
authHTTP 401 or 403 at any step. The server is reachable; the credential is wrong or expired.down
unreachableConnection failed, DNS failed, or HTTP 5xx.down
timeoutThe check ran out of time; the error names the step.down
protocolThe endpoint answered, but not with valid MCP: a JSON-RPC error, a missing tools array, HTML.down
probeThe probe tool is gone, errored, or failed its assertion.down
driftThe server works, but its tool schema changed in a breaking way.driftStatus

Schema drift

The first successful check records the tool list as the baseline. Later checks compare what the server serves with that baseline, and every difference is classified:

SeverityChangesWhat happens
breakingTool removed · parameter removed · new required parameter · optional parameter became required · parameter type changed · enum value removed · prompt removedEvent, and the monitor goes to driftStatus (default degraded) with the diff as its error, which alerts your channels
non_breakingTool added · optional parameter added · parameter became optional · enum value added · constraint or output schema changed · resources changedEvent only. The baseline moves.
descriptionOnly the wording of a tool or parameter description changedEvent only. The baseline moves.

Description changes are listed on their own because a quietly edited tool description is how tool poisoning reaches an agent. Subscribe a webhook to monitor.schema_changed if you want to review each one.

Nested object parameters are compared by path (filters.date). A new required field inside an optional object is not breaking, because callers that don't send the object are unaffected. Key order and a server version bump are never drift.

After a breaking change

The monitor stays in driftStatus until one of three things happens:

  • You accept it. Click Accept as new baseline on the monitor page, or call monitors_schema_accept. Do this once your agents work with the new schema.
  • The server rolls it back. The tool list matches the baseline again and the monitor recovers by itself.
  • driftHoldHours pass (default 24, up to 720). The change is accepted automatically, so a monitor is never degraded forever.
curl -X POST https://upbutler.com/api/v1/monitors/mon_…/schema/accept \
  -H "Authorization: Bearer $UPBUTLER_API_KEY"

driftStatus: "down" opens an incident like any outage, with an AI report that names the tools and parameters. driftStatus: "none" records breaking changes without changing the monitor.

GET /monitors/:id/schema
{
  "baseline": { "hash": "8bc366949f88e4ce…", "tools": [ { "name": "search", "params": { "query": { "type": "string", "required": true } } } ] },
  "pending": {
    "at": "2026-10-09T12:04:11.000Z",
    "until": "2026-10-10T12:04:11.000Z",
    "summary": "tool `search` lost required param `query`"
  },
  "changes": [
    {
      "severity": "breaking",
      "status": "pending",
      "changes": [
        { "severity": "breaking", "type": "param_removed", "tool": "search", "param": "query",
          "text": "tool `search` lost required param `query`" }
      ]
    }
  ]
}
Event: monitor.schema_changed
{
  "type": "monitor.schema_changed",
  "data": {
    "monitor": { "_id": "mon_…", "name": "Acme MCP", "kind": "mcp", "target": "https://mcp.acme.com/mcp" },
    "schemaChange": {
      "severity": "breaking",
      "pending": true,
      "summary": "tool `search` lost required param `query`",
      "changes": [ { "severity": "breaking", "type": "param_removed", "tool": "search", "param": "query", "text": "…" } ]
    }
  }
}

The event is not in the default set channels receive, so it doesn't add noise: breaking changes already alert through monitor.degraded or monitor.down. Add monitor.schema_changed to a channel's events to get all three severities as a tool changelog.

LLM endpoint

curl -X POST https://upbutler.com/api/v1/monitors \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "OpenRouter gpt-5-mini",
    "kind": "llm",
    "url": "https://openrouter.ai/api/v1",
    "intervalSec": 600,
    "llm": {
      "model": "openai/gpt-5-mini",
      "apiKey": "sk-or-…",
      "expectContains": "pong",
      "degradedTtftMs": 3000
    }
  }'

url is the API base URL or the full endpoint. llm.api selects the wire format:

  • openai (default): any OpenAI-compatible /chat/completions. That covers OpenAI, OpenRouter, Azure OpenAI, vLLM, LiteLLM, Ollama, Together, Groq and most gateways.
  • anthropic: the native Messages API (/v1/messages, x-api-key).
Native Anthropic
{
  "name": "Claude Haiku",
  "kind": "llm",
  "url": "https://api.anthropic.com",
  "llm": { "api": "anthropic", "model": "claude-haiku-4-5", "apiKey": "sk-ant-…" }
}

Each check sends the prompt (default Reply with the single word: pong) with stream: true and a small output cap, and records:

MetricMeaning
ttfbMsUntil the response headers arrived
ttftMsTime to first token: until the first content (or reasoning) delta arrived
totalMsUntil the stream finished
chunksContent deltas received

The stream must end with a finish signal (finish_reason or [DONE]; message_stop for Anthropic). A stream that just stops is reported as truncated. Optional rules: expectContains (not case-sensitive), expectRegex, degradedTtftMs, degradedTotalMs.

Error classes

The class is in evidence.llm.errorClass and in the check's error text. The response body is read before the status code, because providers disagree on codes: OpenAI reports an empty wallet as 429 insufficient_quota, Anthropic as 400.

ClassTriggerStatus
rate_limitedHTTP 429 (your key is throttled; Retry-After is recorded)degraded
overloadedHTTP 529 or 503, or an overloaded_error inside the streamdown
model_unavailablemodel_not_found, "deprecated", "decommissioned", HTTP 404 or 410down
billingHTTP 402, insufficient_quota, "credit balance is too low"down
authHTTP 401 or 403down
truncatedThe stream ended without a finish signaldown
timeout · network · provider_error · bad_request · assertionNo answer in time · connection failed · other 5xx · other 4xx · the response failed your ruledown

Like every monitor, a failure only flips the state after failureThreshold consecutive checks, so one 529 doesn't page anyone.

Deprecation signals

Response headers named Sunset, Deprecation or anything containing "deprecat" are recorded with the check. When a Sunset date is within deprecationDays (default 30), the monitor goes degraded with "Model will be retired in N days". When the model is already gone, the check fails as model_unavailable. If the provider answers with a different model than you asked for, both names are shown in the check.

Cost

The completion runs on your key, so your provider bills it. Two limits hold on every plan: the interval is never shorter than 5 minutes, and maxTokens is at most 64 (default 16).

With the default prompt a check is about 25 input tokens and at most 16 output tokens:

IntervalChecks / monthInput tokensOutput tokens (max)At $0.25 / $2 per 1MAt $3 / $15 per 1M
10 min (default)4,320108k69kabout $0.17about $1.36
5 min8,640216k138kabout $0.33about $2.72
1 hour72018k12kabout $0.03about $0.23

Multiply by the number of regions. A failure is re-checked once after 15 seconds, which adds a few completions during an outage. "Test now" sends one completion. Reasoning models may spend the whole output cap on reasoning and return no text; raise maxTokens or drop the response rule for those.

Credit balance

# OpenRouter account credits (needs a management key)
curl -X POST https://upbutler.com/api/v1/monitors \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "OpenRouter credits",
    "kind": "balance",
    "balance": { "preset": "openrouter_credits", "apiKey": "sk-or-…", "minDays": 7, "minAmount": 10 }
  }'

# Any JSON endpoint
curl -X POST https://upbutler.com/api/v1/monitors \
  -H "Authorization: Bearer $UPBUTLER_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "name": "SMS credits",
    "kind": "balance",
    "url": "https://api.example.com/v1/account",
    "headers": { "X-Api-Key": "…" },
    "balance": { "remainingPath": "data.credits.remaining", "unit": "credits", "minDays": 5 }
  }'
PresetReadsKey
openrouter_creditsGET /api/v1/credits: total credits minus total usage of the accountManagement (provisioning) key
openrouter_keyGET /api/v1/key: the remaining spending limit of that key. A key without a limit reports usage only and stays up.The key itself
customAny JSON endpoint. Give remainingPath, or totalPath and usedPath. The key is sent as a bearer token; use headers for other schemes.Whatever the endpoint needs

Thresholds

FieldDefaultEffect
exhaustedAt0down when the remaining amount is at or below it
minAmountoffdegraded when the remaining amount is at or below it
minDays7degraded when the projected run-out is within this many days

How the run-out date is projected

  • Burn per day is what was spent over the last 7 days of readings, divided by the time they span.
  • Only decreases count as spending, so a top-up in the window doesn't hide the burn.
  • There is no projection until there are 3 readings spanning at least an hour, and none when nothing was spent. The amount thresholds work from the first reading.

GET /monitors/:id/balance returns the readings (120 days are kept), the burn per day, daysLeft and runOutAt. The latest values are also on the monitor as balance.last.

Plans

PlanMCP + LLM monitorsBalance monitorsFastest MCP interval
Free1Included5 min
Starter5Included1 min
Pro25Included30 s
Business100Included30 s

All three kinds count toward the plan's monitor total. See Plans & billing.

In AI incident reports

When one of these monitors opens an incident, the analysis is given the failing step, the error class, the step timings and the schema changes of the last 7 days, so a report can say "tool search lost required param query" or "the provider returned 529 for 4 consecutive checks" instead of "the check failed". Keys and header values are never part of it.