Monitoring
AI-stack monitors
Three monitor kinds for what agents depend on. They fail in ways an HTTP 200 check doesn't see: a tool loses a parameter, a stream starts and dies, a model is retired, the credits run out. They use the same confirmation, incidents, status pages, alerts and AI reports as every other monitor.
| Kind | Checks | Default interval | Regions |
|---|---|---|---|
mcp | Handshake, tool list, optional probe call, schema drift | 5 min (plan minimum applies) | Multi-region |
llm | One tiny streamed completion: first-token time, stream completes, error class | 10 min, never under 5 | Multi-region (each region is another completion) |
balance | Remaining credits and projected run-out date | 1 hour, never under 5 min | Main region only |
Create them like any monitor, with POST /monitors, the MCP tool monitors_create, or in the dashboard under Monitors → New monitor. The kind is inferred when you send an mcp, llm or balance object.
MCP server
curl -X POST https://upbutler.com/api/v1/monitors \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "Acme MCP",
"kind": "mcp",
"url": "https://mcp.acme.com/mcp",
"headers": { "Authorization": "Bearer <token for the MCP server>" },
"mcp": {
"probe": { "tool": "ping", "args": {}, "expectContains": "ok" },
"driftStatus": "degraded",
"driftHoldHours": 24
}
}'Each check speaks MCP over Streamable HTTP, and falls back to the older HTTP+SSE transport when the server only offers that:
initialize, thennotifications/initialized.tools/list, followingnextCursoruntil the list is complete. WithlistPrompts/listResources, alsoprompts/listandresources/listwhen the server advertises them.- If
mcp.probeis set: onetools/call. The check fails when the tool is missing, returnsisError, or its text doesn't satisfyexpectContains/expectRegex.
The time of every step is stored with each check (initMs, listMs, callMs), and GET /monitors/:id/metrics returns their p50 and p95.
What a failure means
evidence.mcp.failure | Meaning | Status |
|---|---|---|
auth | HTTP 401 or 403 at any step. The server is reachable; the credential is wrong or expired. | down |
unreachable | Connection failed, DNS failed, or HTTP 5xx. | down |
timeout | The check ran out of time; the error names the step. | down |
protocol | The endpoint answered, but not with valid MCP: a JSON-RPC error, a missing tools array, HTML. | down |
probe | The probe tool is gone, errored, or failed its assertion. | down |
drift | The server works, but its tool schema changed in a breaking way. | driftStatus |
Schema drift
The first successful check records the tool list as the baseline. Later checks compare what the server serves with that baseline, and every difference is classified:
| Severity | Changes | What happens |
|---|---|---|
| breaking | Tool removed · parameter removed · new required parameter · optional parameter became required · parameter type changed · enum value removed · prompt removed | Event, and the monitor goes to driftStatus (default degraded) with the diff as its error, which alerts your channels |
| non_breaking | Tool added · optional parameter added · parameter became optional · enum value added · constraint or output schema changed · resources changed | Event only. The baseline moves. |
| description | Only the wording of a tool or parameter description changed | Event only. The baseline moves. |
Description changes are listed on their own because a quietly edited tool description is how tool poisoning reaches an agent. Subscribe a webhook to monitor.schema_changed if you want to review each one.
Nested object parameters are compared by path (filters.date). A new required field inside an optional object is not breaking, because callers that don't send the object are unaffected. Key order and a server version bump are never drift.
After a breaking change
The monitor stays in driftStatus until one of three things happens:
- You accept it. Click Accept as new baseline on the monitor page, or call
monitors_schema_accept. Do this once your agents work with the new schema. - The server rolls it back. The tool list matches the baseline again and the monitor recovers by itself.
driftHoldHourspass (default 24, up to 720). The change is accepted automatically, so a monitor is never degraded forever.
curl -X POST https://upbutler.com/api/v1/monitors/mon_…/schema/accept \
-H "Authorization: Bearer $UPBUTLER_API_KEY"driftStatus: "down" opens an incident like any outage, with an AI report that names the tools and parameters. driftStatus: "none" records breaking changes without changing the monitor.
{
"baseline": { "hash": "8bc366949f88e4ce…", "tools": [ { "name": "search", "params": { "query": { "type": "string", "required": true } } } ] },
"pending": {
"at": "2026-10-09T12:04:11.000Z",
"until": "2026-10-10T12:04:11.000Z",
"summary": "tool `search` lost required param `query`"
},
"changes": [
{
"severity": "breaking",
"status": "pending",
"changes": [
{ "severity": "breaking", "type": "param_removed", "tool": "search", "param": "query",
"text": "tool `search` lost required param `query`" }
]
}
]
}{
"type": "monitor.schema_changed",
"data": {
"monitor": { "_id": "mon_…", "name": "Acme MCP", "kind": "mcp", "target": "https://mcp.acme.com/mcp" },
"schemaChange": {
"severity": "breaking",
"pending": true,
"summary": "tool `search` lost required param `query`",
"changes": [ { "severity": "breaking", "type": "param_removed", "tool": "search", "param": "query", "text": "…" } ]
}
}
}The event is not in the default set channels receive, so it doesn't add noise: breaking changes already alert through monitor.degraded or monitor.down. Add monitor.schema_changed to a channel's events to get all three severities as a tool changelog.
LLM endpoint
curl -X POST https://upbutler.com/api/v1/monitors \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "OpenRouter gpt-5-mini",
"kind": "llm",
"url": "https://openrouter.ai/api/v1",
"intervalSec": 600,
"llm": {
"model": "openai/gpt-5-mini",
"apiKey": "sk-or-…",
"expectContains": "pong",
"degradedTtftMs": 3000
}
}'url is the API base URL or the full endpoint. llm.api selects the wire format:
openai(default): any OpenAI-compatible/chat/completions. That covers OpenAI, OpenRouter, Azure OpenAI, vLLM, LiteLLM, Ollama, Together, Groq and most gateways.anthropic: the native Messages API (/v1/messages,x-api-key).
{
"name": "Claude Haiku",
"kind": "llm",
"url": "https://api.anthropic.com",
"llm": { "api": "anthropic", "model": "claude-haiku-4-5", "apiKey": "sk-ant-…" }
}Each check sends the prompt (default Reply with the single word: pong) with stream: true and a small output cap, and records:
| Metric | Meaning |
|---|---|
ttfbMs | Until the response headers arrived |
ttftMs | Time to first token: until the first content (or reasoning) delta arrived |
totalMs | Until the stream finished |
chunks | Content deltas received |
The stream must end with a finish signal (finish_reason or [DONE]; message_stop for Anthropic). A stream that just stops is reported as truncated. Optional rules: expectContains (not case-sensitive), expectRegex, degradedTtftMs, degradedTotalMs.
Error classes
The class is in evidence.llm.errorClass and in the check's error text. The response body is read before the status code, because providers disagree on codes: OpenAI reports an empty wallet as 429 insufficient_quota, Anthropic as 400.
| Class | Trigger | Status |
|---|---|---|
rate_limited | HTTP 429 (your key is throttled; Retry-After is recorded) | degraded |
overloaded | HTTP 529 or 503, or an overloaded_error inside the stream | down |
model_unavailable | model_not_found, "deprecated", "decommissioned", HTTP 404 or 410 | down |
billing | HTTP 402, insufficient_quota, "credit balance is too low" | down |
auth | HTTP 401 or 403 | down |
truncated | The stream ended without a finish signal | down |
timeout · network · provider_error · bad_request · assertion | No answer in time · connection failed · other 5xx · other 4xx · the response failed your rule | down |
Like every monitor, a failure only flips the state after failureThreshold consecutive checks, so one 529 doesn't page anyone.
Deprecation signals
Response headers named Sunset, Deprecation or anything containing "deprecat" are recorded with the check. When a Sunset date is within deprecationDays (default 30), the monitor goes degraded with "Model will be retired in N days". When the model is already gone, the check fails as model_unavailable. If the provider answers with a different model than you asked for, both names are shown in the check.
Cost
The completion runs on your key, so your provider bills it. Two limits hold on every plan: the interval is never shorter than 5 minutes, and maxTokens is at most 64 (default 16).
With the default prompt a check is about 25 input tokens and at most 16 output tokens:
| Interval | Checks / month | Input tokens | Output tokens (max) | At $0.25 / $2 per 1M | At $3 / $15 per 1M |
|---|---|---|---|---|---|
| 10 min (default) | 4,320 | 108k | 69k | about $0.17 | about $1.36 |
| 5 min | 8,640 | 216k | 138k | about $0.33 | about $2.72 |
| 1 hour | 720 | 18k | 12k | about $0.03 | about $0.23 |
Multiply by the number of regions. A failure is re-checked once after 15 seconds, which adds a few completions during an outage. "Test now" sends one completion. Reasoning models may spend the whole output cap on reasoning and return no text; raise maxTokens or drop the response rule for those.
Credit balance
# OpenRouter account credits (needs a management key)
curl -X POST https://upbutler.com/api/v1/monitors \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "OpenRouter credits",
"kind": "balance",
"balance": { "preset": "openrouter_credits", "apiKey": "sk-or-…", "minDays": 7, "minAmount": 10 }
}'
# Any JSON endpoint
curl -X POST https://upbutler.com/api/v1/monitors \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "SMS credits",
"kind": "balance",
"url": "https://api.example.com/v1/account",
"headers": { "X-Api-Key": "…" },
"balance": { "remainingPath": "data.credits.remaining", "unit": "credits", "minDays": 5 }
}'| Preset | Reads | Key |
|---|---|---|
openrouter_credits | GET /api/v1/credits: total credits minus total usage of the account | Management (provisioning) key |
openrouter_key | GET /api/v1/key: the remaining spending limit of that key. A key without a limit reports usage only and stays up. | The key itself |
custom | Any JSON endpoint. Give remainingPath, or totalPath and usedPath. The key is sent as a bearer token; use headers for other schemes. | Whatever the endpoint needs |
Thresholds
| Field | Default | Effect |
|---|---|---|
exhaustedAt | 0 | down when the remaining amount is at or below it |
minAmount | off | degraded when the remaining amount is at or below it |
minDays | 7 | degraded when the projected run-out is within this many days |
How the run-out date is projected
- Burn per day is what was spent over the last 7 days of readings, divided by the time they span.
- Only decreases count as spending, so a top-up in the window doesn't hide the burn.
- There is no projection until there are 3 readings spanning at least an hour, and none when nothing was spent. The amount thresholds work from the first reading.
GET /monitors/:id/balance returns the readings (120 days are kept), the burn per day, daysLeft and runOutAt. The latest values are also on the monitor as balance.last.
Plans
| Plan | MCP + LLM monitors | Balance monitors | Fastest MCP interval |
|---|---|---|---|
| Free | 1 | Included | 5 min |
| Starter | 5 | Included | 1 min |
| Pro | 25 | Included | 30 s |
| Business | 100 | Included | 30 s |
All three kinds count toward the plan's monitor total. See Plans & billing.
In AI incident reports
When one of these monitors opens an incident, the analysis is given the failing step, the error class, the step timings and the schema changes of the last 7 days, so a report can say "tool search lost required param query" or "the provider returned 529 for 4 consecutive checks" instead of "the check failed". Keys and header values are never part of it.