Know when your model provider is slow, not just down
LLM API monitoring with your own key. Every check streams one tiny completion from OpenAI, Anthropic, OpenRouter or any compatible endpoint, and measures what your users feel: time to first token, a stream that finishes, and the reason when it does not.
Their status page, and your key
A provider's status page is an announcement. Our own keyless probe of a provider's API edge shows that DNS, TLS and routing work; as each service page says, it does not prove that authenticated requests succeed. Only a call with your key, to your model, tells you whether your product works.
An LLM monitor makes that call: the prompt (by default Reply with the single word: pong) with stream: true and a small output cap. The stream must end with a finish signal; one that just stops is reported as truncated. Add expectContains or expectRegex to check the answer, and degradedTtftMs or degradedTotalMs to be told when it gets slow.
- ttfbMs
- Until the response headers arrived
- ttftMs
- Time to first token: until the first content (or reasoning) delta arrived
- totalMs
- Until the stream finished
- chunks
- Content deltas received
curl -X POST https://upbutler.com/api/v1/monitors \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "OpenRouter gpt-5-mini",
"kind": "llm",
"url": "https://openrouter.ai/api/v1",
"intervalSec": 600,
"llm": {
"model": "openai/gpt-5-mini",
"apiKey": "sk-or-…",
"expectContains": "pong",
"degradedTtftMs": 3000
}
}'{
"name": "Claude Haiku",
"kind": "llm",
"url": "https://api.anthropic.com",
"llm": { "api": "anthropic", "model": "claude-haiku-4-5", "apiKey": "sk-ant-…" }
}429 is not 529: every failure gets a class
The class is in evidence.llm.errorClass and in the check's error text. The response body is read before the status code, because providers disagree on codes: OpenAI reports an empty wallet as 429 insufficient_quota, Anthropic as 400. Like every monitor, a failure only changes the state after failureThreshold consecutive checks, so one 529 pages nobody.
| Class | Trigger | Monitor state |
|---|---|---|
| rate_limited | HTTP 429: your key is throttled. Retry-After is recorded. | degraded |
| overloaded | HTTP 529 or 503, or an overloaded_error inside the stream. | down |
| model_unavailable | model_not_found, "deprecated", "decommissioned", HTTP 404 or 410. | down |
| billing | HTTP 402, insufficient_quota, "credit balance is too low". | down |
| auth | HTTP 401 or 403. | down |
| truncated | The stream ended without a finish signal. | down |
| timeout · network · provider_error · bad_request · assertion | No answer in time · connection failed · other 5xx · other 4xx · the response failed your rule. | down |
Retired models. Response headers named Sunset or Deprecation are recorded with the check. When a Sunset date is within deprecationDays (30 by default) the monitor goes degraded with "Model will be retired in N days". If the provider answers with a different model than you asked for, both names are shown.
What it costs you in tokens
The completion runs on your key, so your provider bills it. Two limits hold on every plan: the interval is never shorter than 5 minutes, and maxTokens is at most 64 (16 by default). With the default prompt a check is about 25 input tokens and at most 16 output tokens.
| Interval | Checks / month | Input tokens | Output tokens (max) | At $0.25 / $2 per 1M | At $3 / $15 per 1M |
|---|---|---|---|---|---|
| 10 min (default) | 4,320 | 108k | 69k | about $0.17 | about $1.36 |
| 5 min | 8,640 | 216k | 138k | about $0.33 | about $2.72 |
| 1 hour | 720 | 18k | 12k | about $0.03 | about $0.23 |
Multiply by the number of regions you check from. A failure is re-checked once after 15 seconds, which adds a few completions during an outage.
curl -X POST https://upbutler.com/api/v1/monitors \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"name": "OpenRouter credits",
"kind": "balance",
"balance": { "preset": "openrouter_credits", "apiKey": "sk-or-…", "minDays": 7, "minAmount": 10 }
}'The outage you can predict: an empty balance
A balance monitor reads your remaining credits once an hour and projects the run-out date from what was spent over the last 7 days of readings. Only decreases count as spending, so a top-up does not hide the burn. It goes degraded when the run-out is within minDays (7 by default) or the amount falls to minAmount, and down when it reaches zero.
Presets for OpenRouter account credits and key limits; any JSON endpoint with a remainingPath works too. Balance monitors are included on every plan. For budgets across providers and agent runs, see spend guard.
Is it them? Check the provider first
Live status of the model providers, read from each provider's own status page where it publishes one, with a JSON verdict your agent can read before it retries. Declare one as a dependency and your incidents are tagged "likely upstream" when theirs began around the same time.
- OpenAI status
- Anthropic (Claude) status
- OpenRouter status
- Google Gemini status
- Groq status
- Mistral status
- All AI services
MCP and LLM monitors share one allowance: 1 on Free, 5 on Starter, 25 on Pro, 100 on Business. Related: MCP server monitoring. Full list on pricing.
LLM monitoring questions
What is LLM API monitoring?
LLM API monitoring sends a real, tiny completion to your model provider on a schedule, with your own key, and records whether the stream starts, how long the first token takes and whether it finishes. It catches what a status page and a plain HTTP check miss: your key being rate limited, a model being retired, an empty credit balance, a stream that starts and dies.
Why not just watch the provider's status page?
A provider status page tells you what the provider has announced, for everyone. It does not know whether authenticated calls with your key, your model and your region succeed. UpButler shows both: the service status catalog reads each provider's own status, and an LLM monitor measures your own calls. When your monitor fails while the provider has an open incident, your incident is tagged "likely upstream" with theirs linked.
What is the difference between a 429 and a 529?
A 429 means your key is being throttled: the provider is fine, you are over a limit. UpButler classifies it as rate_limited and marks the monitor degraded, with Retry-After recorded. A 529 (or 503, or an overloaded_error inside the stream) means the provider is overloaded: that is overloaded, and the monitor goes down after the usual confirmation.
Which providers and models work?
Any OpenAI-compatible /chat/completions endpoint, which covers OpenAI, OpenRouter, Azure OpenAI, vLLM, LiteLLM, Ollama, Together, Groq and most gateways, and the native Anthropic Messages API. Set llm.api to anthropic for the latter.
What does it cost in tokens?
The completion runs on your key, so your provider bills it. With the default prompt a check is about 25 input tokens and at most 16 output tokens: about 4,320 checks a month at the default 10-minute interval, roughly $0.17 a month on a model priced $0.25 / $2 per million tokens. The interval is never shorter than 5 minutes and maxTokens is capped at 64, on every plan.
Is my API key safe?
The key is encrypted at rest (AES-256-GCM) and is never returned by the API, the dashboard, events or webhooks; you see its last four characters. Error text and stored response snippets are scrubbed of the key before they are saved. We recommend a key made for this monitor alone, limited to one cheap model and with a small spending limit where the provider supports it.
Will I know before a model is retired?
Response headers named Sunset or Deprecation are recorded with each check. When a Sunset date is within deprecationDays (30 by default), the monitor goes degraded with "Model will be retired in N days". When the model is already gone, the check fails as model_unavailable.
Can it warn me before my credits run out?
Yes, with a balance monitor. It reads the remaining credits (OpenRouter presets, or any JSON endpoint), works out the burn per day from the last 7 days of readings, and goes degraded when the projected run-out date is within minDays (7 by default). Balance monitors are included on every plan.
Measure the call your users make.
1 LLM monitor and credit monitoring are free. 100 on Business.