Monitoring
Agent Runs
A heartbeat says an agent is alive. Agent Runs say whether it is doing the right thing: each run reports its steps and spend, UpButler flags the ones that are stuck, looping, failing or burning money, pages someone, and every answer it gives the agent says whether to continue.
How it works
- Create a monitor of kind
agent. It has an ingest token (ar_…). One monitor takes many concurrent runs: a fleet shares it and each run names itsagent. - The agent calls start, then progress at every step, then end. No API key is needed for these: the token is the secret, as with heartbeats.
- UpButler applies the detection rules. A flagged run sets the monitor degraded or down through the same state machine as every other monitor: incident, alerts, status page components, AI report, escalation, agent responders.
- Every progress response carries the kill switch:
{"continue": false, "action": "stop", "reason": "…"}when a person, an agent responder or a policy stopped the run.
curl -X POST https://upbutler.com/api/v1/monitors \
-H "Authorization: Bearer $UPBUTLER_API_KEY" \
-H "Content-Type: application/json" \
-d '{"name": "Support triage fleet", "kind": "agent", "agent": {"budget": {"usd": 2, "steps": 60}}}'
# → { "_id": "mon_…", "agent": { "token": "ar_Xf3kq9LmR2vT8wYzN4bC1dE7", "startUrl": "https://upbutler.com/api/v1/runs", … } }With an API key you can skip this step: POST /runs {"agent": "triage-bot"} finds the agent monitor by name and creates it on first use. An agent monitor counts as one monitor of your plan, however many runs it takes.
The run protocol
Start
curl -X POST https://upbutler.com/api/v1/runs \
-H "Content-Type: application/json" \
-d '{
"token": "ar_Xf3kq9LmR2vT8wYzN4bC1dE7",
"agent": "triage-bot-3",
"runId": "job-4812",
"task": "Triage ticket #4812",
"budget": {"usd": 2, "tokens": 400000, "steps": 60, "durationSec": 900},
"expectedDurationSec": 180,
"metadata": {"ticket": 4812}
}'{
"ok": true,
"continue": true,
"id": "run_0n3q8jz4k7m2p9x1c5v",
"token": "rt_Qm7vB2xK9pL4sT1nW8cZ3yH6",
"progressUrl": "https://upbutler.com/api/v1/runs/rt_Qm7vB2xK9pL4sT1nW8cZ3yH6/progress",
"endUrl": "https://upbutler.com/api/v1/runs/rt_Qm7vB2xK9pL4sT1nW8cZ3yH6/end",
"pingUrl": "https://upbutler.com/api/v1/runs/rt_Qm7vB2xK9pL4sT1nW8cZ3yH6/progress",
"monitor": { "id": "mon_0n3q8jy9x1c4v7b2n6m", "name": "Support triage fleet" },
"run": { "id": "run_0n3q8jz4k7m2p9x1c5v", "status": "running", "steps": 0, "usd": 0, "tokens": 0, "flags": [], "budgetUse": [] }
}| Field | Type | Description |
|---|---|---|
token | string | Ingest token of the agent monitor (ar_…). Or send it as "Authorization: Bearer ar_…". Not needed with an API keymax length 100 |
monitor | string | With an API key: the agent monitor (id or name). Created when it does not exist. Default: the agent namemax length 120 |
agent | string | Which agent this is, e.g. "triage-bot". Runs of a fleet share one monitor and differ by agentmax length 80 |
runId | string | Your own id for this run. Starting twice with the same one returns the same run (safe to retry)max length 120 |
task | string | What the run is doing, in a sentence. Shown in alerts and to respondersmax length 500 |
budget | object | Hard limits for this run: usd, tokens, steps, durationSec |
budget.usd | number | Spend in USD, as reported by the agentmax 1,000,000 |
budget.tokens | integer | max 1,000,000,000,000 |
budget.steps | integer | max 1,000,000 |
budget.durationSec | integer | max 2,592,000 |
expectedDurationSec | integer | min 1 · max 2,592,000 |
metadata | map<string, any> | Small JSON object kept with the run (max 4 KB) |
The token can also be sent as Authorization: Bearer ar_…. Starting twice with the same runId returns the same run, so a start is safe to retry.
Progress
curl -X POST https://upbutler.com/api/v1/runs/rt_Qm7vB2xK9pL4sT1nW8cZ3yH6/progress \
-H "Content-Type: application/json" \
-d '{"step": "search", "message": "Looking up similar tickets", "toolCalls": ["search_tickets"], "usd": 0.012, "tokens": 1840}'{
"ok": true,
"continue": true,
"run": {
"id": "run_0n3q8jz4k7m2p9x1c5v", "status": "running", "steps": 4, "usd": 0.31, "tokens": 48210, "flags": [],
"budgetUse": [{ "dimension": "usd", "used": 0.31, "limit": 2, "pct": 15.5 }]
}
}{
"ok": true,
"continue": false,
"action": "stop",
"reason": "Automatic stop: over budget: $2.04 of $2.00",
"run": { "id": "run_0n3q8jz4k7m2p9x1c5v", "status": "running", "steps": 41, "usd": 2.04, "tokens": 391820, "flags": ["budget"], "budgetUse": [ … ] }
}| Field | Type | Description |
|---|---|---|
step | string | number | Step label ("search") or counter (12). Labels count as one step each; a number sets the step countmax length 120 |
message | string | max length 500 |
status | string | The agent's own state label, e.g. "planning"max length 60 |
usd | number | Spend since the last ping (added to the total)min 0 · max 100,000 |
tokens | integer | Tokens since the last ping (added to the total)min 0 · max 10,000,000,000 |
toolCalls | string[] | integer | Tool names called in this step, or just how manymax 50 items · min 0 · max 10,000 |
fingerprint | string | What makes this step "the same as before" for loop detection, e.g. a hash of tool name + arguments. Default: step label + status + message + tool namesmax length 200 |
totals | object | Running totals instead of deltas, for agents that already keep them |
totals.usd | number | min 0 |
totals.tokens | integer | min 0 |
totals.steps | integer | min 0 |
usdandtokensare deltas: what was spent since the last ping. They are what the agent computed; UpButler does not proxy or price your LLM calls. Agents that keep running totals sendtotalsinstead.- A step label counts as one step; a numeric
stepsets the step count. - An empty body is a keepalive: "still working", and a look at the kill switch.
- A plain-text body becomes the
message, and fields may be sent as query parameters, socurl -d "exporting users" "$PROGRESS_URL?usd=0.02"works. - The path takes the run token (
rt_…, no key) or the run id (run_…) together with an API key.
End
curl -X POST https://upbutler.com/api/v1/runs/rt_Qm7vB2xK9pL4sT1nW8cZ3yH6/end \
-H "Content-Type: application/json" \
-d '{"status": "success", "summary": "Routed to billing"}'
# status: success | failed | cancelled (+ "error" for failed, final "usd"/"tokens" not reported yet)Always end a run, also when it failed or was told to stop. A run that never ends is flagged as stuck after 10 min and closed as lost after 1 h of silence (abandonAfterSec); lost runs count as failed. Ending twice is fine.
| Status | Meaning | Counts as failed |
|---|---|---|
running | Started, not ended yet | |
success | Ended by the agent | No |
failed | Ended by the agent with an error | Yes |
cancelled | Ended early by the agent's own decision | No |
stopped | Ended as cancelled after a stop from the kill switch, or went quiet for 2 minutes after receiving one | Yes |
lost | Silent for abandonAfterSec; closed by UpButler | Yes |
The kill-switch contract
Every response to start, progress and signal has these fields:
continue | action | What the agent must do |
|---|---|---|
true | Carry on. | |
false | stop | Stop working now: no further tool calls or LLM calls. Call end with status cancelled. reason says why. |
false | pause | Do nothing, wait retryAfterSec (15 s) and ask again with GET /runs/{token}/signal, until it says continue or stop. |
true | slow_down | Carry on, but wait retryAfterSec between steps. |
Who sets it:
- People: the Stop button on the monitor page or in a run's timeline, the Stop run button in every alert about an agent monitor (next to Acknowledge in email, Slack and Telegram, a link in Discord,
links.stopRunsin webhook payloads; a signed link that opens a confirmation page and works for 3 days), or the API. - Agent responders:
POST /incidents/:id/runs/stopwith their incident token. See below. - Policy: rules listed in
autoStopstop the run the moment they detect. By default that is onlybudget(a hard budget breach). A person can resume a run a policy stopped; the policy leaves that run alone afterwards.
# One run
curl -X POST https://upbutler.com/api/v1/runs/run_0n3q8jz4k7m2p9x1c5v/stop \
-H "Authorization: Bearer $UPBUTLER_API_KEY" -H "Content-Type: application/json" \
-d '{"action": "stop", "reason": "Looping on search_tickets"}'
# Every flagged run of a monitor (or all of them: "onlyFlagged": false)
curl -X POST https://upbutler.com/api/v1/monitors/mon_0n3q8jy9x1c4v7b2n6m/runs/stop \
-H "Authorization: Bearer $UPBUTLER_API_KEY" -H "Content-Type: application/json" \
-d '{"onlyFlagged": true}'
# Changed your mind before the agent acted on it
curl -X POST https://upbutler.com/api/v1/runs/run_0n3q8jz4k7m2p9x1c5v/resume -H "Authorization: Bearer $UPBUTLER_API_KEY"Reading the switch without reporting
curl https://upbutler.com/api/v1/runs/rt_Qm7vB2xK9pL4sT1nW8cZ3yH6/signal
# → { "ok": true, "continue": true, "run": { "id": "run_…", "status": "running" } }For a watchdog thread, or between the steps of one long tool call. A poll is not progress: a run that only polls is still flagged as stuck.
Stop webhook
Set agent.stopWebhookUrl on the monitor for orchestrators that can cancel a job from outside (a queue, Kubernetes, a workflow engine). It receives run.stop_requested (stop, pause, slow down) and run.resumed right away, signed per Standard Webhooks and retried like every delivery. The signing secret is shown once: POST /monitors/:id/runs/webhook-secret (or "new signing secret" on the monitor page).
POST https://agents.example.com/hooks/upbutler-stop
webhook-id: evt_0n3q8jz9d2f4g6h8j0k
webhook-timestamp: 1791590451
webhook-signature: v1,K5oZfzN95Z9UVu1EsfQmfVNQhnkZ2pj9o9NDN/H/pI4=
{
"id": "evt_0n3q8jz9d2f4g6h8j0k",
"type": "run.stop_requested",
"timestamp": "2026-10-10T09:20:51.204Z",
"workspaceId": "ws_0n3q8jy8r5t1y3u7i9o",
"data": {
"monitor": { "id": "mon_0n3q8jy9x1c4v7b2n6m", "name": "Support triage fleet" },
"run": { "id": "run_0n3q8jz4k7m2p9x1c5v", "runId": "job-4812", "agent": "triage-bot-3", "task": "Triage ticket #4812", "status": "running", "steps": 41, "usd": 2.04, "tokens": 391820 },
"signal": { "continue": false, "action": "stop", "reason": "Automatic stop: over budget: $2.04 of $2.00" },
"stop": { "action": "stop", "by": "UpButler", "source": "policy", "at": "2026-10-10T09:20:51.198Z" }
}
}Detection rules
Rules are evaluated on the server, on every ping and once a minute, and they are deterministic: the same reports always give the same result. Each rule maps to down, degraded or off per monitor (agent.severity). Down opens an incident and pages; degraded alerts without an incident.
| Rule | Detects | Setting | Default | Monitor |
|---|---|---|---|---|
missing | No run started within periodSec + graceSec (only with a schedule, and only while nothing is running) | periodSec, graceSec | no schedule | down |
stuck | A running run reported no progress for stuckAfterSec. A poll of /signal is not progress | stuckAfterSec | 10 min | down |
overrun | Running longer than the run’s expectedDurationSec × overrunFactor, even while it keeps reporting | overrunFactor | 3× | degraded |
loop | The last pings are one fingerprint repeated loopRepeats times, or a cycle of 2 to 4 fingerprints repeated that many times | loopRepeats | 5 repeats | down |
budget | Above any limit of the budget: usd, tokens, steps or durationSec (the run’s budget, else the monitor’s) | budget | none | down |
budget_soft | At softBudgetPct of any limit, and not over it yet | softBudgetPct | 80% | degraded |
failure_rate | failed of the last window finished runs ended as failed, stopped or lost (cancelled does not count). 2 successful runs in a row clear it at once | failureRate | 3 of 5 | down |
cost_spike | A running run spent more than factor × the median spend of the last finished runs (up to 20; needs minRuns of them with a cost) | costSpike | 3×, 5 runs | degraded |
- Fingerprints. A ping's fingerprint is the explicit
fingerprintyou send, otherwise its step label +status+message+ tool names (case and whitespace are ignored). A numeric step is a counter and is left out, so an agent that counts up while repeating itself is still caught. Empty pings have no fingerprint. The SDKs'toolCall(name, args)sends a hash of the tool name and its arguments: callingsearchwith different queries is work, calling it with the same query five times is a loop. - Loops end. A loop is what the run is doing now. One different step clears the flag, and the monitor recovers.
- A flag lives as long as its run. When a flagged run ends, its flags go with it and the monitor recovers.
missingclears when the next run starts;failure_rateclears after 2 successful runs in a row, or when the failures leave the window. - Stopping is not fixing. What happens to the incident depends on how the flagged run ended. If it ended in
success, the incident resolves like any recovered monitor. If it wasstopped(by a person, an agent responder or a policy),failedorlost, the monitor recovers but the incident stays open with a note saying why, until a person or an agent responder resolves it. The next detection opens a new incident. - Several runs. The monitor is as bad as its worst run. A second run flagged while the monitor is already down does not page again; subscribe a channel to
run.flaggedif you want one event per detection.
Monitor settings
All under agent on POST /monitors and PATCH /monitors/:id, and in the monitor form.
| Field | Type | Description |
|---|---|---|
periodSec | integer | null | Expect a run to start at least this often (null = no schedule). Same as the top-level periodSecmin 30 · max 2,678,400 · nullable |
graceSec | integer | min 0 · max 86,400 |
budget | object | null | Default budget per run; a run can bring its own. null clears itnullable |
budget.usd | number | Spend in USD, as reported by the agentmax 1,000,000 |
budget.tokens | integer | max 1,000,000,000,000 |
budget.steps | integer | max 1,000,000 |
budget.durationSec | integer | max 2,592,000 |
stuckAfterSec | integer | No progress for this long → stuck (default 600)min 30 · max 86,400 |
overrunFactor | number | Running longer than expectedDurationSec × this → overrun (default 3)min 1 · max 100 |
loopRepeats | integer | The same step, or a cycle of up to 4 steps, repeated this many times → loop (default 5)min 2 · max 50 |
softBudgetPct | number | Warn (degraded) at this share of any budget (default 80)min 1 · max 100 |
failureRate | object | failed of the last window finished runs failed (default 3 of 5) |
failureRate.failedrequired | integer | min 1 · max 50 |
failureRate.windowrequired | integer | min 1 · max 50 |
costSpike | object | Spend above factor × the median of recent runs (default 3×, needs 5 runs) |
costSpike.factorrequired | number | min 1.1 · max 100 |
costSpike.minRunsrequired | integer | min 1 · max 20 |
autoStop | string[] | Rules that set the kill switch on their own (default ["budget"])budgetloopoverruncost_spike |
abandonAfterSec | integer | A run silent for this long is closed as lost (default 3600)min 60 · max 604,800 |
severity | map<string, string> | Per rule: down, degraded or off |
stopWebhookUrl | string (url) | null | Receives a signed POST when a run is told to stop, pause, slow down or resumemax length 2,000 · nullable |
Incidents, alerts and AI reports
An agent monitor is a monitor. When it goes down, an incident opens (and joins an open one on the same status page), your channels and escalation policy are used, components that follow it change status, and reminders, acknowledgement and quick mute work as everywhere else.
- The alert text names the run: Run "Triage ticket #4812" (triage-bot-3) is looping: the same step repeated 5× in a row: search_tickets.
- The AI incident report gets the run context: task, the last steps, spend against budget, the detections and whether a stop was sent and received.
- Stopping a run from anywhere adds a line to the open incident's timeline: who stopped what, and why. The incident then stays open after the run ends (see Stopping is not fixing).
- Events:
run.flagged,run.stop_requested,run.resumed(for channels that subscribe to them; see Events).
Agent responders: one agent on call for another
When the incident pages an agent responder, the handoff packet's monitor section carries the runs, and the incident token can stop them. The action is stop_runs in the responder's allowed actions (on by default for new responders). It is limited to runs of the incident's own agent monitors and, like mute, needs the incident claimed first.
"monitors": [{
"id": "mon_0n3q8jy9x1c4v7b2n6m", "name": "Support triage fleet", "kind": "agent", "state": "down",
"runs": {
"running": 7,
"detection": { "stuckAfterSec": 600, "loopRepeats": 5, "autoStop": ["budget"], "defaultBudget": { "usd": 2, "steps": 60 } },
"flagged": [{
"id": "run_0n3q8jz4k7m2p9x1c5v", "agent": "triage-bot-3", "task": "Triage ticket #4812",
"startedAt": "2026-10-10T09:12:03.000Z", "lastProgressAt": "2026-10-10T09:20:44.000Z",
"steps": 38, "usd": 1.8, "tokens": 352000,
"budgetUse": [{ "dimension": "usd", "used": 1.8, "limit": 2, "pct": 90 }],
"detections": [{ "rule": "loop", "severity": "down", "at": "2026-10-10T09:19:10.000Z", "message": "the same step repeated 5× in a row: search_tickets" }],
"killSwitch": null,
"lastSteps": [{ "at": "…", "type": "progress", "step": "search_tickets", "toolCalls": ["search_tickets"], "usd": 0.05 }, …]
}],
"recent": [{ "id": "run_…", "agent": "triage-bot-1", "status": "success", "endedAt": "…", "usd": 0.13 }, …]
}
}]# With the incident token from the page (the responder must allow "stop_runs", and claim first)
curl -X POST https://upbutler.com/api/v1/incidents/$INCIDENT_ID/claim -H "Authorization: Bearer $INCIDENT_TOKEN"
curl https://upbutler.com/api/v1/incidents/$INCIDENT_ID/runs -H "Authorization: Bearer $INCIDENT_TOKEN"
curl -X POST https://upbutler.com/api/v1/incidents/$INCIDENT_ID/runs/stop \
-H "Authorization: Bearer $INCIDENT_TOKEN" -H "Content-Type: application/json" \
-d '{"reason": "Same fetch five times with no new data; $1.80 of $2.00 spent"}'
# → { "action": "stop", "stopped": [{ "id": "run_…", "agent": "triage-bot-3", "task": "Triage ticket #4812" }], … }After the stop the monitor recovers once the runs end, so POST /incidents/:id/verify passes; the incident itself is the responder's (or a person's) to resolve or hand back, with a note on what was wrong.
Without runIds the flagged runs are stopped; healthy runs of the same fleet are left alone unless named. The incident token cannot call /runs/:id/stop or read other runs.
SDKs
The helpers do the three calls, account for LLM usage and honour the kill switch.
import { withRun, RunStopped } from '@upbutler/sdk';
try {
await withRun(
{ token: process.env.UPBUTLER_RUN_TOKEN!, agent: 'triage-bot', task: `Triage ticket #${id}`, budget: { usd: 2, steps: 60 }, expectedDurationSec: 180 },
async (run) => {
// A tool call: fingerprinted by name + arguments, so the same call again and again is a loop.
const similar = await search(q);
await run.toolCall('search_tickets', { q }); // throws RunStopped when told to stop
// LLM usage: UpButler reads result.usage, it never sees the call. Prices are USD per million tokens.
const res = await run.llm((signal) => anthropic.messages.create({ model, max_tokens: 1024, messages }, { signal }), { inputPerMTok: 3, outputPerMTok: 15 });
await run.step('route', { message: 'billing' }); // sends the buffered usage with it
},
);
} catch (e) {
if (!(e instanceof RunStopped)) throw e; // already ended as cancelled
console.warn('Stopped by UpButler:', e.signal.reason);
}import os
from upbutler import start_run, RunStopped
try:
with start_run(token=os.environ["UPBUTLER_RUN_TOKEN"], agent="triage-bot", task=f"Triage ticket #{id}",
budget={"usd": 2, "steps": 60}, expected_duration_sec=180) as run:
similar = search(q)
run.tool_call("search_tickets", {"q": q}) # raises RunStopped when told to stop
# LLM usage: UpButler reads result.usage, it never sees the call. Prices are USD per million tokens.
res = run.llm(lambda: client.messages.create(model=model, max_tokens=1024, messages=messages),
input_per_mtok=3, output_per_mtok=15)
run.step("route", message="billing") # sends the buffered usage with it
except RunStopped as e: # already ended as cancelled
print("Stopped by UpButler:", e.signal.get("reason"))import { startRun } from '@upbutler/sdk';
const run = await startRun({ token: process.env.UPBUTLER_RUN_TOKEN!, agent: 'triage-bot', onStop: 'return' });
try {
for (const ticket of tickets) {
const signal = await run.step('triage', { message: `ticket ${ticket.id}` });
if (!signal.continue) break; // the contract: stop when told to
await triage(ticket);
}
} finally {
await run.end({ status: run.stopped ? 'cancelled' : 'success' });
}| TypeScript | Python | Does |
|---|---|---|
startRun(opts) · client.runs.start(opts) | start_run(...) · client.runs.start(...) | Starts a run and returns a handle. With a token no API key is needed |
withRun(opts, fn) | with start_run(...) as run: | Ends the run for you: success, failed with the error, or cancelled when stopped (RunStopped is re-raised) |
run.progress(input) · run.step(label) | run.progress(...) · run.step(label) | One ping. Throws RunStopped on stop, waits on pause, sleeps on slow-down |
run.toolCall(name, args) | run.tool_call(name, args) | A ping fingerprinted by tool name + arguments |
run.usage(usage, pricing) · run.llm(call, pricing) | run.usage(usage, input_per_mtok, output_per_mtok) · run.llm(call, ...) | Adds tokens and cost from an Anthropic or OpenAI usage object (or a cost you computed). Sent with the next ping; no request of its own |
run.poll() · run.abortSignal · pollSec | run.poll() | Reads the switch without reporting. abortSignal aborts on stop: pass it to fetch or your LLM client. pollSec polls in the background |
run.end(input) | run.end(status, summary, error) | Ends the run with any usage not sent yet. Safe to call twice |
CLI: wrap any command
export UPBUTLER_RUN_TOKEN=ar_Xf3kq9LmR2vT8wYzN4bC1dE7
# Start → run the command → end with its exit code. A stop sends SIGTERM (SIGKILL after 10 s).
upbutler run --agent nightly-report --task "Build the weekly report" \
--budget-usd 5 --budget-seconds 1800 --expect 600 -- python agent.py
# Inside the command, one line of output is one progress ping (the line is not printed):
echo '::upbutler:: {"step":"export","usd":0.02,"tokens":1800,"toolCalls":["sql_query"]}'
echo '::upbutler:: writing the report'
# The command also gets UPBUTLER_RUN_ID, UPBUTLER_RUN_PROGRESS_URL, UPBUTLER_RUN_SIGNAL_URL, UPBUTLER_RUN_END_URL.Output of the command counts as activity; marker lines carry steps and spend. Between pings the wrapper polls the kill switch every 15 seconds (--interval). The exit code of the command is kept: 0 ends the run as success, anything else as failed with the stderr tail, a stop as cancelled (exit 143). If UpButler is unreachable the command runs unmonitored.
Agent frameworks
These are thin: a ping where the framework tells you a tool ran, usage where it reports it, and a way out when the run is stopped. Adjust them to the version you use.
import { query } from '@anthropic-ai/claude-agent-sdk';
import { withRun } from '@upbutler/sdk';
await withRun({ token: process.env.UPBUTLER_RUN_TOKEN!, agent: 'repo-fixer', task: prompt, budget: { usd: 5 }, pollSec: 20 }, async (run) => {
const abortController = new AbortController();
run.abortSignal.addEventListener('abort', () => abortController.abort()); // stop → abort the agent loop
for await (const msg of query({ prompt, options: { abortController } })) {
if (msg.type === 'assistant') {
for (const block of msg.message.content) {
if (block.type === 'tool_use') await run.toolCall(block.name, block.input); // throws RunStopped
}
}
if (msg.type === 'result') run.usage({ total_cost_usd: msg.total_cost_usd, ...msg.usage });
}
});from agents import Agent, Runner, RunHooks
from upbutler import start_run
class UpButlerHooks(RunHooks):
def __init__(self, run):
self.run = run
async def on_tool_end(self, context, agent, tool, result):
# Fingerprint = tool + what it returned: the same answer again and again is a loop.
self.run.tool_call(tool.name, {"result": str(result)[:2000]}) # raises RunStopped → the run ends
async def on_agent_end(self, context, agent, output):
self.run.usage(tokens=context.usage.total_tokens)
with start_run(token=RUN_TOKEN, agent="support-agent", task=question, budget={"tokens": 400_000}) as run:
result = await Runner.run(agent, question, hooks=UpButlerHooks(run))from upbutler import start_run, tool_fingerprint
with start_run(token=RUN_TOKEN, agent="research-graph", task=topic, budget={"steps": 80}) as run:
for chunk in graph.stream(inputs, stream_mode="updates"):
for node, update in chunk.items():
# One ping per node that ran; the same node producing the same update is a loop.
run.step(node, fingerprint=tool_fingerprint(node, update)) # raises RunStopped → leaves the stream- Claude Agent SDK: tool calls are read from the assistant messages; cost and tokens from the final
resultmessage, so the budget in USD is only known at the end of a query. Use astepsordurationSecbudget for a limit that holds during the run.pollSecplus the abort controller stops a query in the middle of a long turn. - OpenAI Agents SDK: the tool hooks do not carry the arguments, so the fingerprint is the tool plus its result. An exception raised in a hook ends the run, which is how
RunStoppedgets out. - LangGraph: one ping per node update from
stream(stream_mode="updates"). Leaving the loop stops the graph.
Retention and limits
- Runs and their timelines are kept for 30 days after their last activity, and a monitor keeps its newest 1,000 finished runs.
- A run's timeline stores its first 1,000 events; totals, detection and the kill switch keep working past that.
- Up to 500 runs in progress per monitor. 600 pings a minute per run, 600 starts a minute per monitor.
- Detection, the kill switch, the stop webhook and automatic stops are on every plan. Runs started per month:
| Plan | Runs / month | Agent monitors |
|---|---|---|
| Free | 1,000 | Count as monitors (10) |
| Starter | 10,000 | Count as monitors (25) |
| Pro | 100,000 | Count as monitors (75) |
| Business | 1,000,000 | Count as monitors (300) |
Over the quota, start answers 402 plan_limit with the plan that lifts it; pings of runs already going are never refused, and the SDK helpers carry on unmonitored.
API reference
| Endpoint | Auth | Does |
|---|---|---|
POST /runs | ar_ token or API key | Start a run |
POST /runs/:id/progress | rt_ token, or run id + API key | Report a step, read the kill switch |
GET /runs/:id/signal | rt_ token, or run id + API key | Read the kill switch only |
POST /runs/:id/end | rt_ token, or run id + API key | End a run |
GET /monitors/:id/runs | API key | List runs (status, agent) with the fleet summary. MCP: runs_list |
GET /runs/:id | API key | A run with its timeline |
POST /runs/:id/stop · /resume | API key (write) | Set or lift the kill switch. MCP: runs_stop |
POST /monitors/:id/runs/stop | API key (write) | Stop all, or all flagged, runs of a monitor |
GET /incidents/:id/runs · POST /incidents/:id/runs/stop | API key or incident token | The runs behind an incident; stop them |
Field-level detail for every operation is in the REST API reference.