Skip to content

Operations

Running Busbar in production: process configuration, health/readiness, the metrics to watch, circuit-breaker and health-probe behavior, failover/exhaustion outcomes, governance/admin usage, and troubleshooting.

Busbar is a single native binary configured by two YAML files and environment variables.

Env varDefaultPurpose
BUSBAR_CONFIG/etc/busbar/config.yamlPath to the deployment config. The one bootstrap env var — it locates config.yaml itself.
Provider key varsn/aNamed by each provider’s api_key: { env: ... } reference (e.g. ANTHROPIC_KEY).
Token/secret varsn/aAnything referenced via ${VAR} in either file (client tokens, admin token, …).

Operational config moved into config.yaml (1.5.3). Several knobs that used to be env vars now live in config.yaml, so they are reviewable, --validate-checked, and part of the deployment artifact. The old env vars still work for one release (each logs a deprecation warning) but are the migration path, not the home:

Deprecated env varNew config.yaml key
BUSBAR_PROVIDERSproviders_file: (top-level; default providers.yaml next to config.yaml)
BUSBAR_CONFIG_OVERLAYconfig.overlay.file (see config mutability)
BUSBAR_WORKER_THREADSadvanced.worker_threads
BUSBAR_UPSTREAM_HTTP1_ONLYadvanced.upstream_http1_only
BUSBAR_UPSTREAM_H2_PRIOR_KNOWLEDGEadvanced.upstream_h2_prior_knowledge

TOKIO_WORKER_THREADS is still honored as a fallback for advanced.worker_threads.

The top-level config: block governs whether the admin API may change config at runtime and where those changes persist. Durable by default: with nothing specified, config is mutable and admin-API mutations persist to busbar-overlay.json next to config.yaml — so a group or hook provisioned over the admin API survives a restart out of the box.

config:
locked: false # false = mutable (admin API may change config); true = immutable
overlay:
file: busbar-overlay.json # where mutations persist; default = next to config.yaml
  • locked: true — an immutable/GitOps deployment: admin-API config mutations are refused at runtime (edit config.yaml and POST /config/reload instead). The overlay is ignored.
  • Boot invariant — locked XOR a writable overlay. A mutable config with no writable overlay (you set overlay: false, or the config directory is read-only) refuses to boot, with a message telling you to either point config.overlay.file at a writable path or set config.locked: true. This makes “applied but silently lost on restart” impossible to reach.
    • Read-only config dir (e.g. /etc/busbar on a read-only mount): the default overlay path is unwritable, so an unconfigured mutable busbar refuses to boot. Fix: set a writable config.overlay.file, or set config.locked: true (a read-only/GitOps deployment never persists runtime mutations anyway). See the upgrade note.

Worker threads and scaling. Busbar’s request path is CPU-bound (parse, translate, serialize), so throughput scales with worker threads. The default is one worker per available core (available_parallelism, which respects CPU affinity and the cgroup cpuset, but not the CFS cpu.max bandwidth quota, which it cannot see), which gives linear scaling: ~9,750 req/s per core, sub-millisecond, to ~156k on 16 cores in our benchmark. Each worker carries a thread stack and, on glibc, its own malloc arena, so idle memory grows slowly with the count. For a footprint-sensitive sidecar set advanced.worker_threads: 1 (or 2). On a CPU-quota-limited pod (a k8s CPU limit on a many-core node) the default sizes to the node’s full core count and oversubscribes the quota: set advanced.worker_threads to your CPU limit; likewise to cap a shared box, set it to the cores you want Busbar to use. Scale up by default, tune down deliberately. (Before 1.4.0 the default was capped at min(cores, 4), which pinned throughput to ~4 cores regardless of box size, set the value explicitly on older binaries.)

Startup is fail-loud: an unset ${VAR}, an unknown provider reference, an unknown protocol or auth mode, or an invalid on_exhausted action stops the process with a diagnostic. A provider whose key env var is empty logs a warning and runs (its lane will fail auth on first use). auth.chain: [] prints a loud open-relay warning.

The HTTP client uses a 300s request timeout and pools up to 1024 idle keep-alive connections per upstream host.

Validating configuration (busbar --validate)

Section titled “Validating configuration (busbar --validate)”

busbar --validate runs the exact load → resolve → validate pipeline the gateway runs at boot, then exits, without starting the server. It binds no port, writes nothing to disk, spawns no tasks, and makes no network call, so it is safe to run anywhere, including in CI and against a config edited on a live host before you reload it. It does read the files your file: secret references name, because since 1.5.3 it resolves those references rather than only checking their shape.

Terminal window
BUSBAR_CONFIG=./config.yaml BUSBAR_PROVIDERS=./providers.yaml busbar --validate
# ok: config valid — 2 provider(s), 2 model(s), 1 pool(s)
# note: 1 env var(s) referenced but unset here — required at runtime: BUSBAR_CLIENT_TOKEN

Two different mechanisms read the environment, and --validate treats them differently. Read both of the middle bullets before you wire this into CI.

  • Exit 0 = valid; 1 = errors (same diagnostics boot prints: invalid YAML, removed keys, dangling pool/lane references, malformed auth chains, cert-file and base_url/path SSRF violations). Use it as a CI gate: busbar --validate && deploy.
  • ${VAR} interpolation is lenient, and needs no value. A ${VAR} token anywhere in config.yaml that is unset in your shell is reported in a note: (“required at runtime”) rather than failing. (At real boot an unset ${VAR} is still a hard error.)
  • env: and file: secret REFERENCES are resolved, and one that cannot resolve fails the run. A secret reference is the { env: VAR } / { file: /path } form on providers.*.api_key, tls.cert / tls.key / tls.client_ca (and the admin_tls equivalents), auth.signing_key, and the admin-tokens token. Since 1.5.3 --validate reads each one and exits 1 naming the first that fails, where it previously checked only the shape and exited 0. So a CI job needs the same variables and files the deployment has, or a config whose references resolve in that environment. A reference served by a secret PLUGIN is not resolved here, since the plugin may not be loadable. Boot is unchanged: an unresolvable reference logs a warning and Busbar serves, with every request to that provider failing upstream. It checks structure and secret resolution, never upstream reachability.
  • Honors BUSBAR_CONFIG, BUSBAR_PROVIDERS, and --safe-mode exactly as boot does. Because it reuses the boot path, a clean --validate means a clean boot.

Inspecting the SSRF denylist (busbar --print-metadata-blocklist)

Section titled “Inspecting the SSRF denylist (busbar --print-metadata-blocklist)”

Provider base_url values are checked against a cloud-metadata denylist, so a compromised or mistyped config cannot turn Busbar into a reader of your instance credentials. The list the running binary actually enforces is the built-in set plus whatever you added under security.blocked_metadata_hosts, which means it is not something you can read off the config file alone. This flag prints it, one entry per line, and exits 0:

Terminal window
BUSBAR_CONFIG=./config.yaml busbar --print-metadata-blocklist
  • The built-in set always prints, so the flag works before a deployment is wired up.
  • Your security.blocked_metadata_hosts entries are appended when BUSBAR_CONFIG points at a config that reads and parses. If it does not, the flag prints the built-in set alone and says so on stderr rather than handing you a silently incomplete list. Run Busbar normally to see the parse error.
  • It does NOT subtract the allow-overrides. security.allow_metadata_hosts, a provider’s own allow_metadata_hosts, and allow_all_metadata still win at request time, so a host printed here can still be reachable if you unblocked it. See the configuration reference for how the two sides combine.

Busbar terminates TLS natively for the client↔Busbar hop. Add an optional tls block to config.yaml; when it is absent, Busbar serves plain HTTP exactly as before (no behavior change). When present, Busbar handles the TLS handshake itself, no sidecar required.

listen: "0.0.0.0:8443"
tls:
cert: { file: /etc/busbar/tls/fullchain.pem } # PEM cert chain, leaf first (secret reference)
key: { file: /etc/busbar/tls/privkey.pem } # PEM private key (PKCS#8 / PKCS#1 / SEC1)
# client_ca: { file: /etc/busbar/tls/ca.pem } # OPTIONAL: see "Mutual TLS" below

Each of cert, key, and client_ca is a secret reference, not a bare path: the { file: /path } form above reads PEM bytes from disk, and { env: VAR } reads them from an environment variable (or { module: <secret-plugin>, settings: {…} } from a secret backend). The plaintext cert_file/key_file/client_ca_file path keys of earlier releases are gone in 1.5.0.

Certificate & key formats. cert resolves to a PEM certificate chain with the leaf (server) certificate first, followed by any intermediates: exactly what most CAs ship as fullchain.pem. key resolves to the matching PEM private key in PKCS#8 (BEGIN PRIVATE KEY), PKCS#1 (BEGIN RSA PRIVATE KEY), or SEC1 (BEGIN EC PRIVATE KEY) encoding. Busbar advertises http/1.1 over ALPN.

Fail-fast. Any missing, unreadable, or unparseable cert/key/CA file stops the process at startup with a message naming the offending file: a misconfigured certificate can never silently downgrade or half-start the listener. Key bytes are never logged.

Set client_ca (a secret reference resolving to a PEM CA bundle) to require mutual TLS: every client must present a certificate that chains to that CA, or the TLS handshake is rejected before any request is processed. This is transport-level zero-trust: only holders of a cert your CA signed can establish a connection at all, with no service mesh or external proxy. It composes with (and runs before) the normal auth token / virtual-key check. A client with a missing or wrong certificate is dropped at handshake; the rejection is contained to that one connection and never affects the server or other clients.

Certs are loaded once at startup, so rotation always needs a restart — but it does not need a shell. Push the new cert/key/CA through the admin API, then restart in-product:

Terminal window
curl -X PUT http://localhost:8081/api/v1/admin/config/settings \
-H "x-admin-token: $ADMIN_TOKEN" -H 'content-type: application/json' \
--data '{"tls": {"cert": {"file": "..."}, "key": {"file": "..."}, "client_ca": {"file": "..."}}}'
# -> {"reload_to_apply": ["tls"], "note": "... takes effect on the next restart ..."}
curl -X POST http://localhost:8081/api/v1/admin/restart \
-H "x-admin-token: $ADMIN_TOKEN"

The PUT stores the new material durably (overlay-persisted) and reports tls under reload_to_apply — restart-scoped, per the PUT /config/settings table. POST /restart then applies it: it drains through the same graceful-shutdown path a signal takes (in-flight requests finish first), which is exactly why a restart on rotation is safe under live traffic — the same guarantee this section always relied on, now reachable without shelling in. If no process supervisor is detected, the endpoint refuses with 409 conflict unless the request sets confirm: true (an unsupervised exit would leave Busbar down).

Without admin API access (or without a config overlay configured), the file-level fallback still works: replace the PEM files on disk and restart Busbar directly (e.g. systemctl restart busbar).

Reverse proxy alternative. A TLS-terminating reverse proxy (nginx, Caddy, Envoy) in front of a plain-HTTP Busbar still works if you prefer to manage certs there: simply omit the tls block.

When Busbar terminates TLS itself, the native listener bounds the request header-read phase (30 s) in addition to the TLS handshake, so a client that completes the handshake and then trickles request headers one byte at a time cannot pin a connection open indefinitely. This bound applies only to reading the request headers: it never limits a streaming response, so long model completions are unaffected.

The plain-HTTP listener (no tls block) does not apply a header-read timeout. For an edge-facing deployment, either enable the tls block (recommended) or front Busbar with a reverse proxy / load balancer (nginx, Caddy, Envoy, an ALB), which terminates client connections and provides its own slow-client protection. A plain-HTTP Busbar directly exposed to untrusted networks is not recommended.

EndpointAuthMeaning
GET /healthzopen200 ok if any lane is usable; 503 otherwise. Use for liveness/readiness probes.
GET /metricsvirtual keyPrometheus exposition. OPT-IN: mounted only when an export: instance with module: prometheus is configured (with its required settings.buffer_seconds); otherwise the path 404s like any other. Requires a valid key with a non-empty auth.chain, open under chain: []. Restrict at the network layer if unauthenticated scraping is needed.
GET /statsvirtual keyPer-lane health snapshot + pool membership, JSON.

/stats returns, per lane: model, provider, max_concurrent, limit (alias of max_concurrent), inflight, free_slots, available (free permits for a bounded lane, or "unbounded"), at_capacity (true when a bounded lane is at its max_concurrent limit and is therefore shedding/spilling rather than queueing), availability, recovery_hint_ms, breaker_state, ok, err, usable, dead, dead_reason, cooldown_remaining_s, streak, and budget. It is the first place to look when a pool is degraded.

availability renders the shared classify taxonomy — the same one routing dispatches on, so /stats cannot drift from behaviour. It is "available" when the lane would admit a request, or the reason it can’t: "breaker_open", "at_capacity", "dead", "budget_exhausted", "probe_in_flight", or "shedding". recovery_hint_ms is the honest lower bound (ms) on when that lane could next serve (null when available or the reason has no self-recovery, e.g. dead/budget). The breaker (breaker_state: "closed"/"open"/ "half_open") and capacity (at_capacity) axes are exposed INDEPENDENTLY: a saturated Open lane shows breaker_state: "open" AND at_capacity: true — so you can see why such a lane’s breaker never recovers (its recovery probe needs a dispatch it can never win), rather than the signal being collapsed into one string.

Busbar is stateless (apart from governance ledgers, see below), so the robust production shape is N instances behind a load balancer, each configured identically, each health-checked on GET /healthz. Any instance serves any request; lose one and the LB routes around it. On Kubernetes this is replicaCount + the Service/Ingress + a PodDisruptionBudget; on VMs it is N hosts behind an external LB (nginx, HAProxy, or a cloud L4/L7 balancer) probing /healthz.

Three things are worth understanding before you scale out:

  • Circuit-breaker and lane health are per-instance. Each instance learns upstream health independently from its own traffic. This is correct (a lane that’s dead for one instance is usually dead for all) and a new instance re-learns within seconds. Nothing is shared or needs sharing.
  • Session affinity is per-instance. The affinity header pins a session to a lane within one instance. Across instances, an LB that spreads a client’s requests will spread its affinity too. If you depend on affinity, enable sticky sessions at the LB (e.g. by the affinity header / a cookie) so a session lands on the same instance.
  • Governance state defaults to per-instance memory; enforcement is per-node either way. The default store: memory is ephemeral RAM per instance. A cluster-shared store (postgres/valkey) genuinely shares keys and the token ledger across N nodes (but NOT the durable audit log - see below), and each node’s write-behind flush ships ADDITIVE per-(model, tier) token deltas so the store converges on the true fleet totals - but the budget hard cap is still checked from each node’s in-memory counters, so between flushes N nodes splitting traffic can admit up to ~N times a configured cap. For a strict single ceiling, run a single instance (scale vertically); the proxy path itself scales horizontally without this caveat.
  • The durable audit log takes exactly ONE writer. Audit sequence numbers are allocated in-process, so two nodes sharing a store reach for the same numbers and overwrite each other’s entries, breaking the hash chain the next boot verifies. A node that detects another writer logs an error and detaches its durable sink, continuing to audit to its in-memory ring (ephemeral) rather than corrupting the shared log. Point at most one node at a durable audit store; GET /audit is per-instance either way (it serves that node’s in-memory ring, never the store).
  • The signing key (auth.signing_key) must be the SAME secret on every node. It is fleet-shared: every node verifying the same virtual-key tokens must resolve the same ed25519 signing key. A token minted on one instance fails verification on another if they disagree, so point every node’s auth.signing_key at the identical secret reference (never let each instance generate or resolve its own).

So: for a gateway without group limits, scale out freely behind an LB. With limits, either accept the per-node cap semantics over a shared store, or keep enforcement on one instance and scale the box, not the count.

All metrics are Prometheus counters/histograms exposed at /metrics, which is opt-in: with no module: prometheus instance under export: busbar records nothing and does not mount the endpoint. Its settings.buffer_seconds (required when you opt in) sets how many seconds of observations are retained — quantiles cover that window, _sum/_count stay cumulative, and memory is bounded by the window rather than by uptime.

MetricTypeLabelsWatch for
busbar_requests_totalcounteringress_protocol, pool, outcomeoutcome is ok / client_error / exhausted (503) / error. A rising exhausted means pools are running out of healthy members.
busbar_upstream_attempts_totalcounterpool, laneReal upstream calls (re-counted per failover hop).
busbar_upstream_failures_totalcounterpool, lane, dispositiondisposition is transient_upstream / attempt_timeout / hard_down / context_length. Concentration on one lane points at a sick backend.
busbar_breaker_trips_totalcounterpool, laneEach hard-down/trip. Spikes = a backend going down.
busbar_failovers_totalcounterpool, reasonreason is timeout / connect / transient_upstream / attempt_timeout / hard_down / context_length.
busbar_translations_totalcounterfrom, toCross-protocol translation hops.
busbar_request_duration_secondshistogramingress_protocol, poolEnd-to-end latency.
busbar_key_spend_centsgaugekey (+ mint labels)Per-virtual-key derived spend in cents (all-time attribution bucket; spend derives from the token ledger x the current rate card at scrape time).
busbar_key_tokens_totalgaugekey (+ mint labels)Tokens consumed by each virtual key (all-time attribution bucket).
busbar_bucket_spend_centsgaugebucket, group, windowDerived spend per (group, window) enforcement bucket (bucket = group:<name>@<window>).
busbar_bucket_budget_remaining_centsgaugebucket, group, windowBudget cap minus derived spend, only for buckets carrying a budget limit. Enables Prometheus burn-rate alerting per group.
busbar_bucket_tokensgaugebucket, group, window, model, tierPer-(bucket, model, tier) token counters (the raw material for external cost dashboards).
busbar_lane_stategaugepool, laneCircuit-breaker health per lane (the independent breaker axis): 0 = Closed, 1 = HalfOpen, 2 = Open (tripped). Side-effect-free at scrape.
busbar_lane_availablegaugepool, laneUnified availability from the shared classify taxonomy (the same one routing dispatches on): 1 = the lane would admit a request right now, 0 = unavailable for ANY reason (breaker Open, at-capacity, dead, budget, probe-in-flight). Pair with busbar_lane_state (breaker) and busbar_lane_available_permits (capacity) to see which axis is the cause. Replaces the former busbar_lane_at_capacity. Side-effect-free.
busbar_lane_recovery_hint_msgaugepool, laneHonest lower bound (ms) on when an unavailable lane could next serve, from the same recovery_hint_ms that feeds Retry-After: breaker until for an Open lane, the at-capacity floor (2000ms) for a saturated one. 0 when available or the reason has no self-recovery (dead/budget). Side-effect-free.
busbar_lane_inflightgaugepool, laneIn-flight requests (held concurrency permits) per lane — the depth companion to busbar_lane_available. Side-effect-free.
busbar_lane_available_permitsgaugepool, laneFree concurrency permits for a bounded lane (0 = saturated) — the independent capacity axis. Unbounded lanes emit no sample. Side-effect-free.
busbar_pool_queuedgaugepoolRequests currently parked in the on_exhausted: queue bounded wait, per pool. Reads 0 until the queue policy is wired. Side-effect-free.
busbar_route_policy_selections_totalcounterpool, policyRequests where a selection strategy (a native strategy or a gate hook) produced a usable ranked order. Only incremented on a successful Order outcome; abstains and on-error fallbacks are not counted.
busbar_route_policy_rejections_totalcounterpool, policy, statusRequests deliberately rejected by a routing hook’s reject verb (a 4xx to the caller, no upstream dispatched). A guardrail saying no, not a failure.
busbar_webhook_logs_dropped_totalcountern/aRequest-log webhook deliveries shed because the in-flight delivery pool was saturated (a slow/unreachable webhook endpoint). A non-zero rate means request logs are being silently dropped, scale the endpoint or alert.
busbar_file_logs_dropped_totalcountern/aRequest-log file appends shed because that sink’s in-flight append pool was saturated (a slow/stalled filesystem — full disk, hung NFS/EBS mount). A non-zero rate means request-log lines are being dropped, check the mount or alert.
busbar_billing_truncated_totalcountern/aSame-protocol non-stream responses whose body exceeded the translate-body cap, so the terminal usage frame was missed and the request billed zero tokens (the client still got a full response). A non-zero rate signals an over-cap billing gap.

/metrics requires a valid key with a non-empty auth.chain, it is treated as an information-disclosure surface and goes through the same auth check as other routes. Only chain: [] admits scrapes unconditionally. Restrict it at the network layer (firewall, reverse proxy) if you need unauthenticated scraping under your threat model.

The breaker decides health from real request outcomes (passive), with optional active probing layered on top. The disposition pipeline (see architecture.md) decides whether an outcome counts as an upstream fault; this section covers what happens to the lane once it does.

Breaker state is per-(pool, lane): a lane that is a member of more than one pool carries independent Open/Closed/HalfOpen state, streak, cooldown, and error window in each pool, so one pool’s traffic tripping a lane does not bench it for the others. Direct/ad-hoc routes (POST /{provider}/{model}, POST /{model}) and /stats share a single lane-default cell. The concurrency limit and the max_requests lifetime budget are not per-pool, they cap the shared upstream, so they apply across every pool. A successful active health probe (it tests the shared upstream) clears the breaker in every cell for the lane.

probe succeeds → back to Closed probe fails (longer cooldown) trip condition cooldown expires Closed Open HalfOpen

single sub-threshold failure → brief skip, stays Closed

  • Closed: the lane serves traffic. A single upstream failure that does not meet the trip condition still arms a short cooldown (the lane is briefly skipped), but the breaker stays Closed.
  • Open: the lane is tripped and skipped during selection until its cooldown expires.
  • HalfOpen: on cooldown expiry, the next selection attempt transitions the lane to HalfOpen and admits exactly one probe request (single-flight via CAS). A successful probe completes recovery to Closed (streak/error window cleared); a failed probe reopens the lane with an escalated cooldown.

Configured per pool via breaker.trip (see configuration.md):

  • error_rate (default): trips when the failure fraction over window_secs reaches threshold (default 0.5), but never before min_requests (default 5) outcomes have accrued in the window.
  • consecutive: trips on consecutive_n consecutive failures (default 3).

Cooldown grows exponentially with the consecutive failure streak, doubling from base_cooldown_secs up to max_cooldown_secs, with ±10% jitter once the streak is nonzero. A server Retry-After header is always honored as a floor: even if it exceeds max_cooldown_secs. Defaults (no breaker: block): base 15s, max 120s.

  • A transient fault (5xx/timeout/network/overload/rate-limit) drives the trip evaluation and, on trip, opens the breaker: recoverable via the half-open probe.
  • A hard-down fault (billing/quota or auth) opens the breaker immediately with a long sticky cooldown (30 min) rather than waiting for a trip threshold, it does not set a permanent dead flag, so it is still recoverable: a successful active probe (or organic half-open probe) brings it back. An auth hard-down also relays the error to the caller; a billing hard-down fails the request over to another member.

Passive health alone only learns a lane is sick when real traffic hits it, and only recovers it on the next organic request. Active probing (per-provider health: config) adds a background prober:

ModeBehavior
none (default)No probing; pure passive health.
deadPeriodically re-probe only tripped lanes, so a recovered upstream is picked back up promptly.
activePeriodically probe every lane, so a silently-dead upstream trips out before real traffic hits it. Sends a tiny billable one-token request per interval.

Each probing lane gets one background task. interval_secs (default 30) and timeout_secs (default 5) are honored (floored at 1s). The first tick is skipped so Busbar doesn’t probe before any traffic establishes health. A lane with no key is skipped (a guaranteed 401 would only thrash the breaker). A 2xx probe recovers a tripped lane to Closed and increments the lane’s ok counter by one; a failed probe records a transient (which, on a Closed lane in active mode, can trip it out).

For a single request, Busbar will retry across pool members up to the failover max_hops (default 3) and within the timeout_secs budget (default 120). Failover is allowed only before the first upstream byte reaches the client: once streaming has started, a failure cannot fail over (the client holds a partial response); the lane records the breaker fault and the stream terminates with an SSE error event, and the client must retry.

When all members are unusable, the pool’s on_exhausted action decides:

  • reject / status_503 (default): 503 with Retry-After — the soonest genuine member cooldown, or a small saturation floor when exhaustion is pure at-capacity (not the misleading 1).
  • least_bad, serve the soonest-cooldown member that still has a free permit (skipping a saturated one), degraded and logged loudly.
  • { fallback_pool: <name> }, route to another pool (loop-guarded).

If outcome="exhausted" (503) is climbing in busbar_requests_total, check /stats for dead/tripped lanes and consider a fallback_pool or least_bad policy for graceful degradation.

Data-plane callers authenticate with signed, expiring virtual keys (the built-in keys verifier in auth.chain). Keys are managed over the admin API on the separate admin_listen, guarded by auth.admin_auth (the built-in admin-tokens operator credential, sent as Authorization: Bearer <admin_token> or X-Admin-Token: <admin_token>, or an IdP role with admin_scope).

Minting, listing, rotating, and revoking keys — the routes, request/response shapes, the mint-body field reference, and the scope lattice — are owned by the Admin API reference. The limit/group model those keys charge through is owned by Configuration → Virtual keys and enforcement. This guide stays on the operational picture. In brief: POST /api/v1/admin/keys mints a key and returns the signed token once; a key is pure auth (every limit lives on the bound group), it EXPIRES (default 90 days — re-mint or rotate before then), and DELETE puts its subject on the durable revocation denylist immediately.

  • Verification is stateless: signature + expiry + the revocation denylist; policy (group, pools) resolves from the store by the token’s subject, so a PATCH takes effect without re-issuing the credential.
  • Admission walks the bound group’s chain and ANDs every limit of every group: requests (precise, 429 + Retry-After), tokens (best-effort post-paid, 429 + Retry-After), budget (derived spend, the vendor-native quota status with error.type: insufficient_quota; Bedrock signals over-budget as 400), concurrent (in-flight gauge, 429). The rejection names the exact blocking bucket (group + metric + window). A frozen group (enabled: false) rejects with 403.
  • Spend derives from the TOKEN LEDGER: a flat per_request_fee is charged (as +1 request) atomically pre-forward, and the response’s per-(model, tier) token split is ledgered at stream end. Spend = requests x fee + tokens x rate_card rates, recomputed on every check; with no rate card, tokens price at 0 and only the flat fee counts.
  • Ledgers default to in-memory (ephemeral); configure a durable store plugin (store: { module: sqlite|postgres|valkey, settings: {...} }) to persist keys, usage, and the denylist across restarts.

Limit windows are per-process, and the caps are enforced per node even over a shared store (see the fleet caveat above).

SymptomWhere to look
503 on every request/stats, are all lanes dead or in cooldown? Check dead_reason.
A lane stuck dead with billing reasonUpstream wallet/quota; the lane recovers on a successful probe once funded. Consider health.mode: dead.
A lane stuck dead with auth reasonWrong/expired credential behind the provider’s api_key reference.
A few 401s from a Vertex AI or Azure (Entra ID) lane right after startupThe lane’s first OAuth token is still minting. jwt-bearer / oauth-client-credentials lanes fetch an access token in the background at boot (and on every reload); for up to ~1s before it lands, the earliest calls return 401. Clears itself within a second, no action needed. Static-key lanes (bearer / api-key / SigV4) never have this window.
429 from Busbar itselfA group limit blocked. The body’s error.type distinguishes the cause: rate_limit_error = requests/tokens/concurrent limit (the message names group + metric + window); insufficient_quota = a budget limit (Bedrock ingress signals over-budget as 400 instead). Check GET /api/v1/admin/keys/{id}/usage.
403 from BusbarThe virtual key’s allowed_pools doesn’t include the target.
Startup panic: “unset environment variable”A ${VAR} (possibly in a comment) isn’t exported.
Startup panic: “not found in providers.yaml”A config.yaml provider name isn’t in the catalog.
Cross-protocol responses missing fieldsExpected, only the modeled IR subset survives a cross-protocol hop; same-protocol routes are lossless.
High busbar_failovers_total for one laneThat backend is flapping; inspect its busbar_upstream_failures_total disposition.