The post defines an operability stack for production LLM systems built around a single gateway chokepoint that carries a decision_id through metrics, logs, traces and the ledger. Observability adds a decision layer to classic RED metrics: agent_decisions_total (outcome labels), agent_auto_execution_ratio gauge, agent_decision_confidence histogram, agent_decision_duration_seconds histogram, guardrail_blocks_total (layer/rule), judge_invocations_total and judge_disagreements_total counters, llm_tokens_total (direction/model), llm_cost_usd_total (model), llm_call_duration_seconds histogram, shadow_eval_pass_ratio gauge and human_override_rate gauge. Structured logging is split into three tiers (operational, decision‑metadata, regulated/raw) with the decision_id enabling end‑to‑end traceability. Example alerts trigger on a >15% swing in auto_execution_ratio versus a 7‑day baseline, guardrail block rate >3× trailing hour, judge disagreement ratio >0.15, hourly cost increase >budget/24, shadow‑eval pass ratio 8 s for 10 min. SLOs target decision availability ≥99.9%, quality within baseline − 2%, latency p95 under the interaction budget, and cost per 1k decisions within ±20% of plan. Cost control routes every LLM call through the gateway, which estimates tokens, applies per‑tenant token‑per‑minute and daily USD limits (e.g., 200 k tokens/min, $50/day), records actual spend, and surfaces levers: skip the model for deterministic cases, batch, cache, right‑size models, and trim prompts. Routing selects cheap, capable, judge or long‑context models based on decision properties, with fallback chains, circuit‑breaker checks, per‑attempt timeouts and idempotency keys. A runtime kill switch offers LIVE, HUMAN_ONLY and HALTED modes, propagated via shared state with granular per‑tenant/capability scoping and a fail‑to‑safe default. The platform relies on hexagonal architecture (ports/adapters), delegated short‑lived identity grants with tenant‑matching checks, least‑privilege IAM, managed secrets, an append‑only audit datastore, Redis cache for rate‑limit/kill‑switch state, and encrypted object storage for logs and artifacts. Together these pieces make cost observable, routing controllable, failures isolatable, and the system operable at scale.
Read original
stackoverflow/blog