The article argues that tracking only model versions is insufficient for reliable production AI and proposes a versioned release manifest that captures all components influencing behavior: application code, model snapshot, prompt, retrieval index, embedding model, preprocessing pipeline, and runtime settings such as token limits, batching, timeouts, and resource placement. Each manifest entry references immutable artifacts or configuration stores, with secrets stored only as references. An evaluation gate must assess the full end‑to‑end path—retrieval, generation, and output validation—using a versioned dataset that includes ordinary queries, observed failures, ambiguous requests, missing‑evidence cases, and authorization‑boundary attempts; acceptance criteria forbid any access‑control violation, require latency and cost budgets to hold under a representative workload, and demand that key task slices not regress beyond a defined tolerance. Load testing should vary input and output length distributions, concurrency, arrival bursts, warm and cold cache states, and measure time‑to‑first‑token, subsequent token pace, and queue depth, drawing on vLLM and Triton metrics. Rollback is implemented as a routing decision that preserves compatible dependency versions (e.g., index, caches) and drains in‑flight generations, with explicit stop conditions and ownership defined at component interfaces. The core practice is to treat the manifest, a tested evaluation suite, a representative load test, release‑aware tracing, and a rehearsed switch‑back as the minimal viable MLOps workflow before adding further platform complexity.
Read original
stackoverflow/blog