LLM Deployment: CI/CD, staging and one-step rollback
How to ship LLM changes safely - CI with eval gates, shadow mode, a sticky canary with an automatic verdict, config-driven rollback and provider fallback that does not make things worse. With runnable code.
An LLM change - a new prompt, a model swap, a retrieval tweak - is a deploy like any other, and deserves the same safety net: tested in CI, rolled out progressively, reversible in one step. Skipping that is how a one-line prompt edit becomes a production incident.
The two questions
Prototype: “It runs on my machine.” Production: “Can we ship this change safely - and undo it in seconds if it regresses?”
A demo is deployed once, by hand. Production needs every change to flow through a pipeline that catches regressions before users do and can revert without a firefight.
What breaks without it
- Prompt edits straight to prod. Someone tweaks the live prompt to fix one complaint and silently regresses ten others - and with no version history, nobody can tell which edit did it.
- No staging. Changes are validated in production, on real users.
- No rollback. A bad release means a frantic redeploy of application code instead of flipping one value.
- No fallback. Your sole provider has an outage and the whole feature is down.
- Big-bang releases. A change ships to 100% of traffic at once; there’s no small blast radius to learn from.
The controls that matter
- CI with an eval gate - every change runs the eval set and can’t merge on a regression.
- A staging environment that mirrors production (same retrieval, same config).
- Progressive rollout - canary or percentage-based, so a bad change hits few users.
- One-step rollback - driven by config, not a code redeploy.
- Provider / model fallback - automatic failover when a provider errors or times out.
- Change management - for regulated contexts, a documented, auditable process.
Instrument it: CI with an eval gate
Wire evals into the pipeline so quality is a merge condition, not a hope:
# .github/workflows/deploy.yml (sketch)
on:
pull_request:
paths: ['prompts/**', 'src/**', 'evals/**']
jobs:
evaluate:
steps:
- run: python run_evals.py --data evals/
# fail if pass_rate drops below threshold or a "critical" case regresses
deploy-canary:
needs: evaluate
if: github.ref == 'refs/heads/main'
steps:
- run: ./deploy.sh --canary 5 # 5% of traffic first
Instrument it: rollback by config, not redeploy
Load the active prompt and model from configuration so a revert is a value change, not a deploy:
# config (env var, feature flag, or small config service)
ACTIVE = {"prompt_version": "rag-answer-v4", "model": "claude-sonnet-5"}
# rollback = point back at the previous known-good version, instantly
# ACTIVE = {"prompt_version": "rag-answer-v3", "model": "claude-sonnet-5"}
Pair this with a canary check: route a slice of traffic to the new version, watch latency, error rate and evals, and promote or roll back automatically. The next three sections make that concrete.
Shadow first, then canary
An eval set tests the cases you thought of. Real traffic tests the rest. For a risky change - a new model, a rewritten prompt, a different retriever - run it in shadow mode before any user sees it: send a copy of live requests to the candidate, serve the stable answer as usual, and store both for comparison.
- Compare them offline with your judge questions and literal checks, plus latency and tokens. Disagreements are the interesting rows - read a sample by hand.
- Shadowing doubles model spend for the shadowed slice, so shadow a sample, not everything, and for days rather than weeks.
- Never shadow anything with side effects. A candidate agent must run with its write tools stubbed out.
- Shadow copies are still customer data: same redaction, access and retention rules as production traces.
Once the shadow comparison is clean, move to a canary that real users see.
Instrument it: a sticky percentage rollout
Assign users to the canary by hashing a stable id, not by random choice per request - otherwise one person bounces between two behaviours within a conversation. Salt the hash per rollout so each release samples a different 5%. The same file also decides the canary automatically. This runs as-is:
import hashlib
def bucket(user_id: str, salt: str) -> float:
"""Stable position in [0, 100) - the same user always lands in the same place."""
h = hashlib.sha256(f"{salt}:{user_id}".encode()).digest()
return int.from_bytes(h[:8], "big") / 2**64 * 100
def variant(user_id: str, rollout: dict) -> str:
return rollout["candidate"] if bucket(user_id, rollout["salt"]) < rollout["percent"] else rollout["stable"]
ROLLOUT = {"salt": "rag-answer-v5", "percent": 5,
"stable": "rag-answer-v4", "candidate": "rag-answer-v5"}
users = [f"user-{i}" for i in range(100_000)]
share = sum(variant(u, ROLLOUT) == "rag-answer-v5" for u in users) / len(users)
print(f"share on candidate: {share:.2%}")
print("user-42 is sticky:", {variant("user-42", ROLLOUT) for _ in range(5)})
def canary_verdict(stable: dict, canary: dict) -> str:
"""Compare the canary with the stable slice over the same time window."""
if canary["requests"] < 2_000:
return "wait: not enough canary traffic yet"
checks = {
"error rate": canary["error_rate"] <= stable["error_rate"] * 1.2 + 0.001,
"p95 latency": canary["p95_ms"] <= stable["p95_ms"] * 1.15,
"eval pass rate": canary["eval_pass"] >= stable["eval_pass"] - 0.02,
"cost/request": canary["cost"] <= stable["cost"] * 1.10,
}
failed = [name for name, ok in checks.items() if not ok]
return "promote" if not failed else "roll back: " + ", ".join(failed)
stable = {"requests": 90_000, "error_rate": 0.004, "p95_ms": 2400, "eval_pass": 0.91, "cost": 0.0130}
print(canary_verdict(stable, {"requests": 4_800, "error_rate": 0.004, "p95_ms": 2550, "eval_pass": 0.92, "cost": 0.0126}))
print(canary_verdict(stable, {"requests": 4_800, "error_rate": 0.005, "p95_ms": 2500, "eval_pass": 0.86, "cost": 0.0151}))
Its actual output:
share on candidate: 4.91%
user-42 is sticky: {'rag-answer-v4'}
promote
roll back: eval pass rate, cost/request
The second canary is the instructive one: its error rate and latency are fine, so a classic service canary would have promoted it. It fails on the two numbers only an LLM deploy needs - quality on sampled live traffic, and cost per request. The thresholds here are illustrative; set yours from how much each metric normally wobbles day to day, and compare canary and stable over the same window so time-of-day effects cancel out.
Fallbacks that don’t make things worse
Provider fallback sounds simple and goes wrong in predictable ways - retry storms, a fallback model nobody evaluated, 400 errors “fixed” by quietly switching providers. A sketch of the shape that avoids them:
RETRYABLE = {429, 500, 502, 503, 504}
def complete(request, chain=("primary", "secondary"), attempts=2, timeout_s=20):
for target in chain: # e.g. another region first, then another provider
for attempt in range(attempts):
try:
return call(target, request, timeout=timeout_s)
except ProviderError as e:
if e.status not in RETRYABLE:
raise # a 400 is your bug; retrying elsewhere only hides it
sleep(min(2 ** attempt, 8) + random.random()) # backoff with jitter
return degraded_response(request) # canned answer or human handoff, logged as degraded
Three rules go with it. Every model in the chain passes the same eval set, with its own prompt variant if it needs one. Fallback events are logged and alerted on, because “we’ve been on the secondary for six hours” is an incident too. And when a provider is clearly down, a circuit breaker skips it for a while instead of paying the timeout on every request.
A release checklist
Before a prompt, model or retrieval change goes to 100%:
- The eval set passed in CI, with no critical case regressed.
- Token count and cost per request were compared with the current version.
- Shadow comparison reviewed (for model or retriever changes).
- Canary ran long enough to cover the daily traffic shape, and the verdict was “promote”.
- The previous version is still deployable - the rollback target exists.
- The change, its owner and the eval result are recorded in the change log.
Minimal vs mature
| Aspect | Minimal | Production-grade |
|---|---|---|
| Testing | Manual check | Eval gate in CI |
| Environments | Prod only | Staging mirrors prod; shadow for risky changes |
| Rollout | All at once | Sticky canary with an automatic verdict |
| Rollback | Code redeploy | One-step, config-driven |
| Resilience | Single provider | Evaluated fallback chain, circuit breaker |
Where this lives in a real system
Provider failover and routing belong in a control point - see the multi-provider gateway reference architecture. Rollback depends on treating prompts as versioned artifacts, covered in prompt versioning with GitHub, and the same pipeline carries you through a vendor retiring your model - see the migration playbook. And the deployment items in the Production Checklist are the bar to clear before you ship.