Writing

Articles & Guides

Deep dives on each layer of running language models in production - from first principles to the operational details.

Latest

Newest guides first.

Everything else, most recently published at the top.

LLMOps

A short history of LLMOps: from hidden technical debt to agents

How LLMOps took shape, 2015-2026: the MLOps warning about technical debt, models turning into APIs, the 2023 production scramble, standards and regulation, and the agent era. Every milestone dated and linked to its primary source.

Sep 27, 2026 · Read →
LLMOps

The breakthroughs behind LLMOps: what each one changed in production

The papers and launches, 2017-2024, that made large language models capable, cheap to run, open and connected - and the operational work each one created. Every result taken from its primary source.

Sep 27, 2026 · Read →
Deployment

Your model will be retired: an ops playbook for LLM migrations

Every model you build on has a deprecation calendar. A practical playbook for migrating LLM models without breaking production - inventory, pinning, eval gates, canary rollout and one-step rollback.

Jul 19, 2026 · Read →
Governance

Operating LLM agents in production: budgets, gates and escalation

An agent is an LLM with hands. An ops guide to running agents safely in production - least-privilege tools, approval gates for irreversible actions, step and cost budgets, per-step tracing and a real human escalation path.

Jul 19, 2026 · Read →
Evaluation

How to build your first eval dataset

A practical, step-by-step guide to building an LLM eval dataset from real traffic - what a row looks like, how to score it, how many cases you need, and how to wire it into CI.

Jun 20, 2026 · Read →
Observability

What to log in an LLM trace

A field-by-field guide to what belongs in a production LLM trace - request IDs, prompt versions, retrieval, tokens, latency, cost and outcome - plus what to redact.

Jun 19, 2026 · Read →
Evaluation

How to calculate hallucination rate

A practical method for measuring LLM hallucination (faithfulness) rate in production - how to define it, sample it, judge it, and track it over time.

Jun 18, 2026 · Read →
Prompt management

Prompt versioning with GitHub

A concrete workflow for versioning LLM prompts in GitHub - file layout, pull-request review, eval gating in CI, and one-step rollback - without buying a dedicated tool.

Jun 17, 2026 · Read →
RAG operations

RAG freshness monitoring checklist

A focused checklist for keeping a RAG index fresh in production - detecting stale content, missing documents, embedding drift and re-index failures before users do.

Jun 16, 2026 · Read →
Governance

LLM incident response template

A ready-to-adapt incident response template for LLM applications - severity levels, the first 15 minutes, mitigation levers unique to LLMs, and a post-mortem structure.

Jun 15, 2026 · Read →
Observability

How to choose between Langfuse, LangSmith, Braintrust and Helicone

A decision framework for picking an LLM observability and evals platform - how Langfuse, LangSmith, Braintrust and Helicone differ, and which fits your team.

Jun 14, 2026 · Read →
Governance

What your CTO should ask before approving an LLM launch

The questions a technical leader should ask before signing off on shipping an LLM feature to production - covering evals, observability, cost, security, rollback and governance.

Jun 13, 2026 · Read →
Governance

LLM Governance: Audit trails, approvals and explainability

A practical guide to governing LLM applications - audit trails, approval gates, risk tiers, an evidence pack, a retention schedule, and what the EU AI Act asks for - so you can answer for the system, not just run it.

Jun 8, 2026 · Read →
Deployment

LLM Deployment: CI/CD, staging and one-step rollback

How to ship LLM changes safely - CI with eval gates, shadow mode, a sticky canary with an automatic verdict, config-driven rollback and provider fallback that does not make things worse. With runnable code.

Jun 8, 2026 · Read →
Evaluation

LLM Evaluation: How to test prompts, RAG and agents

A production-grade guide to LLM evaluation - what breaks without it, what to measure, how to write an LLM-as-judge and an eval runner, and how to gate releases the way you gate code on tests.

Jun 1, 2026 · Read →
Prompt management

Prompt Versioning: Why prompts should be treated like code

A one-line prompt edit can move behaviour as much as a model swap. A production guide to treating prompts as versioned artifacts - change-tested, attributable and rollbackable in one step.

May 31, 2026 · Read →
RAG operations

RAGOps: How to monitor retrieval quality

Most "the model is wrong" bugs are retrieval bugs. A production guide to measuring retrieval quality, building a retrieval eval set, hybrid search with rank fusion, citation checks, chunking experiments and permission-aware retrieval.

May 30, 2026 · Read →
Cost control

LLM Cost Control: Tokens, caching and model routing

Token spend compounds quietly until a finance review forces a panic. A production guide to unit economics, a worked cost model, token audits, caching, routing, batching and budget enforcement - before the bill, not after.

May 29, 2026 · Read →
Security

LLM Security: Prompt injection, data leakage and guardrails

An LLM with tools and data access is an attack surface. A practical guide to prompt injection, data leakage and excessive agency - with guardrail code, output sanitising, a blast-radius map and a red-team set to run in CI.

May 28, 2026 · Read →
Deployment

LLMOps Checklist: From prototype to production

The narrative companion to the production-readiness checklist - the two questions that separate a demo from a system, walked across all eight layers of the LLMOps stack.

May 27, 2026 · Read →