What changed, and when.
A reference is only as good as its maintenance. This is the dated record of every meaningful update - pricing re-verifications included, since those happen on a schedule, not on inspiration.
-
New: the breakthroughs behind LLMOps
A companion to the history: the papers and launches from 2017 to 2024 that made models capable (the Transformer, scaling laws, in-context learning, chain-of-thought, instruction tuning, Chinchilla, reasoning models), cheaper to run (LoRA and QLoRA, FlashAttention, quantization, speculative decoding, PagedAttention, mixture of experts, prompt caching), open (LLaMA, Llama 2) and connected (RAG, ReAct, function calling, Structured Outputs, MCP) - and the operational work each one created. Every figure is taken from the paper's abstract or the launch post itself.
-
Practical depth for the layer guides
The guides for cost control, observability, security, deployment, prompt management, RAG operations, governance and evaluation roughly doubled in length, almost all of it hands-on. New: a worked cost model for a support assistant (caching, routing and trimming take a $10,140 monthly bill to $4,022), a runnable OpenTelemetry trace, dashboard and alert tables, output sanitising against Markdown-image exfiltration, a blast-radius map and a red-team set, shadow mode, a sticky canary with an automatic verdict and a fallback chain, a strict prompt loader, hybrid search with rank fusion and citation checks, risk tiers, an evidence pack and a retention schedule, repeated eval runs and swap-order pairwise judging, and four incident playbooks. Eight examples say they run as-is; each was executed and its published output is the real output. Also fixed: on phones, wide tables and long inline code made article pages wider than the screen, so browsers zoomed the whole page out - tables now scroll inside the text column, and all 42 pages were checked at 375 pixels; four sentences that rendered as stray bullet points; one code sample that was not valid Python, and the migration playbook's tokenizer figure, which now quotes Anthropic's pricing page directly (approximately 30% more tokens).
-
New: a short history of LLMOps, 2015-2026
How the discipline took shape, from the 2015 paper on hidden technical debt in machine learning systems, through models becoming APIs and the 2023 production scramble, to standards, the EU AI Act and agents - with a timeline in which every milestone is dated, linked to its primary source and tagged with the stack layer it gave rise to. It deliberately does not name a coinage date for the word "LLMOps": there is no reliable record of one.
-
Every article re-read and corrected
A full quality pass over all 22 guides. Reading times are now computed from the text: the hand-set values overstated most articles two- to three-fold. The evaluation guide and all five reference architectures still used must_not_include checks of the kind this changelog retracted on September 11 - phrases such as "invented exceptions" that no model ever writes, so the check passed the answers it existed to catch. They now separate literal checks from judge questions, and every regex was tested against right and wrong answers. The hallucination-rate guide gains confidence intervals, with output from a script: a rise from 3.0% to 5.5% is noise at 200 samples (p = 0.215) and a real regression at 1,000 (p = 0.006). The security guide now uses the OWASP Top 10 for LLM Applications 2025 names and adds system prompt leakage and unbounded consumption. The OpenTelemetry example uses gen_ai.provider.name, which replaced gen_ai.system, and the conventions' source link now points to their new repository. The governance guide adds the EU AI Act's record-keeping and human-oversight articles, with the high-risk dates as moved by the AI Omnibus (December 2027 and August 2028). The cost guide adds batch APIs (50% off at OpenAI and Anthropic) and uses per-model cache rates instead of a flat 10%. RAGOps now distinguishes hit rate from recall. Code samples no longer name retired models, and three unsourced "most common cause" claims were removed.
-
Tools directory: two removals, six additions, ownership changes flagged
Removed Humanloop, which joined Anthropic and sunset its platform, and Protect AI, now part of Palo Alto Networks with its open-source LLM Guard repository archived. Added vLLM, LiteLLM, pgvector, DeepEval, NeMo Guardrails and MLflow, each checked against its own documentation today. Corrected Arize Phoenix: it is source-available under the Elastic License 2.0, not open source, though still free to self-host. Cards can now carry a sourced notice for things that are not attributes: Langfuse was acquired by ClickHouse (MIT core and self-hosting continue), Helicone joined Mintlify and runs in maintenance mode (so it no longer carries a best-for tag and has left the cost-tracking list), Lakera is now part of Check Point, Replicate announced it is joining Cloudflare, and DeepEval sends usage telemetry unless you opt out. Each tool's attributes now live in one place, so a correction reaches every category it appears in. The Langfuse vs LangSmith vs Braintrust vs Helicone guide and the prompt-versioning guide were updated to match.
-
Industry definitions: seven platforms, each re-checked
The homepage now shows the platforms that publish their own LLMOps page - Google Cloud, IBM, Microsoft, Red Hat, Databricks, MLflow and Oracle - each linking to the vendor's own page, and the What is LLMOps article paraphrases all seven. Every paraphrase was checked against the vendor's page today. One correction: Databricks does not publish a one-line definition, so it is now described as documenting LLMOps as its MLOps workflow adapted for LLMs rather than as "breaking it into" components. MLflow now links to its dedicated LLMOps page instead of its homepage.
-
Claude Opus 5.5, GPT-6 Sol/Luna and Grok 4.7 join the calculator
Swapped Claude Opus 5 for the newly launched Claude Opus 5.5, Anthropic's new recommended starting point, at a lower $4/$20 with cache hits at $0.20 (5% of input). Added OpenAI's GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50), replacing GPT-5.6 Terra, GPT-5.6 Luna and GPT-5.4 nano, which they match or undercut. Swapped Grok 4.6 for xAI's new flagship Grok 4.7 at the same $2/$6. Claude Fable 5.1, Sonnet 5, Haiku 4.5, GPT-6 Astra, Gemini 3.8 Flash and Gemini 3.1 Pro confirmed unchanged.
-
Tools directory: SOC 2 now says which offering is certified
SOC 2 was a yes/no flag, which cannot be answered honestly - certification covers a vendor's hosted service, not software you run yourself, so a listing marked both self-hostable and SOC 2 asserted something no vendor claims. Qdrant states it plainly: its compliance table marks self-hosted "not applicable". The badge now reads "SOC 2 · cloud" and never transfers to a self-hosted deployment. Listings checked this pass carry the date and the source page they were checked against; the rest are marked unverified and show no badge either way rather than carrying forward an unchecked flag. Also corrected: LangSmith and Braintrust both support OpenTelemetry and were listed as not supporting it. Each source is its own direct link to the verification page, separate from the link to the vendor.
-
Eval-dataset guide rewritten, with tests for its own checks
The guide offered `must_not_include: ["invented exceptions"]` as a machine-checkable assertion. It is not one - a model never emits that phrase, so the check passed on exactly the answers it existed to catch. A first fix earlier today made the opposite mistake: regexes on "45-day" and "just this once" that rejected correct refusals quoting those words. The guide now separates literal checks (what a regex can decide) from judge questions (anything that depends on meaning), and ships a runnable harness that tests every check against hand-labelled right and wrong answers, with its actual output. Against that labelled set, the original checks rejected two of three correct answers and passed three of four wrong ones. Separately, the homepage dashboard is now labelled illustrative, since its figures are an example of what to monitor rather than measurements from any service we operate.
-
New flagships from Anthropic, OpenAI and Google in the calculator
Swapped Claude Fable 5 for Claude Fable 5.1, Anthropic's new recommended top model, at the same $10/$50. Added OpenAI's new GPT-6 Astra flagship at $10/$50 and Google's Gemini 3.8 Flash at $1.50/$7.50 standard (an introductory $0.75/$3.75 runs through Dec 31, 2026). Removed GPT-5.6 Sol: OpenAI publishes only a promotional $4/$20 rate for it and has never posted a standard price, so listing one would have meant guessing. Opus 5, Sonnet 5, Haiku 4.5, GPT-5.6 Terra/Luna, GPT-5.4 nano, Gemini 3.1 Pro and Grok 4.6 all confirmed unchanged. The calculator also now uses each model's own verified cached-input rate, editable per calculation, instead of assuming 10% for everyone - Claude Fable 5.1 caches at 2.5% of its input price and Grok 4.6 at 25%.
-
Model prices re-verified across all four vendors
Claude Sonnet 5 dropped to its now-permanent $2/$10 rate (the planned Sep 1 increase to $3/$15 was cancelled). Swapped Grok 4.5 for the newly launched Grok 4.6, xAI's new recommended flagship, at the same $2/$6 rate. OpenAI and Google prices confirmed unchanged (GPT-5.6 Sol has a time-boxed promo at $4/$20 through at least Nov 21, 2026; kept its $5/$30 standard rate pending confirmation of the post-promo price).
-
Model prices re-verified across all four vendors
Swapped Claude Opus 4.8 for the newly launched Claude Opus 5 (same $5/$25 rate); cut GPT-5.6 Terra to $2/$12 and GPT-5.6 Luna to $0.20/$1.20 on updated OpenAI pricing. Google and xAI prices confirmed unchanged.
-
Model prices re-verified across all four vendors
Added Claude Sonnet 5, GPT-5.6 Sol/Terra/Luna and Grok 4.5; removed the GPT-4.1 family and Grok 4.1, which vendors no longer list.
-
Two new articles: model-retirement migrations and agent operations
An ops playbook for LLM model migrations (Opus 4.1 retires Aug 5), and a guide to running agents with budgets, gates and escalation.
-
Four practitioner how-tos published
Eval datasets, prompt versioning with GitHub, hallucination-rate measurement, and what to log in an LLM trace.
-
RAG freshness monitoring checklist
Field-by-field checklist for keeping a RAG index true, with an inline pipeline diagram.
-
Learn LLMOps page launched
Thirteen live-verified courses, programs, certifications, books and podcasts, mapped to stack layers - plus site search, RSS feed and llms.txt.
-
Claude Fable 5 added to the cost calculator
Verified at $10/$50 per MTok alongside a full pricing re-check.
-
Sources & citations system
Central register of live-verified references, cited on the homepage, stack page, glossary and every article. Layer colour system and stack diagrams shipped the same week.
-
LLMOps.si launched
The stack guide, production checklist, cost calculator, maturity score, glossary, tools directory, five reference architectures and the first ten articles.
Pricing entries are produced by a scheduled verification pass against vendor pricing pages. Spotted something stale anyway? Tell us.