Changelog

What changed, and when.

A reference is only as good as its maintenance. This is the dated record of every meaningful update - pricing re-verifications included, since those happen on a schedule, not on inspiration.

  1. Sep 27, 2026 Articles

    New: the breakthroughs behind LLMOps

    A companion to the history: the papers and launches from 2017 to 2024 that made models capable (the Transformer, scaling laws, in-context learning, chain-of-thought, instruction tuning, Chinchilla, reasoning models), cheaper to run (LoRA and QLoRA, FlashAttention, quantization, speculative decoding, PagedAttention, mixture of experts, prompt caching), open (LLaMA, Llama 2) and connected (RAG, ReAct, function calling, Structured Outputs, MCP) - and the operational work each one created. Every figure is taken from the paper's abstract or the launch post itself.

  2. Sep 27, 2026 Articles

    Practical depth for the layer guides

    The guides for cost control, observability, security, deployment, prompt management, RAG operations, governance and evaluation roughly doubled in length, almost all of it hands-on. New: a worked cost model for a support assistant (caching, routing and trimming take a $10,140 monthly bill to $4,022), a runnable OpenTelemetry trace, dashboard and alert tables, output sanitising against Markdown-image exfiltration, a blast-radius map and a red-team set, shadow mode, a sticky canary with an automatic verdict and a fallback chain, a strict prompt loader, hybrid search with rank fusion and citation checks, risk tiers, an evidence pack and a retention schedule, repeated eval runs and swap-order pairwise judging, and four incident playbooks. Eight examples say they run as-is; each was executed and its published output is the real output. Also fixed: on phones, wide tables and long inline code made article pages wider than the screen, so browsers zoomed the whole page out - tables now scroll inside the text column, and all 42 pages were checked at 375 pixels; four sentences that rendered as stray bullet points; one code sample that was not valid Python, and the migration playbook's tokenizer figure, which now quotes Anthropic's pricing page directly (approximately 30% more tokens).

  3. Sep 27, 2026 Articles

    New: a short history of LLMOps, 2015-2026

    How the discipline took shape, from the 2015 paper on hidden technical debt in machine learning systems, through models becoming APIs and the 2023 production scramble, to standards, the EU AI Act and agents - with a timeline in which every milestone is dated, linked to its primary source and tagged with the stack layer it gave rise to. It deliberately does not name a coinage date for the word "LLMOps": there is no reliable record of one.

  4. Sep 27, 2026 Articles

    Every article re-read and corrected

    A full quality pass over all 22 guides. Reading times are now computed from the text: the hand-set values overstated most articles two- to three-fold. The evaluation guide and all five reference architectures still used must_not_include checks of the kind this changelog retracted on September 11 - phrases such as "invented exceptions" that no model ever writes, so the check passed the answers it existed to catch. They now separate literal checks from judge questions, and every regex was tested against right and wrong answers. The hallucination-rate guide gains confidence intervals, with output from a script: a rise from 3.0% to 5.5% is noise at 200 samples (p = 0.215) and a real regression at 1,000 (p = 0.006). The security guide now uses the OWASP Top 10 for LLM Applications 2025 names and adds system prompt leakage and unbounded consumption. The OpenTelemetry example uses gen_ai.provider.name, which replaced gen_ai.system, and the conventions' source link now points to their new repository. The governance guide adds the EU AI Act's record-keeping and human-oversight articles, with the high-risk dates as moved by the AI Omnibus (December 2027 and August 2028). The cost guide adds batch APIs (50% off at OpenAI and Anthropic) and uses per-model cache rates instead of a flat 10%. RAGOps now distinguishes hit rate from recall. Code samples no longer name retired models, and three unsourced "most common cause" claims were removed.

  5. Sep 27, 2026 Data

    Tools directory: two removals, six additions, ownership changes flagged

    Removed Humanloop, which joined Anthropic and sunset its platform, and Protect AI, now part of Palo Alto Networks with its open-source LLM Guard repository archived. Added vLLM, LiteLLM, pgvector, DeepEval, NeMo Guardrails and MLflow, each checked against its own documentation today. Corrected Arize Phoenix: it is source-available under the Elastic License 2.0, not open source, though still free to self-host. Cards can now carry a sourced notice for things that are not attributes: Langfuse was acquired by ClickHouse (MIT core and self-hosting continue), Helicone joined Mintlify and runs in maintenance mode (so it no longer carries a best-for tag and has left the cost-tracking list), Lakera is now part of Check Point, Replicate announced it is joining Cloudflare, and DeepEval sends usage telemetry unless you opt out. Each tool's attributes now live in one place, so a correction reaches every category it appears in. The Langfuse vs LangSmith vs Braintrust vs Helicone guide and the prompt-versioning guide were updated to match.

  6. Sep 27, 2026 Data

    Industry definitions: seven platforms, each re-checked

    The homepage now shows the platforms that publish their own LLMOps page - Google Cloud, IBM, Microsoft, Red Hat, Databricks, MLflow and Oracle - each linking to the vendor's own page, and the What is LLMOps article paraphrases all seven. Every paraphrase was checked against the vendor's page today. One correction: Databricks does not publish a one-line definition, so it is now described as documenting LLMOps as its MLOps workflow adapted for LLMs rather than as "breaking it into" components. MLflow now links to its dedicated LLMOps page instead of its homepage.

  7. Sep 27, 2026 Pricing

    Claude Opus 5.5, GPT-6 Sol/Luna and Grok 4.7 join the calculator

    Swapped Claude Opus 5 for the newly launched Claude Opus 5.5, Anthropic's new recommended starting point, at a lower $4/$20 with cache hits at $0.20 (5% of input). Added OpenAI's GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50), replacing GPT-5.6 Terra, GPT-5.6 Luna and GPT-5.4 nano, which they match or undercut. Swapped Grok 4.6 for xAI's new flagship Grok 4.7 at the same $2/$6. Claude Fable 5.1, Sonnet 5, Haiku 4.5, GPT-6 Astra, Gemini 3.8 Flash and Gemini 3.1 Pro confirmed unchanged.

  8. Sep 11, 2026 Data

    Tools directory: SOC 2 now says which offering is certified

    SOC 2 was a yes/no flag, which cannot be answered honestly - certification covers a vendor's hosted service, not software you run yourself, so a listing marked both self-hostable and SOC 2 asserted something no vendor claims. Qdrant states it plainly: its compliance table marks self-hosted "not applicable". The badge now reads "SOC 2 · cloud" and never transfers to a self-hosted deployment. Listings checked this pass carry the date and the source page they were checked against; the rest are marked unverified and show no badge either way rather than carrying forward an unchecked flag. Also corrected: LangSmith and Braintrust both support OpenTelemetry and were listed as not supporting it. Each source is its own direct link to the verification page, separate from the link to the vendor.

  9. Sep 11, 2026 Articles

    Eval-dataset guide rewritten, with tests for its own checks

    The guide offered `must_not_include: ["invented exceptions"]` as a machine-checkable assertion. It is not one - a model never emits that phrase, so the check passed on exactly the answers it existed to catch. A first fix earlier today made the opposite mistake: regexes on "45-day" and "just this once" that rejected correct refusals quoting those words. The guide now separates literal checks (what a regex can decide) from judge questions (anything that depends on meaning), and ships a runnable harness that tests every check against hand-labelled right and wrong answers, with its actual output. Against that labelled set, the original checks rejected two of three correct answers and passed three of four wrong ones. Separately, the homepage dashboard is now labelled illustrative, since its figures are an example of what to monitor rather than measurements from any service we operate.

  10. Sep 11, 2026 Pricing

    New flagships from Anthropic, OpenAI and Google in the calculator

    Swapped Claude Fable 5 for Claude Fable 5.1, Anthropic's new recommended top model, at the same $10/$50. Added OpenAI's new GPT-6 Astra flagship at $10/$50 and Google's Gemini 3.8 Flash at $1.50/$7.50 standard (an introductory $0.75/$3.75 runs through Dec 31, 2026). Removed GPT-5.6 Sol: OpenAI publishes only a promotional $4/$20 rate for it and has never posted a standard price, so listing one would have meant guessing. Opus 5, Sonnet 5, Haiku 4.5, GPT-5.6 Terra/Luna, GPT-5.4 nano, Gemini 3.1 Pro and Grok 4.6 all confirmed unchanged. The calculator also now uses each model's own verified cached-input rate, editable per calculation, instead of assuming 10% for everyone - Claude Fable 5.1 caches at 2.5% of its input price and Grok 4.6 at 25%.

  11. Aug 25, 2026 Pricing

    Model prices re-verified across all four vendors

    Claude Sonnet 5 dropped to its now-permanent $2/$10 rate (the planned Sep 1 increase to $3/$15 was cancelled). Swapped Grok 4.5 for the newly launched Grok 4.6, xAI's new recommended flagship, at the same $2/$6 rate. OpenAI and Google prices confirmed unchanged (GPT-5.6 Sol has a time-boxed promo at $4/$20 through at least Nov 21, 2026; kept its $5/$30 standard rate pending confirmation of the post-promo price).

  12. Aug 1, 2026 Pricing

    Model prices re-verified across all four vendors

    Swapped Claude Opus 4.8 for the newly launched Claude Opus 5 (same $5/$25 rate); cut GPT-5.6 Terra to $2/$12 and GPT-5.6 Luna to $0.20/$1.20 on updated OpenAI pricing. Google and xAI prices confirmed unchanged.

  13. Jul 19, 2026 Pricing

    Model prices re-verified across all four vendors

    Added Claude Sonnet 5, GPT-5.6 Sol/Terra/Luna and Grok 4.5; removed the GPT-4.1 family and Grok 4.1, which vendors no longer list.

  14. Jul 19, 2026 Articles

    Two new articles: model-retirement migrations and agent operations

    An ops playbook for LLM model migrations (Opus 4.1 retires Aug 5), and a guide to running agents with budgets, gates and escalation.

  15. Jun 20, 2026 Articles

    Four practitioner how-tos published

    Eval datasets, prompt versioning with GitHub, hallucination-rate measurement, and what to log in an LLM trace.

  16. Jun 16, 2026 Articles

    RAG freshness monitoring checklist

    Field-by-field checklist for keeping a RAG index true, with an inline pipeline diagram.

  17. Jun 11, 2026 Feature

    Learn LLMOps page launched

    Thirteen live-verified courses, programs, certifications, books and podcasts, mapped to stack layers - plus site search, RSS feed and llms.txt.

  18. Jun 10, 2026 Pricing

    Claude Fable 5 added to the cost calculator

    Verified at $10/$50 per MTok alongside a full pricing re-check.

  19. Jun 8, 2026 Data

    Sources & citations system

    Central register of live-verified references, cited on the homepage, stack page, glossary and every article. Layer colour system and stack diagrams shipped the same week.

  20. Jun 5, 2026 Feature

    LLMOps.si launched

    The stack guide, production checklist, cost calculator, maturity score, glossary, tools directory, five reference architectures and the first ten articles.

Pricing entries are produced by a scheduled verification pass against vendor pricing pages. Spotted something stale anyway? Tell us.