← All articles
LLMOps

A short history of LLMOps: from hidden technical debt to agents

How LLMOps took shape, 2015-2026: the MLOps warning about technical debt, models turning into APIs, the 2023 production scramble, standards and regulation, and the agent era. Every milestone dated and linked to its primary source.

September 27, 2026 · 8 min read · fundamentals · history

LLMOps didn’t launch as a product category. It accumulated, one production problem at a time, as language models moved out of research papers and into APIs anyone could call. Answers nobody could trace, invoices nobody predicted, prompts nobody could roll back, a model retired from under a running system: each of the eight layers of the LLMOps stack is, in effect, the lesson from one of those moments.

This is that story in order, with every milestone dated and linked to its primary source. One thing it deliberately doesn’t do is name the day the word “LLMOps” was coined. There is no reliable record of one, and the practices mattered well before the name settled. Today seven major platforms publish their own definition of it. For the technical side of the story - the papers that made models capable, cheap to run and connected - see the companion piece, the breakthroughs behind LLMOps.

  1. 2015-2019 The warning and the architecture

    1. 2015 Hidden Technical Debt in Machine Learning Systems 
    2. Jun 2017 Attention Is All You Need - the Transformer 
  2. 2020-2022 The model becomes an API

    1. May 2020 GPT-3: tasks specified in the prompt  Prompt management
    2. May 2020 Retrieval-augmented generation (RAG)  RAG operations
    3. Jun 2020 The OpenAI API opens 
    4. Oct 2022 LangChain 0.0.1: chains of prompts, retrieval and tools 
    5. Nov 2022 ChatGPT 
  3. 2023 The production scramble

    1. Jan 2023 NIST AI Risk Management Framework 1.0  Governance
    2. Mar 2023 gpt-3.5-turbo API at a tenth of the old price  Cost control
    3. Apr 2023 Building LLM applications for production (Chip Huyen) 
    4. Jun 2023 Judging LLM-as-a-Judge (MT-Bench)  Evaluation
    5. Jun 2023 Function calling: models that act  Security
    6. Jun 2023 vLLM and PagedAttention  Deployment
    7. Jul 2023 LangSmith: from prototype to production  Observability
    8. 2023 OWASP Top 10 for LLM Applications begins  Security
  4. 2024 Standards and structural levers

    1. Jul 2024 NIST Generative AI Profile (AI 600-1)  Governance
    2. Aug 2024 EU AI Act enters into force  Governance
    3. Nov 2024 Model Context Protocol open-sourced  Security
    4. Dec 2024 OpenTelemetry for Generative AI  Observability
    5. Dec 2024 Prompt caching generally available (Anthropic)  Cost control
  5. 2025-2026 Agents, calendars and consolidation

    1. Feb 2025 EU AI Act: prohibitions and AI literacy apply  Governance
    2. Nov 2025 Replicate announces it is joining Cloudflare 
    3. Jan 2026 Langfuse acquired by ClickHouse 
    4. Mar 2026 Helicone joins Mintlify, enters maintenance mode 
    5. Jul 2026 AI Omnibus defers high-risk AI Act obligations  Governance
    6. Aug 2026 Claude Opus 4.1 retired  Deployment

2015-2019: the warning and the architecture

In 2015 D. Sculley and colleagues published Hidden Technical Debt in Machine Learning Systems at NeurIPS. Its argument was blunt: machine learning lets you build complex prediction systems quickly, but those quick wins are not free. Real-world ML systems tend to run up large, ongoing maintenance costs, from entanglement, hidden feedback loops, undeclared consumers, data dependencies, configuration and changes in the outside world. The model code is a small part of the system, and most of the risk lives in the rest.

That is the clearest early statement of the idea MLOps was built on, and LLMOps inherits it whole. Two years later, Attention Is All You Need introduced the Transformer, the architecture behind today’s large language models. It had no operations story yet. Everything that follows depends on it.

2020-2022: the model becomes an API

In May 2020 the GPT-3 paper showed a single large model performing new tasks from a few examples written into its input. That quietly moved behaviour out of trained weights and into text: the prompt became a program, and programs need versioning, review and rollback. The same month, Lewis et al. described retrieval-augmented generation - grounding a model’s answer in documents fetched at query time. Once RAG reached production, “was the right context retrieved, and is it still true?” became an operational question of its own (RAGOps).

The structural shift came in June 2020, when OpenAI opened an API to its models. From then on most teams building with language models would never train the model they depended on. MLOps’ central object - a pipeline that produces a model artifact - gave way to a system composed around someone else’s model, where everything you control lives outside the weights. That difference is the core of LLMOps vs MLOps.

In October 2022 LangChain published its first release on PyPI. Frameworks that chain prompts, retrieval and tool calls turned a single request into a multi-step program - which is why a trace, not a log line, became the unit of LLM observability. A month later ChatGPT arrived, and demand for LLM features stopped being hypothetical.

2023: the production scramble

2023 is when the gap between demo and production became everybody’s problem.

Cost became a live metric. In March 2023 OpenAI opened gpt-3.5-turbo to developers at $0.002 per 1,000 tokens, which it described as ten times cheaper than its existing GPT-3.5 models. Cheap tokens invited volume, and volume turned spend from a line item into a per-request, per-feature number to watch (cost control).

The problems got named. In April, Chip Huyen’s essay Building LLM applications for production set out the gap plainly: impressive demos are easy, production-ready systems are hard. It pointed at the ambiguity of natural-language prompts, cost and latency, and the need to version and evaluate prompts systematically. Read today, it is a sketch of half the stack.

Evaluation found its workhorse. In June, Zheng et al. reported that strong LLM judges such as GPT-4 agreed with human preferences over 80% of the time - about as often as humans agree with each other - while documenting position, verbosity and self-enhancement biases. LLM-as-a-judge became the standard way to score open-ended output, and calibrating the judge against human labels became part of the practice (how to build your first eval dataset).

Models started to act. Also in June, OpenAI added function calling: the model could return a structured call to a function the developer described. From here a model could do things, not just say them - and a prompt injection could become an action instead of an embarrassing reply (LLM security).

Serving became engineering. vLLM launched with PagedAttention, reporting up to 24x higher throughput than Hugging Face Transformers. For teams hosting open models, inference efficiency became a deployment and cost discipline in its own right.

Tooling and vocabulary arrived. In July LangChain announced LangSmith, pitched as closing the gap between prototype and production - the same framing this site uses - alongside a wave of tracing and evaluation platforms (see the tools directory). OWASP started its Top 10 for LLM Applications, giving security teams a shared list with prompt injection at the top. And NIST’s AI Risk Management Framework 1.0, released that January, gave governance a common reference point.

2024: standards and structural levers

2024 was less about new problems than about standard ways to handle them.

Regulation became concrete. The EU AI Act entered into force on 1 August 2024, with obligations phasing in over the following years. Days earlier, on 26 July, NIST released its Generative AI Profile (NIST AI 600-1) to help organisations identify the risks specific to generative AI.

Tools got a standard plug. In November, Anthropic open-sourced the Model Context Protocol, a standard for connecting assistants to the systems where data lives. Integrating tools got cheaper - and every connected tool became something to scope and secure.

Observability got a shared schema. In December the OpenTelemetry project described its generative-AI semantic conventions, covering traces, metrics and events, with the first instrumentation library targeting OpenAI’s Python client. The conventions are still marked as in development, and have since moved to their own repository, but they gave traces a vendor-neutral shape (what to log in an LLM trace).

Cost control got architecture. Prompt caching, introduced by Anthropic in public beta and generally available on its API by December 2024, was pitched at cutting costs by up to 90% for long prompts. Cost control stopped being only “use a cheaper model” and became a design question: keep prefixes stable, measure the cache hit rate.

2025-2026: agents, calendars and consolidation

Regulation phased in. The EU AI Act’s prohibited practices and AI-literacy obligations applied from 2 February 2025 and the obligations for general-purpose AI models from 2 August 2025. The Act became applicable on 2 August 2026, with exceptions. The largest of those exceptions moved: the AI Omnibus, Regulation (EU) 2026/1744, in force since 27 July 2026, deferred the high-risk obligations to 2 December 2027 for Annex III systems and 2 August 2028 for AI embedded in regulated products. What those obligations ask for - automatic logging over a system’s lifetime and effective human oversight - is exactly the governance layer.

Models got calendars. Vendors now retire models on a published schedule. Anthropic’s deprecation table shows Claude Sonnet 4 and Claude Opus 4 retired on 15 June 2026 and Claude Opus 4.1 on 5 August 2026, with at least 60 days’ notice for publicly released models. A model became a dependency with a deprecation calendar, and migrating off one became routine operations work (the migration playbook).

The first wave of tools consolidated. Replicate announced it was joining Cloudflare in November 2025. ClickHouse acquired Langfuse in January 2026, which kept its MIT-licensed core and self-hosting. Helicone joined Mintlify in March 2026 and moved into maintenance mode. Humanloop’s team joined Anthropic and sunset its platform; Lakera became part of Check Point; Protect AI became part of Palo Alto Networks. Within about three years, much of the first generation of LLMOps tooling had changed hands. The practices didn’t.

Agents raised the stakes. Once a model can call tools in a loop, the operational question changes from “is the answer right?” to “is the action safe, bounded and reversible?” Budgets, approval gates and escalation paths are the newest layer of the practice (operating LLM agents in production).

What the history teaches

Line the milestones up against the stack and every layer turns out to be an answer to a specific pressure:

LayerThe pressure that created itMilestone
Prompt managementBehaviour moved from weights into textGPT-3 and few-shot prompting (2020)
EvaluationOpen-ended output with no label to check againstLLM-as-a-judge (2023)
ObservabilityOne request became a multi-step programLangChain (2022), LangSmith (2023), OTel GenAI (2024)
Cost controlCheap tokens at enormous volumegpt-3.5-turbo pricing (2023), prompt caching (2024)
RAG operationsAnswers grounded in documents that changeThe RAG paper (2020)
SecurityModels that call toolsFunction calling and the OWASP LLM Top 10 (2023)
GovernanceAccountability to customers and regulatorsNIST AI RMF (2023), EU AI Act (2024-2028)
DeploymentSelf-hosted serving and vendor retirementsvLLM (2023), model retirements (2025-2026)

Three lessons come out of that table.

The model was never the whole system. Sculley and colleagues said it about machine learning in 2015, and every milestone since has added more system around the model, not less.

Tools are temporary; practices compound. Several of the first-generation tools changed owners within three years. Versioned prompts, an eval set built from real failures and traces in an open format survive a change of vendor - which is the strongest argument for owning them yourself.

The next layer is already forming. Agents are to 2026 what RAG was to 2023: the pattern everyone is shipping and the place the next set of failure modes will come from.

Where to go next

If the history made the case, the present tense is in What is LLMOps? and the eight-layer stack. To see which of those layers your own system is still missing, take the Maturity Score or work through the Production Checklist.

Get the Production Checklist → Explore the Stack →