The breakthroughs behind LLMOps: what each one changed in production
The papers and launches, 2017-2024, that made large language models capable, cheap to run, open and connected - and the operational work each one created. Every result taken from its primary source.
A short history of LLMOps follows how the practice took shape. This is the layer underneath it: the technical breakthroughs that changed what language models could do, what they cost to run, who could run them and what they could connect to.
Each one is here for the same reason. It didn’t only solve a problem - it created operational work. A capability breakthrough opened a new failure mode; an efficiency breakthrough moved the cost curve and made new deployments possible; an integration breakthrough widened what a model could touch, and therefore what could go wrong. Numbers come from each paper’s abstract or the announcement itself, not from secondary coverage.
-
Capability What models can do
- Jun 2017 The Transformer
- Jan 2020 Scaling laws for language models
- May 2020 GPT-3 and in-context learning Prompt management
- Jan 2022 Chain-of-thought prompting Prompt management
- Mar 2022 InstructGPT: instruction following from human feedback Evaluation
- Mar 2022 Chinchilla: compute-optimal training Cost control
- Sep 2024 Reasoning models (OpenAI o1) Cost control
-
Efficiency What they cost to run
- Jun 2021 LoRA: low-rank adaptation Deployment
- May 2022 FlashAttention Cost control
- Aug 2022 LLM.int8(): 8-bit inference Deployment
- Oct 2022 GPTQ: 3- and 4-bit quantization Deployment
- Nov 2022 Speculative decoding Deployment
- May 2023 QLoRA Deployment
- Sep 2023 PagedAttention (vLLM) Deployment
- Jan 2024 Mixtral: sparse mixture of experts Cost control
- Dec 2024 Prompt caching generally available (Anthropic) Cost control
-
Access Who can run them
- Feb 2023 LLaMA: open foundation models Deployment
- Jul 2023 Llama 2: free for commercial use Governance
-
Integration What they connect to
- May 2020 Retrieval-augmented generation RAG operations
- Oct 2022 ReAct: reasoning and acting Observability
- Jun 2023 Function calling Security
- Aug 2024 Structured Outputs Evaluation
- Nov 2024 Model Context Protocol Security
Capability: what models can do
The Transformer (2017)
Attention Is All You Need replaced recurrence with attention, and every large language model since is built on it. One property of that design still shows up on every invoice: the time and memory cost of self-attention grows quadratically with sequence length. Context length has been an operational dial - cost, latency, what fits - from the first day.
Scaling laws (2020)
Kaplan et al. found that language-model loss falls as a power law with model size, dataset size and training compute, across more than seven orders of magnitude. Returns from scale became predictable, and predictability justifies budgets. The consequence for operations was indirect but enormous: the most capable models became too large and too expensive to train for all but a few labs, so most teams would rent them through an API rather than own them.
GPT-3 and in-context learning (2020)
Language Models are Few-Shot Learners showed one large model performing new tasks from a handful of examples written into its input, with no retraining. Behaviour moved from weights into text. From that moment the prompt was a program - and a program nobody versions, reviews or tests is a production incident waiting for a date (prompt management).
Chain-of-thought prompting (2022)
Wei et al. showed that prompting a model to produce intermediate reasoning steps significantly improves its performance on arithmetic, commonsense and symbolic reasoning. Quality became something you could buy with output tokens, which made it a cost and latency trade-off rather than a fixed property of the model.
InstructGPT and human feedback (2022)
Ouyang et al. fine-tuned GPT-3 with reinforcement learning from human feedback. In human evaluations, outputs from the 1.3B-parameter InstructGPT model were preferred to those of the 175B-parameter GPT-3, despite having 100× fewer parameters. Following instructions by default is what made models usable in products at all. It also meant a model’s behaviour depends on post-training you don’t see - so two versions behind the same model name can behave differently, and every upgrade needs an eval run.
Chinchilla: compute-optimal training (2022)
Kaplan’s results had pointed toward very large models trained on relatively modest data. Hoffmann et al. revised that: for compute-optimal training, model size and training tokens should scale equally. Their 70B-parameter Chinchilla, trained on four times more data with the same compute as the 280B Gopher, outperformed it. Smaller models trained on more data are cheaper to serve at a given quality - the economics behind the small-model tier that routing depends on.
Reasoning models (2024)
OpenAI’s o1 was trained with large-scale reinforcement learning to use its chain of thought productively, and OpenAI reported that its performance improves both with more training compute and with more time spent thinking at inference. “Think longer” became a setting. For operations that means three things. Reasoning tokens are billed: OpenAI’s documentation states that they aren’t visible through the API but are billed as output tokens, and that a request which hits its output limit mid-reasoning can cost money without returning any visible answer. Latency varies with the problem, not just the prompt length. And the effort or thinking-budget parameter belongs in versioned config and in your eval runs, like any other model setting.
Efficiency: what they cost to run
LoRA and QLoRA (2021, 2023)
LoRA freezes the model and trains small low-rank matrices instead. Compared with fully fine-tuning GPT-3 175B, it cut the number of trainable parameters by 10,000× and the GPU memory requirement by 3×. QLoRA went further, fine-tuning a 65B-parameter model on a single 48GB GPU while preserving full 16-bit fine-tuning performance. Fine-tuning went from a research project to an option - and brought classical MLOps back into the picture: training data to curate, adapters to version, a model artifact to evaluate (LLMOps vs MLOps).
FlashAttention (2022)
Dao et al. made attention IO-aware - tiling the computation to cut reads and writes between GPU memory levels - while keeping it exact. They reported a 3× training speedup on GPT-2 at 1K sequence length, and longer usable context. Long context windows followed, and with them a new operational failure mode: context bloat, where “it fits” quietly replaces “it’s needed”.
Quantization: LLM.int8() and GPTQ (2022)
LLM.int8() cut the memory needed for inference by half while retaining full-precision performance, and loaded a 175B-parameter checkpoint in 8-bit without degradation. GPTQ quantized 175B-parameter models to 3 or 4 bits per weight in about four GPU hours, with negligible accuracy loss against the uncompressed baseline. Self-hosting suddenly needed far fewer GPUs. The catch is operational: a quantized model is a different artifact from its full-precision parent, and it deserves its own eval run before it replaces anything.
Speculative decoding (2022)
Leviathan et al. let a small draft model propose several tokens that the large model verifies in parallel, and reported a 2-3× speedup on T5-XXL with identical outputs and no retraining. Latency improvements rarely come without a quality trade-off; this one did, and serving engines such as vLLM now support it. Its gain depends on how often the draft is right, so it’s a number to measure on your own traffic rather than assume.
PagedAttention and vLLM (2023)
Kwon et al. borrowed virtual-memory paging from operating systems to manage the key-value cache, cutting its waste to near zero and letting requests share it. vLLM improved throughput by 2-4× at the same latency compared with the serving systems it was measured against. Throughput per GPU is cost per request, so this moved self-hosted serving from “possible” to “competitive”.
Mixtral and sparse experts (2024)
In Mixtral, each token is routed to two experts per layer: it has access to 47B parameters but uses 13B active parameters during inference. Compute per token tracks the active parameters, while memory must still hold all of them - a split that matters when you size serving hardware and compare cost per token.
Prompt caching (2024)
Reusing a processed prompt prefix across requests, pitched by Anthropic at up to 90% lower cost for long prompts, turned prompt layout into a cost decision: stable content first, per-request content last, and a cache hit rate to watch alongside latency.
Access: who can run them
LLaMA and Llama 2 (2023)
Meta’s LLaMA models, from 7B to 65B parameters, were trained only on publicly available datasets, and the 13B model outperformed the 175B GPT-3 on most benchmarks. They were released to the research community. Llama 2, five months later, was free for research and commercial use. Together with quantization and vLLM, open weights made self-hosting a real choice - keeping data in your own environment, pinning a model that no vendor can retire - and made the serving, scaling and upgrade work yours (deployment).
Integration: what they connect to
Retrieval-augmented generation (2020)
Lewis et al. combined a generator with a retriever that fetches documents at query time. Grounding answers in your own content is what made LLMs useful for most business questions - and it added a whole subsystem to operate: an index to keep fresh, retrieval quality to measure, permissions to enforce (RAGOps).
ReAct: reasoning and acting (2022)
Yao et al. had models interleave reasoning traces with actions that query external sources, so each step could inform the next. That loop is the shape of today’s agents. Operationally it replaced “one request, one response” with a multi-step trajectory - which is why agents need per-step tracing, budgets and a way out.
Function calling (2023)
OpenAI let developers describe functions and have the model return a JSON object with arguments to call them. Tool use became a standard API feature, and the model gained the ability to act. Prompt injection stopped being a text problem and became an action problem (LLM security).
Structured Outputs (2024)
OpenAI’s Structured Outputs constrained model output to match developer-supplied JSON Schemas, and on OpenAI’s own complex schema-following evaluation its new model scored 100% with the feature enabled. Parsing failures stopped being a routine error class. It also made the cheapest kind of eval far more useful: once output is guaranteed to be structured, fields can be checked literally instead of judged.
Model Context Protocol (2024)
Anthropic open-sourced MCP as a standard for connecting assistants to the systems where data lives. Integrating a new tool no longer meant writing a new adapter for every application - and every connected server became part of the attack surface to scope, authenticate and log (operating agents).
The pattern
| Breakthrough | Year | What it changed in production |
|---|---|---|
| Transformer | 2017 | Context length became a cost and latency dial |
| Scaling laws | 2020 | The best models became rented, not owned |
| In-context learning | 2020 | Prompts became programs to version and test |
| Chain-of-thought | 2022 | Quality could be bought with output tokens |
| Instruction tuning (RLHF) | 2022 | Usable by default; behaviour varies by version |
| Compute-optimal training | 2022 | Small, cheap models good enough to route to |
| Reasoning models | 2024 | Effort became a billed, versioned setting |
| LoRA / QLoRA | 2021-23 | Fine-tuning became an option, with MLOps attached |
| FlashAttention | 2022 | Long context became affordable - and bloat-prone |
| Quantization | 2022 | Self-hosting on far fewer GPUs; new artifacts to evaluate |
| Speculative decoding | 2022 | Lower latency with identical outputs |
| PagedAttention | 2023 | Self-hosted throughput became competitive |
| Mixture of experts | 2024 | Compute and memory costs split apart |
| Prompt caching | 2024 | Prompt layout became a cost decision |
| Open weights | 2023 | Self-hosting and data control became a real choice |
| RAG | 2020 | A retrieval subsystem to measure and keep fresh |
| ReAct | 2022 | Multi-step trajectories to trace and budget |
| Function calling | 2023 | Models could act; injection could too |
| Structured Outputs | 2024 | Literal checks on structured fields |
| MCP | 2024 | A standard tool surface - and attack surface |
Read down the right-hand column and most of the LLMOps stack appears on its own. None of these breakthroughs made operations unnecessary; each one moved the hard part somewhere new. That is the most useful thing the history teaches about the next breakthrough: ask what it will make possible, and then ask what it will make someone have to operate.