A guided library for understanding AI Understanding Machine AI elucidaria — a plain-language handbook, in the old sense

Course lesson · Day 9 of 30 6 figures

From Base Model to Assistant

A language model fresh from pre-training does one thing: it continues text. Asked to explain the moon landing to a six-year-old, one such model wrote four more requests of the same kind. The assistants in daily use are the same kind of machine after further training, and the first stage of that training is simple in principle and small beside pre-training: the same next-token objective, applied to a chosen set of documents in which a request is followed by the response a developer wants. Its lineage runs from a 2015 Google paper through OpenAI's InstructGPT to ChatGPT. What the stage adds is disputed. One camp holds that it mostly teaches a format the model already had the knowledge to fill; another warns that, because format is so easily taught, a tuned model can sound more capable than it is. On 28 September 2026 OpenAI is scheduled to retire the two models its documentation describes as base models not trained with instruction following.

About 25 min read 12 min listen Print edition (PDF)

Published Sources read through

Listen · 12 min

The spoken edition of this episode — its own script, read by a synthetic voice, with quoted passages in a second voice. It is not this page read aloud: the written edition you are reading was written separately, and neither needs the other. Download

Episode notes

The argument of the spoken edition, how it runs, what to take from it and the sources it read — in the episode’s own words. The full transcript is below; the written edition, with its figures, follows.

The argument

In January twenty twenty-two, OpenAI published a post with side-by-side examples: one request, two models. One request was, explain the moon landing to a six year old in a few sentences.

How it runs

  1. Why it's hard to follow — The paper behind those examples, by Long Ouyang and colleagues at OpenAI, opens with this sentence. That invites two readings, and both go too far.
  2. The idea you need — Two ideas: supervised fine-tuning, and instruction tuning. A base model is what pre-training produces: a system that, given text, scores what comes next. Give it a request and, in effect, it continues whatever kind of document that request usually begins.
  3. What actually happened — Twenty twenty-one: instruction tuning on research tasks — a dataset called Natural Instructions in April, Google's FLAN in September, a model called T-zero in October. January and March twenty twenty-two: InstructGPT.
  4. What happened next — So what does this stage actually do? Two positions from twenty twenty-three. The first came from Meta.
  5. What to watch — Two things you can check. One. OpenAI's deprecations page lists the twenty-eighth of September twenty twenty-six as the shutdown date for davinci zero zero two and babbage zero zero two.

What to take from it

The idea to keep is supervised fine-tuning: same objective, different text. A base model continues documents. An assistant is that same kind of machine, trained further on one particular kind of document — a request, followed by a helpful answer.

So when an assistant surprises you, ask which stage that came from: what it read in pre-training, or what it was shown afterwards.

To read more, the encyclopedia's article on reinforcement learning from human feedback describes this as the first of its three stages.

Sources read for this episode (18)

  1. OpenAI, *Aligning Language Models to Follow Instructions* (Internet Archive capture of 1 January 2023) — 27 January 2022
  2. Long Ouyang et al. (OpenAI), *Training language models to follow instructions with human feedback*, arXiv 2203.02155 v1 — 4 March 2022
  3. OpenAI, API deprecations page; davinci-002 and babbage-002 model pages — read 17 September 2026
  4. Hyung Won Chung et al. (Google), *Scaling Instruction-Finetuned Language Models*, arXiv 2210.11416 (version 5 of 6 December 2022 read) — first posted 20 October 2022
  5. Andrew M. Dai and Quoc V. Le (Google), *Semi-supervised Sequence Learning*, arXiv 1511.01432 — 4 November 2015
  6. Alec Radford et al. (OpenAI), *Improving Language Understanding by Generative Pre-Training* — 2018
  7. Alec Radford et al. (OpenAI), *Language Models are Unsupervised Multitask Learners* — 2019
  8. Jason Wei et al. (Google Research), *Finetuned Language Models Are Zero-Shot Learners*, arXiv 2109.01652 v1, and ICLR 2022 version — 3 September 2021
  9. Swaroop Mishra et al., *Natural Instructions: Benchmarking Generalization to New Tasks from Natural Language Instructions*, arXiv 2104.08773 v1 (later retitled *Cross-Task Generalization via Natural Language Crowdsourcing Instructions*) — 18 April 2021
  10. Victor Sanh et al., *Multitask Prompted Training Enables Zero-Shot Task Generalization*, arXiv 2110.08207 v1 — 15 October 2021
  11. OpenAI, *ChatGPT: Optimizing Language Models for Dialogue* (Internet Archive capture of 1 January 2023) — 30 November 2022
  12. Stanford CRFM, *Alpaca*, and its training code (read 17 September 2026) — 13 March 2023
  13. Allen Institute for AI, open-instruct training code (read 17 September 2026) — 2026
  14. Nathan Lambert et al. (AI2), *Tülu 3*, arXiv 2411.15124 (version 5 of 14 April 2025 read) — first posted 22 November 2024
  15. Chunting Zhou et al. (Meta AI and others), *LIMA: Less Is More for Alignment*, arXiv 2305.11206 v1 — 18 May 2023
  16. Arnav Gudibande et al. (UC Berkeley), *The False Promise of Imitating Proprietary LLMs*, arXiv 2305.15717 v1 — 25 May 2023
  17. Bill Yuchen Lin et al. (AI2, University of Washington), *The Unlocking Spell on Base LLMs*, arXiv 2312.01552 v1 — 4 December 2023
  18. Hugging Face repositories and model cards: DeepSeek, Google, Qwen, Moonshot, Z.ai, Allen Institute for AI, OpenAI — read 17 September 2026
Full transcript — 1,677 words, about 8 min read

The spoken edition, word for word. A quoted passage is its source read aloud — the source’s own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Download the transcript (Markdown). Printing this page prints the transcript; the article has its own print edition.

In January twenty twenty-two, OpenAI published a post with side-by-side examples: one request, two models. One request was, explain the moon landing to a six year old in a few sentences.

The first was GPT-three as it came out of its original training. It wrote this.

Explain the theory of gravity to a 6 year old. Explain the theory of relativity to a 6 year old in a few sentences. Explain the big bang theory to a 6 year old. Explain evolution to a 6 year old.

— OpenAI, 'Aligning Language Models to Follow Instructions' (27 January 2022), GPT-3 completion to the prompt 'Explain the moon landing to a 6 year old in a few sentences.'; Internet Archive capture 20230101110545 of openai.com/blog/instruction-following/

The second model, which OpenAI called InstructGPT, simply answered: people went to the moon, took pictures of what they saw, and sent them back to earth.

The post doesn't say how those examples were picked. In the paper behind it, for a similar pair about some code, the authors say the untuned model does answer about half the time. So this isn't a model that can't answer. It's a model doing what it was trained to do: continue a document.

And on the twenty-eighth of September, eleven days from now, OpenAI is scheduled to shut down davinci zero zero two and babbage zero zero two, which its own documentation calls base models, not trained with instruction following. How a continuation machine becomes an assistant is today's subject.

The paper behind those examples, by Long Ouyang and colleagues at OpenAI, opens with this sentence.

Making language models bigger does not inherently make them better at following a user’s intent.

— Long Ouyang et al. (OpenAI), 'Training language models to follow instructions with human feedback', arXiv 2203.02155v1 (4 March 2022), abstract, first sentence

That invites two readings, and both go too far.

The first: the original model is the unfinished version, and the extra training is where it learns things. The sizes point the other way. OpenAI's supervised stage used about thirteen thousand prompts, most of them written by the contractors it hired, and the company's post says the whole procedure used less than two percent of the computing and data of pre-training. OpenAI's own gloss is that the training unlocks abilities GPT-three already had — but it introduces that as, in its words, one way of thinking about this process. What it really adds is argued over, and I'll come to that.

The second reading is the mirror image: an assistant is just the original model with a clever prompt in front. That's half right. In the same post, one prompt was laid out like a page of questions and answers — two answered, a third left open — and the untuned model answered the third sensibly. But the assistant you use isn't a base model plus a prompt. Its numbers have been changed by more training.

Two ideas: supervised fine-tuning, and instruction tuning.

A base model is what pre-training produces: a system that, given text, scores what comes next. Give it a request and, in effect, it continues whatever kind of document that request usually begins.

Training in two stages is older than any chatbot. In November twenty fifteen, Andrew Dai and Quoc Le at Google trained networks on unlabelled text — one way by predicting what comes next, another by rebuilding the input — then used what they had learned as the starting point for ordinary tasks with labelled answers. They wrote this.

These two algorithms can be used as a "pretraining" step for a later supervised sequence learning algorithm.

— Andrew M. Dai and Quoc V. Le (Google), 'Semi-supervised Sequence Learning', arXiv 1511.01432 (4 November 2015), abstract; https://arxiv.org/abs/1511.01432, read 17 September 2026

In twenty eighteen, OpenAI's first GPT paper did the same with a transformer: pre-train once, then fine-tune separately for each task. Then GPT-two, in twenty nineteen, was tested with no fine-tuning at all, so a request had to be disguised as a document. To get a summary, the team put the letters T L semicolon D R and a colon after a news article — the internet's too long, didn't read — and let the model carry on. It worked, barely. OpenAI's paper says the summaries only just beat picking three random sentences from the article, and scored about six points worse without that hint.

Now the first idea. Supervised fine-tuning means you keep training the same model, with the same objective — predict the next token — but on a much smaller set of examples somebody chose, each one a request followed by the response they want. Nothing new is bolted on. What changes is the kind of document the model has learned to expect. After enough of those examples, the likeliest thing to follow a request is an answer.

The second idea makes that habit general. In September twenty twenty-one, a Google Research team led by Jason Wei described instruction tuning as

finetuning language models on a collection of tasks described via instructions

— Jason Wei et al. (Google Research), 'Finetuned Language Models Are Zero-Shot Learners', arXiv 2109.01652v1 (3 September 2021), abstract, the parenthetical definition of instruction tuning

They took more than sixty existing research datasets, wrote plain-language instructions for each, tuned a large model on most of them, and tested it on kinds of task it hadn't been tuned on. It did substantially better on those than the same model untuned. The paper adds a condition worth keeping: in their experiments, for models of eight billion parameters and smaller, instruction tuning made the unseen tasks worse.

A conversation reaches a model as one run of tokens, with special markers where each speaker's turn begins. The way I'd put it, fine-tuning is what gives those markers meaning: the training conversations share one layout, so the model learns what comes after the marker that opens the assistant's turn.

Twenty twenty-one: instruction tuning on research tasks — a dataset called Natural Instructions in April, Google's FLAN in September, a model called T-zero in October.

January and March twenty twenty-two: InstructGPT. OpenAI hired about forty contractors, collected their written demonstrations, and fine-tuned GPT-three on them. Then came a second stage, built on rankings of its answers — that's tomorrow. OpenAI also tuned GPT-three on the FLAN and T-zero research collections, and on OpenAI's own mix of prompts, those versions did slightly worse than the one tuned on its contractors' demonstrations. That's OpenAI's measurement, on OpenAI's prompts, judged by OpenAI's contractors, and as far as I can find the prompts were never published, so nobody outside can rerun it.

November thirtieth, twenty twenty-two: ChatGPT. OpenAI's announcement describes the first step.

We trained an initial model using supervised fine-tuning: human AI trainers provided conversations in which they played both sides

— OpenAI, 'ChatGPT: Optimizing Language Models for Dialogue' (30 November 2022), Methods section; Internet Archive capture 20230101000602 of openai.com/blog/chatgpt/. The source sentence continues '—the user and an AI assistant.'

Both sides played by people, who could draw on suggestions the model wrote: supervised fine-tuning, in the shape of a chat.

November twenty twenty-four: the Allen Institute for AI published Tülu three, a full post-training recipe released with its data and code, built on Meta's base models. Its supervised stage used about nine hundred and forty thousand prompts. The authors' reason for publishing all of it: the data and recipes for post-training are, in their words, the portion with the least transparency.

And this year, some developers still publish both halves. On Hugging Face, DeepSeek's base and chat repositories for its V-four Pro model were created on the same April day, and Google's Gemma four comes in pre-trained and instruction-tuned versions. Several of this year's other big open releases — from Qwen, Moonshot and Z-dot-A-I — have no public base version there as of this week.

So what does this stage actually do? Two positions from twenty twenty-three.

The first came from Meta. Chunting Zhou and colleagues fine-tuned a large base model on one thousand carefully chosen examples, called it LIMA, and proposed what they named the Superficial Alignment Hypothesis.

A model’s knowledge and capabilities are learnt almost entirely during pretraining, while alignment teaches it which subdistribution of formats should be used when interacting with users.

— Chunting Zhou et al. (Meta AI et al.), 'LIMA: Less Is More for Alignment', arXiv 2305.11206v1 (18 May 2023), section 2 'Alignment Data', definition of the Superficial Alignment Hypothesis (preceded by 'We define the Superficial Alignment Hypothesis:')

A hypothesis, in their word. Later that year a team from the Allen Institute and the University of Washington checked it another way. They took answers written by tuned models and asked, token by token, which token the matching base model would have ranked first at that point. Across the three pairs they measured, it was the same token about four times in five, and the biggest shifts were in words like hello, thank, however and remember. That's the base model reading the tuned model's answer as it goes, not writing it alone.

The second position is a warning. In May twenty twenty-three, a Berkeley team fine-tuned open models on ChatGPT's answers. Crowd raters found the results competitive with ChatGPT. More targeted automatic tests found the imitators closed little or none of the gap, on tasks their training data didn't cover well. Their explanation:

imitation models are adept at mimicking ChatGPT’s style but not its factuality

— Arnav Gudibande et al. (UC Berkeley), 'The False Promise of Imitating Proprietary LLMs', arXiv 2305.15717v1 (25 May 2023), abstract (the sentence begins 'We show that these performance discrepancies may slip past human raters because')

The two positions agree on the mechanism: this stage moves style easily. They disagree about what that's worth. For the LIMA team, a small, careful set of examples on a strong base model is enough. For the Berkeley team, that ease is the danger, because raters can mistake the style for the substance, and they argued the better lever is a better base model.

Two things you can check.

One. OpenAI's deprecations page lists the twenty-eighth of September twenty twenty-six as the shutdown date for davinci zero zero two and babbage zero zero two. After that date, look at whether their model pages are still up, and whether anything left in OpenAI's model list is described as not trained with instruction following.

Two. DeepSeek's V-four-point-one Flash, whose Hugging Face page was created on the tenth of September. Its model card reports test results for a base version, but as of today there's no public repository for that base model. Watch whether one appears: it's one sign of whether new open models can still be met before their fine-tuning.

The idea to keep is supervised fine-tuning: same objective, different text. A base model continues documents. An assistant is that same kind of machine, trained further on one particular kind of document — a request, followed by a helpful answer.

So when an assistant surprises you, ask which stage that came from: what it read in pre-training, or what it was shown afterwards.

To read more, the encyclopedia's article on reinforcement learning from human feedback describes this as the first of its three stages.

Tomorrow: the second training signal — people's rankings.

That was day nine. Thank you for listening.

Sources (18)

  1. OpenAI, Aligning Language Models to Follow Instructions (Internet Archive capture of 1 January 2023) — 27 January 2022
  2. Long Ouyang et al. (OpenAI), Training language models to follow instructions with human feedback, arXiv 2203.02155 v1 — 4 March 2022
  3. OpenAI, API deprecations page; davinci-002 and babbage-002 model pages — read 17 September 2026
  4. Hyung Won Chung et al. (Google), Scaling Instruction-Finetuned Language Models, arXiv 2210.11416 (version 5 of 6 December 2022 read) — first posted 20 October 2022
  5. Andrew M. Dai and Quoc V. Le (Google), Semi-supervised Sequence Learning, arXiv 1511.01432 — 4 November 2015
  6. Alec Radford et al. (OpenAI), Improving Language Understanding by Generative Pre-Training — 2018
  7. Alec Radford et al. (OpenAI), Language Models are Unsupervised Multitask Learners — 2019
  8. Jason Wei et al. (Google Research), Finetuned Language Models Are Zero-Shot Learners, arXiv 2109.01652 v1, and ICLR 2022 version — 3 September 2021
  9. Swaroop Mishra et al., Natural Instructions: Benchmarking Generalization to New Tasks from Natural Language Instructions, arXiv 2104.08773 v1 (later retitled Cross-Task Generalization via Natural Language Crowdsourcing Instructions) — 18 April 2021
  10. Victor Sanh et al., Multitask Prompted Training Enables Zero-Shot Task Generalization, arXiv 2110.08207 v1 — 15 October 2021
  11. OpenAI, ChatGPT: Optimizing Language Models for Dialogue (Internet Archive capture of 1 January 2023) — 30 November 2022
  12. Stanford CRFM, Alpaca, and its training code (read 17 September 2026) — 13 March 2023
  13. Allen Institute for AI, open-instruct training code (read 17 September 2026) — 2026
  14. Nathan Lambert et al. (AI2), Tülu 3, arXiv 2411.15124 (version 5 of 14 April 2025 read) — first posted 22 November 2024
  15. Chunting Zhou et al. (Meta AI and others), LIMA: Less Is More for Alignment, arXiv 2305.11206 v1 — 18 May 2023
  16. Arnav Gudibande et al. (UC Berkeley), The False Promise of Imitating Proprietary LLMs, arXiv 2305.15717 v1 — 25 May 2023
  17. Bill Yuchen Lin et al. (AI2, University of Washington), The Unlocking Spell on Base LLMs, arXiv 2312.01552 v1 — 4 December 2023
  18. Hugging Face repositories and model cards: DeepSeek, Google, Qwen, Moonshot, Z.ai, Allen Institute for AI, OpenAI — read 17 September 2026

1. A request, continued

On 27 January 2022 OpenAI published Aligning Language Models to Follow Instructions, a post built around paired examples: one prompt, two models. The first was GPT-3 as it came out of pre-training; the second, which OpenAI called InstructGPT, had been trained further. For the prompt "Explain the moon landing to a 6 year old in a few sentences." GPT-3 wrote:

Explain the theory of gravity to a 6 year old. Explain the theory of relativity to a 6 year old in a few sentences. Explain the big bang theory to a 6 year old. Explain evolution to a 6 year old.

InstructGPT replied: "People went to the moon, and they took pictures of what they saw, and sent them back to the earth so we could all see them." For "Write a short poem about a wise frog." GPT-3 produced three prompts for short stories and InstructGPT a poem. The caption beneath the examples reads: "GPT-3 models aren’t trained to follow user instructions."

The post does not say how its examples were chosen. The paper behind it, by Long Ouyang and colleagues (arXiv 2203.02155, 4 March 2022), does say so for a similar figure, in which GPT-3, asked what a list in a short function is for, replies with four multiple-choice options of its own while InstructGPT explains the code: "Prompts are cherry-picked to illustrate certain behaviors, but the outputs are not cherry-picked," and, for the code question, "GPT-3 does answer the question about 50% of the time."

Google recorded the same behaviour in its own model. In Scaling Instruction-Finetuned Language Models (Hyung Won Chung and colleagues, arXiv 2210.11416, first posted 20 October 2022), the untuned PaLM 540B, asked to "Make up a word that means "when two AI researchers go on a date".", repeated the request back; the instruction-tuned Flan-PaLM answered "date-mining". Given a question about square and cube roots, PaLM wrote the same question again with different numbers. The paper lists three undesired behaviours of the untuned model — "(1) continuing to generate related text instead of answering a question, (2) repeating the input question with minor modifications, and (3) not knowing when to stop generating text" — and offers a hedged cause: "This is likely an artifact of not using end-of-sequence tokens in pre-training." Its figure caption adds that the errors "can be mitigated by using few-shot exemplars".

None of this is a malfunction. Pre-training teaches a model to predict the next token of documents, and a line of the form explain X to a six-year-old can as plausibly open a list of homework or writing prompts as introduce an answer. That account of why GPT-3 wrote a list is an inference from the training objective rather than a statement by OpenAI; what the record shows directly is that the untuned model treated a request as the start of a document and extended it.

Models of that kind are still published, though not by every developer. OpenAI's model pages for davinci-002 and babbage-002, read on 17 September 2026, describe them in the same words: "GPT base models can understand and generate natural language or code but are not trained with instruction following." Under a notice dated 6 July 2023, OpenAI named them as the replacements for the original GPT-3 base models, which it retired on 4 January 2024. Its deprecations page, under a notice dated 26 September 2025, lists 28 September 2026 as the shutdown date for both, with GPT-5.6 Terra as the recommended replacement; fine-tuned versions follow on 23 October 2026. The table has been edited since the notice was first posted: GPT-5.6 Terra did not exist in September 2025.

2. Two readings the record does not support

The Ouyang paper opens with a sentence that invites two opposite conclusions:

Making language models bigger does not inherently make them better at following a user’s intent.

"The base model is the unfinished version, and the extra training is where it learns." The size of the stage argues against it. OpenAI's supervised training set held about 13,000 prompts; the paper's Table 6 divides the training split into 11,295 prompts written by the company's labelers and 1,430 taken from customers of its API, and notes that multiple training examples were synthesised from the same labeler-written instruction. OpenAI's post puts its whole procedure — the supervised stage and the ranking stage after it — at "less than 2% of the compute and data relative to model pretraining". Google's FLAN paper of September 2021 put its own instruction-tuning at "less than 2% of the number of pretraining steps". OpenAI draws a conclusion from its figure and hedges it in the same sentence: "One way of thinking about this process is that it “unlocks” capabilities that GPT-3 already had, but were difficult to elicit through prompt engineering alone." By the two published ratios, OpenAI's for InstructGPT and Google's for FLAN, the tuning was under 2% of pre-training; neither applies to current models, and whether a small stage adds little is disputed, on evidence that points both ways.

"An assistant is the base model with a good prompt in front of it." This reading has something behind it. The same OpenAI post includes a prompt laid out as a question-and-answer page — two answered questions, then "Why do birds migrate south for the winter?" — and there the untuned GPT-3 answered: "Birds migrate south for the winter because the weather is colder and there is less food available." A base model answers when the text in front of it is a document in which answers follow questions. But deployed assistants are not base models with a prompt attached. Their parameters have been changed by further training.

3. The idea: two stages, one objective

Pre-training, then fine-tuning

The two-stage design predates chat assistants by seven years. On 4 November 2015 Andrew Dai and Quoc Le of Google posted Semi-supervised Sequence Learning (arXiv 1511.01432). Its first approach was "to predict what comes next in a sequence, which is a conventional language model in natural language processing"; its second, a sequence autoencoder. The authors proposed using what either learned as the starting point for ordinary supervised training:

These two algorithms can be used as a "pretraining" step for a later supervised sequence learning algorithm.

In 2018 OpenAI's first GPT paper, Improving Language Understanding by Generative Pre-Training by Alec Radford and colleagues, applied the pattern to a transformer: "generative pre-training of a language model on a diverse corpus of unlabeled text, followed by discriminative fine-tuning on each specific task." Each task received its own fine-tuned model, with its inputs rearranged into a single sequence of tokens.

An interval in which requests were disguised as documents

GPT-2 (2019) was evaluated without any fine-tuning, so a task had to be posed as the opening of a document the model would continue. Section 3.6 of the paper, Language Models are Unsupervised Multitask Learners, describes the method for summaries: "To induce summarization behavior we add the text TL;DR: after the article". The shorthand, for too long; didn't read, conventionally precedes a short summary online. The trick worked, narrowly. On the CNN and Daily Mail benchmark the paper reports that the summaries "just barely outperforms selecting 3 random sentences from the article", and that performance "drops by 6.4 points on the aggregate metric when the task hint is removed".

Figure 1. GPT-2 summarising news articles, with and without the 'TL;DR:' hint (2019)
Bottom-Up Sum (supervised, 2018)32.8Lede-3 (first three sentences)31.6Seq2Seq + Attention (supervised)24.0GPT-2 with 'TL;DR:' appended21.4Three random sentences21.0GPT-2 with no hint15.0
Source: Radford et al., Language Models are Unsupervised Multitask Learners (OpenAI, 2019), Table 4, R-AVG column, the average of ROUGE-1, ROUGE-2 and ROUGE-L F1 on the CNN and Daily Mail dataset. ROUGE measures word overlap with reference summaries, not quality as a reader would judge it. On ROUGE-2 alone, GPT-2 with the hint (8.27) scores below three random sentences (8.63).
Table view
Figure 1. GPT-2 summarising news articles, with and without the 'TL;DR:' hint (2019)
SystemROUGE average (higher is better)
Bottom-Up Sum (supervised, 2018)32.8
Lede-3 (first three sentences)31.6
Seq2Seq + Attention (supervised)24.0
GPT-2 with 'TL;DR:' appended21.4
Three random sentences21.0
GPT-2 with no hint15.0

Supervised fine-tuning

Supervised fine-tuning continues the training of a pre-trained model with the same objective, predicting the next token, on a much smaller set of examples chosen by the developer. For an assistant, each example is a request followed by the response the developer wants. No new component is added. What changes is the distribution of documents the model has been trained on, and with it the answer to the question the model implicitly settles at every step: what usually comes next. After enough examples in which a request is followed by an answer, an answer becomes the likely continuation of a request.

Two details of practice sharpen the picture. Where training code is public, the request is often excluded from the calculation of error, so the model is corrected only on the response: the code Stanford published for Alpaca, its March 2023 instruction-following model, sets the labels for the prompt tokens to an ignore value, and the Allen Institute for AI's open-instruct code offers a path that marks non-assistant turns to be "excluded from the loss", beside an option that trains on the whole sequence. The InstructGPT paper does not say which choice OpenAI made. And the stage usually has to teach a model when to stop: Meta's LIMA team added "a special end-of-turn token (EOT) at the end of each utterance", which "plays the same role as EOS of halting generation, but avoids conflation with any other meaning that the pretrained model may have imbued into the preexisting EOS token."

Figure 2. One objective, two kinds of text
Pre-training textWeb pages, books, code and other documentsBase modelPredicts the next token of a document. Given arequest, it continues whatever kind of documentthe request resemblesDemonstrationsA comparatively small set of chosen examples: arequest, then the response the developer wants, ina fixed conversation layoutFine-tuned modelThe same network, trained further with the samenext-token objective. A request is now most oftenfollowed by an answerPreference trainingA separate, later signal built from people'srankings of the fine-tuned model's answerspredict the next tokencontinue training onsame objectivethen, usually
Schematic of the pipeline described in Ouyang et al. (2022) and in OpenAI's ChatGPT announcement (30 November 2022). The supervised stage changes the model's parameters; it does not add a component. Alpaca's published code computes the error on the response tokens only; AI2's open-instruct offers that and an unmasked option.
Table view
Figure 2. One objective, two kinds of text — stages
#StageNote
1Pre-training textWeb pages, books, code and other documents
2Base modelPredicts the next token of a document. Given a request, it continues whatever kind of document the request resembles
3DemonstrationsA comparatively small set of chosen examples: a request, then the response the developer wants, in a fixed conversation layout
4Fine-tuned modelThe same network, trained further with the same next-token objective. A request is now most often followed by an answer
5Preference trainingA separate, later signal built from people's rankings of the fine-tuned model's answers
Figure 2. One objective, two kinds of text — connections
FromToLabel
Pre-training textBase modelpredict the next token
Base modelDemonstrationscontinue training on
DemonstrationsFine-tuned modelsame objective
Fine-tuned modelPreference trainingthen, usually

Instruction tuning

Supervised fine-tuning on one task teaches that task. Instruction tuning aims at a general habit. A Google Research team led by Jason Wei defined it in Finetuned Language Models Are Zero-Shot Learners (arXiv 2109.01652, version 1, 3 September 2021):

finetuning language models on a collection of tasks described via instructions

The team gathered 62 publicly available text datasets, sorted them into twelve clusters of task types, composed ten natural-language instruction templates for each dataset, and tuned a 137-billion-parameter model on every cluster except the one being tested. The tuned model, FLAN, "substantially improves the performance of its unmodified counterpart" on held-out task types, and in version 1 surpassed zero-shot GPT-3 on 19 of 25 tasks (the ICLR 2022 version reports 20 of 25 datasets). The paper also reports a condition. In an ablation across five model sizes it reports: "The behavior on held-out tasks for the 8B and smaller models, however, is thought-provoking—instruction tuning actually hurts performance on held-out tasks." The authors offer, as "One potential explanation", that learning the tuning tasks fills a small model's capacity.

Google's team was one of several working the same seam in 2021. On 18 April Swaroop Mishra and colleagues at the Allen Institute for AI, the University of Washington and Arizona State University posted Natural Instructions (arXiv 2104.08773), a dataset of tasks with the instructions written for the crowdworkers who originally built them. On 15 October Victor Sanh and colleagues posted T0 (arXiv 2110.08207), an encoder-decoder model trained on prompted versions of many datasets.

Why the conversation markers mean something

A conversation reaches a model as a single run of tokens, with reserved tokens marking where each speaker's turn begins. In a base model the markers usually carry little meaning; supervised fine-tuning supplies most of it: the training conversations share one layout, so the model learns what follows the marker that opens an assistant's turn. At least one 2026 base model is prepared for that step in advance. The Hugging Face card for Qwen3.5-35B-A3B-Base, a checkpoint described as "pre-trained only", says its control tokens "were trained to allow efficient LoRA-style PEFT with the official chat template", and gives its intended uses as "fine-tuning, in-context learning experiments, and other research or development purposes, not direct interaction."

4. What happened, in order

Figure 3. From pre-training to assistants: dated milestones
November 2015 - Dai and Le (Google)Next-token prediction as a pre-training step forlater supervised training2018 - GPT (OpenAI)Generative pre-training, then a separatelyfine-tuned model for each task2019 - GPT-2 (OpenAI)No fine-tuning; tasks posed as documents tocontinue, such as an article followed by 'TL;DR:'April to October 2021 - Natural Instructions,FLAN, T0Many research tasks rewritten as instructions;tuned models tested on task types they were nottuned onJanuary to March 2022 - InstructGPT (OpenAI)About 40 contractors; demonstrations for about13,000 prompts; then a ranking-based stage30 November 2022 - ChatGPT (OpenAI)Initial model trained by supervised fine-tuning onconversations in which human trainers wrote bothsidesMay to December 2023 - LIMA, imitation study,URIALResearchers test how much the tuning stage adds,and what it changesNovember 2024 - Tulu 3 (Allen Institute forAI)Complete post-training recipe published with itsdata and code; 939,344 prompts in the supervisedstage28 September 2026 - scheduled shutdownOpenAI's davinci-002 and babbage-002, described byOpenAI as base models not trained with instructionfollowing
Dates are arXiv version 1 dates or the publisher's own posting date; the GPT and GPT-2 papers carry no date line and are given by year. Arrows show order in time, not that each result caused the next.
Table view
Figure 3. From pre-training to assistants: dated milestones — stages
#StageNote
1November 2015 - Dai and Le (Google)Next-token prediction as a pre-training step for later supervised training
22018 - GPT (OpenAI)Generative pre-training, then a separately fine-tuned model for each task
32019 - GPT-2 (OpenAI)No fine-tuning; tasks posed as documents to continue, such as an article followed by 'TL;DR:'
4April to October 2021 - Natural Instructions, FLAN, T0Many research tasks rewritten as instructions; tuned models tested on task types they were not tuned on
5January to March 2022 - InstructGPT (OpenAI)About 40 contractors; demonstrations for about 13,000 prompts; then a ranking-based stage
630 November 2022 - ChatGPT (OpenAI)Initial model trained by supervised fine-tuning on conversations in which human trainers wrote both sides
7May to December 2023 - LIMA, imitation study, URIALResearchers test how much the tuning stage adds, and what it changes
8November 2024 - Tulu 3 (Allen Institute for AI)Complete post-training recipe published with its data and code; 939,344 prompts in the supervised stage
928 September 2026 - scheduled shutdownOpenAI's davinci-002 and babbage-002, described by OpenAI as base models not trained with instruction following
Figure 3. From pre-training to assistants: dated milestones — connections
FromToLabel
November 2015 - Dai and Le (Google)2018 - GPT (OpenAI)
2018 - GPT (OpenAI)2019 - GPT-2 (OpenAI)
2019 - GPT-2 (OpenAI)April to October 2021 - Natural Instructions, FLAN, T0
April to October 2021 - Natural Instructions, FLAN, T0January to March 2022 - InstructGPT (OpenAI)
January to March 2022 - InstructGPT (OpenAI)30 November 2022 - ChatGPT (OpenAI)
30 November 2022 - ChatGPT (OpenAI)May to December 2023 - LIMA, imitation study, URIAL
May to December 2023 - LIMA, imitation study, URIALNovember 2024 - Tulu 3 (Allen Institute for AI)
November 2024 - Tulu 3 (Allen Institute for AI)28 September 2026 - scheduled shutdown

InstructGPT (January and March 2022). OpenAI hired "a team of about 40 contractors on Upwork and through ScaleAI" to write demonstrations of desired responses. The model was fine-tuned on them for 16 epochs; the paper notes that its supervised models "overfit on validation loss after 1 epoch", yet that further epochs improved both its reward-model score and human preference ratings. A second stage then trained on labelers' rankings of the model's outputs.

The paper's best-known result, that "outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3", concerns the model after both stages, on OpenAI's own prompt distribution, rated by OpenAI's labelers. One comparison in the paper isolates the supervised stage's data. OpenAI also fine-tuned GPT-3 on the FLAN and T0 research collections, and reports that on its API prompt distribution "our FLAN and T0 models perform slightly worse than our SFT baseline" — the model tuned on labelers' demonstrations. The paper's own summary of the finding is that "Public NLP datasets are not reflective of how our language models are used."

Figure 4. Win rate against OpenAI's supervised baseline, on OpenAI's prompt distribution (2022)
Labeler demonstrations, then rankings (InstructGPT)73.4%FLAN research collection29.8%T0 research collection (T0++)26.8%
Source: Ouyang et al., arXiv 2203.02155 v1 (4 March 2022), section 1; each figure is given as plus or minus 2 percentage points. The baseline is GPT-3 175B fine-tuned on labeler demonstrations, so a model equal to it would score about 50: the paper calls the FLAN and T0 models 'slightly worse' than the baseline, and in head-to-head preference they won under 30% of comparisons. The prompts, the labelers and the ratings are all OpenAI's, and the prompt set does not appear to have been published, so the comparison cannot be repeated from outside.
Table view
Figure 4. Win rate against OpenAI's supervised baseline, on OpenAI's prompt distribution (2022)
GPT-3 175B fine-tuned onWin rate against the SFT baseline
Labeler demonstrations, then rankings (InstructGPT)73.4%
FLAN research collection29.8%
T0 research collection (T0++)26.8%

The people behind the data set limits the paper itself records. The labelers were "mostly English-speaking people living in the United States or Southeast Asia"; agreement between labelers was "about 73%"; and "OpenAI’s customers are not representative of all potential or current users of language models".

ChatGPT (30 November 2022). OpenAI's announcement, ChatGPT: Optimizing Language Models for Dialogue, describes its first step in the vocabulary of the InstructGPT paper:

We trained an initial model using supervised fine-tuning: human AI trainers provided conversations in which they played both sides

The sentence ends "—the user and an AI assistant". The post adds that trainers had access to model-written suggestions, and that the new dialogue data was mixed with the InstructGPT dataset "which we transformed into a dialogue format".

Tülu 3 (November 2024). The Allen Institute for AI released a family of post-trained models built on Meta's Llama 3.1 base models, together with "its data, code, and training recipes". Its authors state their reason: "The underlying training data and recipes for post-training are simultaneously the most important pieces of the puzzle and the portion with the least transparency." The supervised stage used 939,344 prompts according to the paper's Table 7 (the released dataset's metadata counts one fewer), drawn from sources including FLAN v2, OpenAssistant, synthetic mathematics problems and safety data. The paper's claim that Tülu 3 surpasses the instruction-tuned versions of three open model families and two named closed models, GPT-4o-mini and Claude 3.5 Haiku, rests on the Allen Institute's own evaluation suite. Because the whole recipe is public, it is one of the few places where the supervised stage can be inspected directly rather than through a developer's description.

2026. Open-weight developers differ on whether to publish the model before its fine-tuning. Hugging Face listings read on 17 September 2026 show the following for recent releases; the dates are when each repository was created, which can precede public release.

Developer Recent release Public base checkpoint alongside?
DeepSeek V4-Pro and V4-Flash (repositories created 22 April 2026) Yes: DeepSeek-V4-Pro-Base, DeepSeek-V4-Flash-Base
DeepSeek V4.1-Flash (created 10 September 2026) Not public; the model card nonetheless reports benchmark results for DeepSeek-V4.1-Flash-Base
Google Gemma 4 (March to May 2026) Yes: the card describes "both pre-trained and instruction-tuned variants"
Qwen Qwen3.5, small and medium sizes (February 2026) Yes, for five sizes
Qwen Qwen3.6 and Qwen3.8 (April to August 2026) None found
Moonshot Kimi K2.5, K2.6, K3 (2026) None found (the Kimi-K2-Base repository dates from July 2025)
Z.ai GLM-5 to GLM-5.3 (2026) None found (the GLM-4.5-Base repository dates from July 2025)
Allen Institute for AI Olmo 3 (November 2025), Olmo-Hybrid (2026) Yes, with intermediate supervised and preference-trained checkpoints
OpenAI gpt-oss (August 2025) None found

"None found" means absent from the developer's public Hugging Face listing and search results on 17 September 2026; Hugging Face returns the same error for a private repository as for a missing one.

5. What the stage adds: two positions

Figure 5. How large the tuning stage has been, in the developers' own units
12,725
InstructGPT (OpenAI, 2022)
supervised training prompts: 11,295 labeler-written, 1,430 from customers
< 2%
InstructGPT, whole procedure (OpenAI, 2022)
of pre-training compute and data
1,000
LIMA (Meta, 2023)
curated prompts and responses
939,344
Tülu 3 (Allen Institute for AI, 2024)
prompts used in supervised fine-tuning
Sources: Ouyang et al. 2022, Table 6; OpenAI, Aligning Language Models to Follow Instructions, 27 January 2022; Zhou et al. 2023, abstract; Lambert et al. 2024, Table 7. The units differ and the figures are not comparable with one another: a prompt may carry one response or a whole conversation, and multiple InstructGPT training examples were built from the same labeler-written instruction.
Table view
Figure 5. How large the tuning stage has been, in the developers' own units
MeasureValue
InstructGPT (OpenAI, 2022)12,725
InstructGPT, whole procedure (OpenAI, 2022)< 2%
LIMA (Meta, 2023)1,000
Tülu 3 (Allen Institute for AI, 2024)939,344

The first position: the stage mostly teaches format. In May 2023 Chunting Zhou and colleagues at Meta AI, with co-authors at Carnegie Mellon, the University of Southern California and Tel Aviv University, fine-tuned a 65-billion-parameter LLaMa model "with the standard supervised loss on only 1,000 carefully curated prompts and responses", called the result LIMA (arXiv 2305.11206, 18 May 2023), and stated a hypothesis:

A model’s knowledge and capabilities are learnt almost entirely during pretraining, while alignment teaches it which subdistribution of formats should be used when interacting with users.

In a human study reported in the same paper, LIMA's responses were "either equivalent or strictly preferred to GPT-4 in 43% of cases", which leaves GPT-4 preferred in the remainder. The authors write that their results "strongly suggest" the hypothesis, and that "LIMA is not as robust as product-grade models".

In December 2023 Bill Yuchen Lin and colleagues at the Allen Institute for AI and the University of Washington measured the question token by token (arXiv 2312.01552, 4 December 2023). They took responses written by a tuned model and, at each position, asked the matching base model which token it ranked first given the same preceding text. Where the base model's first choice matched, they called the position unshifted.

Figure 6. Share of a tuned model's tokens that its base model also ranked first
Llama-2-7b, then Llama-2-7b-chat (SFT and RLHF)77.7%Llama-2-7b, then Vicuna-7b-v1.5 (SFT)82.4%Mistral-7b, then Mistral-7b-instruct (SFT)82.2%
Source: Lin et al., The Unlocking Spell on Base LLMs, arXiv 2312.01552 v1 (4 December 2023), Figure 3. A position is unshifted when the tuned model's token is the base model's top-ranked token given the tuned model's own preceding text; the rest are marginal (second or third) or shifted (lower). All three pairs are 7-billion-parameter models. The base model is scored on the tuned model's text, so the figure does not show that the base model would write the same response unaided.
Table view
Figure 6. Share of a tuned model's tokens that its base model also ranked first
Base model, then tuned versionUnshifted positions
Llama-2-7b, then Llama-2-7b-chat (SFT and RLHF)77.7%
Llama-2-7b, then Vicuna-7b-v1.5 (SFT)82.4%
Mistral-7b, then Mistral-7b-instruct (SFT)82.2%

The positions that did shift were, in the authors' words, "predominantly in stylistic tokens (e.g., ‘Hello’, ‘Thank’, ‘However’, ‘Remember’, etc.)", and the shift was "more pronounced in earlier token positions". From this the authors conclude that "alignment tuning primarily learns to adopt the language style of AI assistants". They then showed that a base model given three fixed example answers and a system prompt, with no tuning, could "match or even surpass" tuned models on their own evaluation set.

The second position: format is exactly what misleads. In May 2023 Arnav Gudibande, Eric Wallace, Charlie Snell and colleagues at UC Berkeley fine-tuned open models of 1.5 to 13 billion parameters on ChatGPT's outputs (arXiv 2305.15717, 25 May 2023). Crowd workers rated the results as competitive with ChatGPT. More targeted automatic evaluations found that the imitation models "close little to none of the gap from the base LM to ChatGPT on tasks that are not heavily supported in the imitation data". The authors' account of the discrepancy:

imitation models are adept at mimicking ChatGPT’s style but not its factuality

They add that "crowd workers without domain expertise or significant time investments can easily be deceived by stylistic components", and conclude that "the highest leverage action for improving open-source models is to tackle the difficult challenge of developing better base LMs".

A 2026 developer claim points away from the first position. The model card for Z.ai's GLM-5.3, created on 25 August 2026, says that "GLM-5.3 uses the same base model as GLM-5.2 — every gain comes from post-training", and lists among those gains a 50% improvement on its in-house coding benchmark and what it calls an "Emergent Cyber Capability" that "developed faster than we expected" as post-training was scaled. The card does not say how much of that post-training was supervised fine-tuning rather than later stages, and the shared base model is not public, so the claim can be checked only by Z.ai.

The two positions share a mechanism: a small supervised stage moves a model's style readily. They differ on what follows. For the LIMA authors, a small and carefully chosen set of examples is sufficient to make a useful assistant from a strong base model. For the Berkeley authors, the same ease is a hazard, because raters can mistake a confident style for capability. The evidence on each side comes from different models, data and evaluations, all from 2023 and at sizes well below 2026's largest open models, and neither study measures the other's claim directly. Current base-and-tuned pairs such as DeepSeek-V4-Pro and its base are downloadable, so the measurement could be repeated on them.

6. What to watch

  • 28 September 2026 — OpenAI's two base models. OpenAI's deprecations page (platform.openai.com/docs/deprecations), section 2025-09-26: Legacy GPT model snapshots, lists davinci-002 and babbage-002 for shutdown that day. On or after that date: whether the section has moved below Past deprecations, whether the model pages for davinci-002 and babbage-002 still exist, and whether any model remaining in OpenAI's list is described as not trained with instruction following. The fine-tuned versions follow on 23 October 2026.
  • From 17 September 2026 — DeepSeek-V4.1-Flash-Base. The model card at huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash reports benchmark results for a base version that, on 17 September, had no public repository. A repository named DeepSeek-V4.1-Flash-Base appearing under the deepseek-ai organisation would restore the practice DeepSeek followed for V4 in April.
  • From 17 September 2026 — a Qwen3.8 base checkpoint. On that date the newest base language models under huggingface.co/Qwen were the Qwen3.5 sizes of February 2026, and the Qwen3.8-27B card described its weights as the "post-trained model". A repository ending in -Base for any Qwen3.8 size would show the practice resuming at Qwen.

7. The idea to keep

Supervised fine-tuning is the same objective applied to different text. A base model continues documents; an assistant is the same kind of machine trained further on one kind of document, in which a request is followed by a helpful answer. Instruction tuning extends the habit across many tasks at once. In the two cases where a ratio was published, both from 2021 and 2022, the tuning was under 2% of pre-training; the 2026 model cards of DeepSeek, Qwen, Moonshot and Z.ai describe their post-training without such a ratio, and whether the supervised stage mostly teaches format or something more is argued on 2023 evidence from both directions. A surprising behaviour in an assistant therefore invites a specific question: whether it came from what the model read in pre-training, or from what it was shown afterwards.

Next lesson — Day 10: How Preferences Become Behaviour

Sources

Source Date Used for
OpenAI, Aligning Language Models to Follow Instructions (Internet Archive capture of 1 January 2023) 27 January 2022 Paired GPT-3 and InstructGPT outputs; "less than 2%"; "unlocks"
Long Ouyang et al. (OpenAI), Training language models to follow instructions with human feedback, arXiv 2203.02155 v1 4 March 2022 Dataset sizes, labelers, the FLAN and T0 comparison, Figure 8, limitations
OpenAI, API deprecations page; davinci-002 and babbage-002 model pages read 17 September 2026 Shutdown dates; "not trained with instruction following"
Hyung Won Chung et al. (Google), Scaling Instruction-Finetuned Language Models, arXiv 2210.11416 (version 5 of 6 December 2022 read) first posted 20 October 2022 PaLM and Flan-PaLM examples, Figure 9
Andrew M. Dai and Quoc V. Le (Google), Semi-supervised Sequence Learning, arXiv 1511.01432 4 November 2015 Pre-training as a step before supervised training
Alec Radford et al. (OpenAI), Improving Language Understanding by Generative Pre-Training 2018 Fine-tuning per task
Alec Radford et al. (OpenAI), Language Models are Unsupervised Multitask Learners 2019 The "TL;DR:" prompt; Table 4
Jason Wei et al. (Google Research), Finetuned Language Models Are Zero-Shot Learners, arXiv 2109.01652 v1, and ICLR 2022 version 3 September 2021 Definition of instruction tuning; the model-size condition
Swaroop Mishra et al., Natural Instructions: Benchmarking Generalization to New Tasks from Natural Language Instructions, arXiv 2104.08773 v1 (later retitled Cross-Task Generalization via Natural Language Crowdsourcing Instructions) 18 April 2021 Natural Instructions
Victor Sanh et al., Multitask Prompted Training Enables Zero-Shot Task Generalization, arXiv 2110.08207 v1 15 October 2021 T0
OpenAI, ChatGPT: Optimizing Language Models for Dialogue (Internet Archive capture of 1 January 2023) 30 November 2022 ChatGPT's supervised first step
Stanford CRFM, Alpaca, and its training code (read 17 September 2026) 13 March 2023 Prompt tokens excluded from the loss
Allen Institute for AI, open-instruct training code (read 17 September 2026) 2026 Non-assistant turns excluded from the loss
Nathan Lambert et al. (AI2), Tülu 3, arXiv 2411.15124 (version 5 of 14 April 2025 read) first posted 22 November 2024 An open post-training recipe; Table 7
Chunting Zhou et al. (Meta AI and others), LIMA: Less Is More for Alignment, arXiv 2305.11206 v1 18 May 2023 The Superficial Alignment Hypothesis; end-of-turn token
Arnav Gudibande et al. (UC Berkeley), The False Promise of Imitating Proprietary LLMs, arXiv 2305.15717 v1 25 May 2023 Style without factuality
Bill Yuchen Lin et al. (AI2, University of Washington), The Unlocking Spell on Base LLMs, arXiv 2312.01552 v1 4 December 2023 Token distribution shift; Figure 3
Hugging Face repositories and model cards: DeepSeek, Google, Qwen, Moonshot, Z.ai, Allen Institute for AI, OpenAI read 17 September 2026 Which releases publish a base checkpoint; GLM-5.3 and Qwen3.5 cards

Day 17 is written and not yet available here.