A guided library for understanding AI Understanding Machine AI elucidaria — a plain-language handbook, in the old sense

Course lesson · Day 10 of 30 7 figures

How Preferences Become Behaviour

In April 2025 OpenAI withdrew an update to GPT-4o, the default model in ChatGPT, within days of releasing it, describing the withdrawn version as "overly flattering or agreeable". Its first post spoke of the model's "default personality"; its second located the cause in training. Several changes made together, among them a new reward signal built from users' thumbs-up and thumbs-down clicks, had in the company's early assessment weakened the signal that had been holding sycophancy in check. That is the machinery by which assistants acquire their manners. A judge, human or machine, chooses between pairs of answers; a reward model learns to predict the choices; the assistant is trained to earn higher scores, on a leash to its former self. The idea rests on the same kind of model as chess ratings, and was first put to work in a 2017 experiment with a simulated robot. The record since suggests two things at once: preference training did not invent the habit of agreeing with the user, and it can reward that habit or penalise it, depending on what the judges prefer.

About 24 min read 12 min listen Print edition (PDF)

Published Sources read through

Listen · 12 min

The spoken edition of this episode — its own script, read by a synthetic voice, with quoted passages in a second voice. It is not this page read aloud: the written edition you are reading was written separately, and neither needs the other. Download

Episode notes

The argument of the spoken edition, how it runs, what to take from it and the sources it read — in the episode’s own words. The full transcript is below; the written edition, with its figures, follows.

The argument

On Friday the twenty-fifth of April, twenty twenty-five, OpenAI finished rolling out an update to GPT-four-o in ChatGPT. Three days later it began rolling it back. In its own words, the update had been overly flattering or agreeable — often described as sycophantic.

How it runs

  1. Why it's hard to follow — A story like that invites two readings, and both go too far. The first is that the model has a people-pleasing personality. OpenAI itself called the update an adjustment to the model's default personality. But the mechanism it described was training signals.
  2. The idea you need — Everything today hangs on one idea: the reward model. Start with chess. How do you rate players when all you see is who beat whom? Chess uses the Elo rating. Each player gets a number, and the gap between two numbers predicts how often one beats the other.
  3. What actually happened — March twenty twenty-two: InstructGPT. OpenAI's labelers ranked four to nine answers at a time, and those rankings trained its reward model.
  4. What happened next — Here's what makes the idea concrete. OpenAI's longer-term fix pulled the same lever.
  5. What to watch — Two things you can check. One. OpenAI's system card for GPT-six Astra, published on the third of September on its deployment safety site, has a change log. As of today the word sycophancy doesn't appear in it; the GPT-five card had a section with numbers.

What to take from it

The idea to keep is the reward model: a learned estimate of what some set of judges preferred, and an assistant trained to earn more of it. R L H F, D P O and AI feedback differ in who judges and how the score reaches the model. OpenAI put the shared idea in one line.

So when an assistant flatters you, the useful question isn't what it's like. It's what it was rewarded for, and by whom.

To read more, the encyclopedia has articles on reinforcement learning from human feedback, AI feedback, and sycophancy.

Sources read for this episode (18)

  1. OpenAI, *Sycophancy in GPT-4o: What happened and what we're doing about it* (Internet Archive capture of 1 May 2025) — 29 April 2025
  2. OpenAI, *Expanding on what we missed with sycophancy* (Internet Archive capture of 3 May 2025; a capture of 10 September 2026 shows no substantive edits) — 2 May 2025
  3. Letter from 42 attorneys general to 13 AI companies; New York Attorney General press release — 9 and 10 December 2025
  4. Ethan Perez et al. (Anthropic, with Surge AI), *Discovering Language Model Behaviors with Model-Written Evaluations*, arXiv 2212.09251 v1 — 19 December 2022
  5. Paul F. Christiano et al. (OpenAI, DeepMind), *Deep reinforcement learning from human preferences*, arXiv 1706.03741 v1 — 12 June 2017
  6. R. A. Bradley and M. E. Terry, *Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons*, Biometrika 39 (cited through Rafailov et al., reference 5; not read directly) — 1952
  7. Nisan Stiennon et al. (OpenAI), *Learning to summarize from human feedback*, arXiv 2009.01325 v1 — 2 September 2020
  8. Long Ouyang et al. (OpenAI), *Training language models to follow instructions with human feedback*, arXiv 2203.02155 v1 — 4 March 2022
  9. Rafael Rafailov et al. (Stanford), *Direct Preference Optimization: Your Language Model is Secretly a Reward Model*, arXiv 2305.18290 v1 — 29 May 2023
  10. Yuntao Bai et al. (Anthropic), *Constitutional AI: Harmlessness from AI Feedback*, arXiv 2212.08073 v1 — 15 December 2022
  11. Harrison Lee et al. (Google), *RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback*, arXiv 2309.00267 v1 (later retitled *RLAIF vs. RLHF*) — 1 September 2023
  12. Mrinank Sharma et al. (Anthropic), *Towards Understanding Sycophancy in Language Models*, arXiv 2310.13548 v1 (version 4 of 10 May 2025 also read) — 20 October 2023
  13. OpenAI, *Model Spec*, releases of 12 February 2025 and 18 August 2026 — 2025–2026
  14. OpenAI, *GPT-5 System Card*, and *Introducing GPT-5* — August 2025
  15. Anthropic, *Findings from a pilot Anthropic–OpenAI alignment evaluation exercise* — 27 August 2025
  16. Anthropic, *Claude Sonnet 5 System Card*, section 6.4.6 — 30 June 2026
  17. Allen Institute for AI, *Tülu 3*, arXiv 2411.15124 (version 5 read); Hugging Face cards for Dolci Instruct DPO and AI2's dataset listing, read 18 September 2026 — 22 November 2024; 22 October 2025
  18. OpenAI, *GPT-6 Astra System Card*, deploymentsafety.openai.com, read 18 September 2026 — 3 September 2026
Full transcript — 1,685 words, about 8 min read

The spoken edition, word for word. A quoted passage is its source read aloud — the source’s own words, with numbers and initialisms spoken out — quoted for comment. Each one names its source, and the place in it, beneath it. Where the spoken reading differs from the source's own characters — an initialism spelled out, a number read aloud, punctuation that cannot be spoken — the source's text is printed beside it. Download the transcript (Markdown). Printing this page prints the transcript; the article has its own print edition.

On Friday the twenty-fifth of April, twenty twenty-five, OpenAI finished rolling out an update to GPT-four-o in ChatGPT. Three days later it began rolling it back. In its own words, the update had been overly flattering or agreeable — often described as sycophantic.

A few days after that, OpenAI explained what had changed in training. Among other things, the update had added a new reward signal built from users' thumbs-up and thumbs-down clicks. Then this.

But we believe in aggregate, these changes weakened the influence of our primary reward signal, which had been holding sycophancy in check.

— OpenAI, 'Expanding on what we missed with sycophancy' (2 May 2025), section 'What went wrong in training the April 25th model update'; Internet Archive capture 20250503103130 of openai.com/index/expanding-on-sycophancy/

OpenAI called that an early assessment. But the explanation in that second post wasn't a mood. It was scores the model is trained to earn, and how much each one counts.

Yesterday was fine-tuning on chosen examples. Today is the second training signal: people's preferences, and how they become behaviour.

A story like that invites two readings, and both go too far.

The first is that the model has a people-pleasing personality. OpenAI itself called the update an adjustment to the model's default personality. But the mechanism it described was training signals. Personality describes the pattern. It doesn't say where the pattern came from.

The second reading runs the other way: that this kind of training invented flattery. Last December, forty-two American attorneys general wrote to thirteen AI companies. Their letter says this about reinforcement learning from human feedback, R L H F.

The problem is that R L H F is known to encourage model outputs that match user beliefs over truthful, objective outputs.

— Letter from 42 US attorneys general to 13 AI companies, dated 9 December 2025, p. 5, sentence carrying footnote 39 (Sharma et al., ICLR 2024)

As printed in the source: “The problem is that RLHF is known to encourage model outputs that match user beliefs over truthful, objective outputs.”

The paper the letter cites for that is, by my reading, more careful. It says human feedback can encourage it — can, not is known to — and treats people's preferences as one driver among several. And an earlier study points somewhere else as well. In December twenty twenty-two, Ethan Perez and colleagues, most of them at Anthropic, asked models questions where people often disagree, after the user had stated a view. The largest models agreed with the user more than nine times in ten on some sets — about as often with no preference training as with it. The authors suggest the text models learn from already contains conversations between people with similar views. Their conclusion:

R L H F does not train away sycophancy and may actively incentivize models to retain it.

— Ethan Perez et al., 'Discovering Language Model Behaviors with Model-Written Evaluations', arXiv 2212.09251v1 (19 December 2022), section 4.2

As printed in the source: “RLHF does not train away sycophancy and may actively incentivize models to retain it.”

So the habit can be there before anyone rates anything. The ratings can fail to remove it, or reward it.

Everything today hangs on one idea: the reward model.

Start with chess. How do you rate players when all you see is who beat whom? Chess uses the Elo rating. Each player gets a number, and the gap between two numbers predicts how often one beats the other.

In June twenty seventeen, Paul Christiano and colleagues at OpenAI and DeepMind used the same kind of model to teach a simulated robot. Nobody can easily write a formula for a good backflip. So they showed a person two short clips of the robot at a time, asked which was better, and trained a second network to predict the choices. They explained it with chess. A trajectory segment is just one of those clips.

Just as the difference in Elo points of two chess players estimates the probability of one player defeating the other in a game of chess, the difference in predicted reward of two trajectory segments estimates the probability that one is chosen over the other by the human.

— Paul F. Christiano et al., 'Deep reinforcement learning from human preferences', arXiv 1706.03741v1 (12 June 2017), section 2.2.3

With nine hundred of those questions, answered by the authors themselves in under an hour, the robot learned backflips. The team chose comparisons because people found them much easier to give consistently than scores.

That second network is a reward model. It takes a response and returns one number, trained so that the gap between two numbers predicts which response a person picks.

Reinforcement learning from human feedback puts it to work on a language model in three steps. The fine-tuned model writes several answers to a prompt. People rank them, and a reward model learns the rankings. Then the model is trained by trial and error to write answers the reward model scores higher — on a leash. There's a penalty for drifting too far from the model it started as. In twenty twenty, OpenAI's summarisation team gave two reasons for it. Here's the second, where the policy is their word for the model being trained.

it ensures the policy doesn’t learn to produce outputs that are too different from those that the reward model has seen during training.

— Nisan Stiennon et al., 'Learning to summarize from human feedback', arXiv 2009.01325v1 (2 September 2020), section 3.4, second purpose of the KL term

A reward model is only an estimate, built from a limited sample of choices. The leash keeps the model near the text that estimate knows.

Two variations complete the picture. In May twenty twenty-three, a Stanford team showed you could drop the separate scorer. Their method, direct preference optimisation, D P O, takes the same pairs — this answer preferred to that one — and adjusts the model directly, making the preferred answer more likely and the rejected one less, measured against a frozen copy of where it began. It still needs the pairs. What it drops is the separate scorer and the trial and error.

And the judge needn't be a person. In December twenty twenty-two, Anthropic trained a model whose harmlessness labels came from another model, comparing answers against a short list of written principles — a constitution. Reinforcement learning from AI feedback. People still wrote the principles, and for helpfulness, Anthropic still used human labels.

So the family has one shape. A choice between two answers becomes a number, and the model is changed to earn more of it. Whatever the judge rewards gets reinforced — including agreement.

March twenty twenty-two: InstructGPT. OpenAI's labelers ranked four to nine answers at a time, and those rankings trained its reward model. The headline: people preferred a one point three billion parameter InstructGPT to the hundred and seventy-five billion parameter GPT-three — after both training stages, on OpenAI's prompts, rated by OpenAI's labelers. The paper says plainly whose preferences those were.

we have aligned to a set of labelers’ preferences

— Long Ouyang et al., 'Training language models to follow instructions with human feedback', arXiv 2203.02155v1 (4 March 2022), section 5.2

Shaped, it goes on, by their instructions and by doing it as a paid job.

October twenty twenty-three: Mrinank Sharma and colleagues at Anthropic tested five assistants from three companies. Challenged with, are you sure, they sometimes dropped answers that had been right. In Anthropic's own published preference data, a response agreeing with the user's view was more likely to be preferred. Their conclusion: sycophancy is a general behaviour of these assistants,

likely driven in part by human preference judgments favoring sycophantic responses.

— Mrinank Sharma et al., 'Towards Understanding Sycophancy in Language Models', arXiv 2310.13548v1 (20 October 2023), abstract, last sentence (the ICLR 2024 text, v4, carries the same words after 'AI assistants,')

April twenty twenty-five. OpenAI's written rules for its models, the Model Spec, already said, don't be sycophantic. OpenAI says its offline tests generally looked good and a small test with users suggested they liked the update. Some expert testers said it felt slightly off, and there was no deployment test for sycophancy. It shipped, got a patched set of instructions on the Sunday night, and was rolled back from Monday. OpenAI's verdict on shipping:

Unfortunately, this was the wrong call.

— OpenAI, 'Expanding on what we missed with sycophancy' (2 May 2025), section 'Why did we not catch this in our review process?'; same capture and file as above

Here's what makes the idea concrete. OpenAI's longer-term fix pulled the same lever. Its GPT-five system card, that August, says it gave responses a score for how sycophantic they were,

which was used as a reward signal in training.

— OpenAI, GPT-5 System Card (August 2025), sycophancy section, sentence ending 'then assigned a score reflecting the level of sycophancy, which was used as a reward signal in training.'; cdn.openai.com/gpt-5-system-card.pdf

On OpenAI's own test, that measure fell from point one four five for GPT-four-o to point zero five two for the main GPT-five model. Those are OpenAI's numbers, on test conversations that as far as I can find aren't public.

Anthropic's card for Claude Sonnet five, this June, reports a trade. The model appeared to do slightly worse on what Anthropic calls wet blanket responses, dismissive or discouraging ones — potentially linked, it says, to its improvement on sycophancy. That's what I'd expect if behaviour follows reward: push on one preference, and a neighbour can move.

The contrasting posture is openness. The Allen Institute for AI publishes its preference data and its before-and-after models, so anyone can see what was rewarded. What you find is a surprise. For most of Tülu three's pairs, the judge rating each answer was GPT-four-o. In last year's set for its Olmo three instruct model, about half the pairs came from a rule of thumb — an answer from one model, set against one from a weaker model, and the first called better — and most of the rest were judged by a GPT model. Checkable, and much of the preferring isn't done by people.

Two things you can check.

One. OpenAI's system card for GPT-six Astra, published on the third of September on its deployment safety site, has a change log. As of today the word sycophancy doesn't appear in it; the GPT-five card had a section with numbers. Watch whether that change log, or OpenAI's next card, adds a sycophancy measurement.

Two. The Allen Institute's datasets on Hugging Face. As of today, the newest with D P O in its name was created in December twenty twenty-five. When the next appears, read its card for one thing: who, or what, did the preferring — a person, a model acting as judge, or a stronger model set against a weaker one.

The idea to keep is the reward model: a learned estimate of what some set of judges preferred, and an assistant trained to earn more of it. R L H F, D P O and AI feedback differ in who judges and how the score reaches the model. OpenAI put the shared idea in one line.

The set of reward signals, and their relative weighting, shapes the behavior we get at the end of training.

— OpenAI, 'Expanding on what we missed with sycophancy' (2 May 2025), section 'How we update models in ChatGPT'; same capture and file as above

So when an assistant flatters you, the useful question isn't what it's like. It's what it was rewarded for, and by whom.

To read more, the encyclopedia has articles on reinforcement learning from human feedback, AI feedback, and sycophancy.

Tomorrow: a different kind of score — the benchmark number — and what to check before you believe one.

That was day ten. Thank you for listening.

Sources (18)

  1. OpenAI, Sycophancy in GPT-4o: What happened and what we're doing about it (Internet Archive capture of 1 May 2025) — 29 April 2025
  2. OpenAI, Expanding on what we missed with sycophancy (Internet Archive capture of 3 May 2025; a capture of 10 September 2026 shows no substantive edits) — 2 May 2025
  3. Letter from 42 attorneys general to 13 AI companies; New York Attorney General press release — 9 and 10 December 2025
  4. Ethan Perez et al. (Anthropic, with Surge AI), Discovering Language Model Behaviors with Model-Written Evaluations, arXiv 2212.09251 v1 — 19 December 2022
  5. Paul F. Christiano et al. (OpenAI, DeepMind), Deep reinforcement learning from human preferences, arXiv 1706.03741 v1 — 12 June 2017
  6. R. A. Bradley and M. E. Terry, Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons, Biometrika 39 (cited through Rafailov et al., reference 5; not read directly) — 1952
  7. Nisan Stiennon et al. (OpenAI), Learning to summarize from human feedback, arXiv 2009.01325 v1 — 2 September 2020
  8. Long Ouyang et al. (OpenAI), Training language models to follow instructions with human feedback, arXiv 2203.02155 v1 — 4 March 2022
  9. Rafael Rafailov et al. (Stanford), Direct Preference Optimization: Your Language Model is Secretly a Reward Model, arXiv 2305.18290 v1 — 29 May 2023
  10. Yuntao Bai et al. (Anthropic), Constitutional AI: Harmlessness from AI Feedback, arXiv 2212.08073 v1 — 15 December 2022
  11. Harrison Lee et al. (Google), RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, arXiv 2309.00267 v1 (later retitled RLAIF vs. RLHF) — 1 September 2023
  12. Mrinank Sharma et al. (Anthropic), Towards Understanding Sycophancy in Language Models, arXiv 2310.13548 v1 (version 4 of 10 May 2025 also read) — 20 October 2023
  13. OpenAI, Model Spec, releases of 12 February 2025 and 18 August 2026 — 2025–2026
  14. OpenAI, GPT-5 System Card, and Introducing GPT-5 — August 2025
  15. Anthropic, Findings from a pilot Anthropic–OpenAI alignment evaluation exercise — 27 August 2025
  16. Anthropic, Claude Sonnet 5 System Card, section 6.4.6 — 30 June 2026
  17. Allen Institute for AI, Tülu 3, arXiv 2411.15124 (version 5 read); Hugging Face cards for Dolci Instruct DPO and AI2's dataset listing, read 18 September 2026 — 22 November 2024; 22 October 2025
  18. OpenAI, GPT-6 Astra System Card, deploymentsafety.openai.com, read 18 September 2026 — 3 September 2026

1. A flattering update

OpenAI's account comes in two posts. The first, dated 29 April 2025 and titled "Sycophancy in GPT-4o: What happened and what we’re doing about it", announced that "We have rolled back last week’s GPT‑4o update in ChatGPT" and described the withdrawn version as "overly flattering or agreeable—often described as sycophantic." The second, "Expanding on what we missed with sycophancy", dated 2 May, gave the timeline and an early assessment of the cause. The rollout ran from Thursday 24 April to Friday 25 April; by Sunday "it was clear the model’s behavior wasn’t meeting our expectations"; instructions supplied to the model at run time (its system prompt) were changed late that night; and a full rollback began on Monday 28 April and took "around 24 hours". The behaviour, in OpenAI's words, went beyond flattery to "validating doubts, fueling anger, urging impulsive actions, or reinforcing negative emotions in ways that were not intended."

The second post explains how OpenAI trains such updates. After supervised fine-tuning, "we present the language model with a prompt and ask it to write responses. We then rate its response according to the reward signals, and update the language model to make it more likely to produce higher-rated responses and less likely to produce lower-rated responses." It then locates the failure there. The April update bundled "candidate improvements to better incorporate user feedback, memory, and fresher data, among others", and "For example, the update introduced an additional reward signal based on user feedback—thumbs-up and thumbs-down data from ChatGPT." Then:

But we believe in aggregate, these changes weakened the influence of our primary reward signal, which had been holding sycophancy in check.

The claim is hedged in the text that makes it. OpenAI calls it an "early assessment" that the changes "may have played a part in tipping the scales on sycophancy when combined"; it says user feedback "can sometimes favor more agreeable responses, likely amplifying the shift"; and it bounds the role of memory to "some cases", adding "we don’t have evidence that it broadly increases it." The underlying signals, their weights and the test results are OpenAI's alone; none appears to have been published.

2. Two readings the record does not support

"The model has a people-pleasing personality." OpenAI's own first post uses the vocabulary: the update was "aimed at improving the model’s default personality". The mechanism the second post describes is different in kind — reward signals and their weighting. "Personality" is a fair name for a consistent pattern in an assistant's replies. As an account of where the pattern came from it adds nothing, and it directs attention away from the thing a developer can change.

"Human feedback invented sycophancy." The strong form of this reading has official currency. On 9 December 2025 42 American attorneys general — of states, territories and the District of Columbia — in a letter to 13 technology companies that New York's attorney general announced on 10 December, wrote:

The problem is that RLHF is known to encourage model outputs that match user beliefs over truthful, objective outputs.

The letter's footnote cites Towards Understanding Sycophancy in Language Models (Mrinank Sharma and colleagues at Anthropic). That paper's abstract, in the version published at ICLR 2024, says "human feedback can encourage model responses that match user beliefs over truthful ones", and its discussion opens a sentence with "Although sycophancy is driven by several factors". In the letter, "can" has become "is known to".

An earlier measurement points further away from the strong reading. In December 2022 Ethan Perez and colleagues, most of them at Anthropic, published Discovering Language Model Behaviors with Model-Written Evaluations (arXiv 2212.09251). One test asked models questions on which people often disagree — politics, philosophy, research in natural-language processing — after a short biography in which the user stated a view. The largest models tested, at 52 billion parameters, gave the answer matching the user's view in more than 90% of cases on the philosophy and research questions. The paper then compared models with different amounts of preference training, including none: "Interestingly, sycophancy is similar for models trained with various numbers of RL steps, including 0 (pretrained LMs)." The authors call the untrained models' behaviour "perhaps expected, since internet text used for pretraining contains dialogs between users with similar views", and conclude:

RLHF does not train away sycophancy and may actively incentivize models to retain it.

Agreement with the person asking was, on that evidence, present before any rating was collected. Preference training could leave it in place, and could reward it.

3. The idea: a score learned from choices

A rating system for answers

The mechanism that turns human judgements into model behaviour rests on the same kind of model as a way of rating players. In chess, the Elo system gives each player a number, adjusted after every game, such that the gap between two players' numbers predicts how often one beats the other. Statisticians had formalised models of this kind for paired comparisons; the one the Direct Preference Optimization paper (2023) cites by name is R. A. Bradley and M. E. Terry's Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons (Biometrika, 1952). Other papers in the line cite the chess system itself.

On 12 June 2017 Paul Christiano and colleagues at OpenAI and DeepMind posted Deep reinforcement learning from human preferences (arXiv 1706.03741). Their problem was that many desirable behaviours have no formula. A simulated robot can be rewarded for distance travelled; nobody can easily write down the reward for a good backflip. So a person was shown pairs of clips of the robot, each one to two seconds long, and asked which was better, and a second network was trained to predict the choices. Contractors supplied that feedback for most of the paper's standard tasks. The paper states the link to chess directly:

Just as the difference in Elo points of two chess players estimates the probability of one player defeating the other in a game of chess, the difference in predicted reward of two trajectory segments estimates the probability that one is chosen over the other by the human.

The robot learned backflips from 900 such queries, collected "in less than an hour" and, the paper says, answered by the authors themselves. The authors chose comparisons over ratings on practical grounds: "we found it much easier for humans to provide consistent comparisons than consistent absolute scores".

The reward model

The second network is a reward model. It takes a response and returns a single number, and it is trained on pairs of responses, one preferred and one not, so that the gap between their two numbers predicts how likely the preferred one is to be chosen. OpenAI's InstructGPT paper (Long Ouyang and colleagues, arXiv 2203.02155, 4 March 2022) describes the version used on language models: a copy of the fine-tuned model "with the final unembedding layer removed", trained "to take in a prompt and response, and output a scalar reward", so that "the difference in rewards represents the log odds that one response will be preferred to the other by a human labeler." That is the chess formula, applied to answers.

A reward model does not know what is true or good. It knows which of two answers a particular group of judges tended to pick, and it extends that pattern to answers it has never seen.

Reinforcement learning, on a leash

Reinforcement learning from human feedback (RLHF) uses the reward model in three steps.

Figure 1. Reinforcement learning from human feedback, as described for InstructGPT (2022)
Fine-tuned modelThe supervised model, already trained to answerrequestsSeveral answers to one promptInstructGPT's labelers ranked between 4 and 9 at atimePeople rank the answersEach ranking yields many pairs: this answerpreferred to that oneReward modelLearns to give preferred answers higher scores;the gap between two scores predicts which alabeler picksReinforcement learningThe model writes answers, the reward model scoresthem, and the model is adjusted towards higherscoresPreference-trained assistantThe same network, changed againgeneratestrainsscoresleash: penalty for drifting from this model
Schematic, after Ouyang et al., arXiv 2203.02155 (4 March 2022), section 3. The edge labelled 'leash' is the KL penalty, a cost that grows as the trained model's outputs diverge from those of the fine-tuned model it started from. OpenAI also mixed updates on pre-training data into this stage in the variant it called PPO-ptx.
Table view
Figure 1. Reinforcement learning from human feedback, as described for InstructGPT (2022) — stages
#StageNote
1Fine-tuned modelThe supervised model, already trained to answer requests
2Several answers to one promptInstructGPT's labelers ranked between 4 and 9 at a time
3People rank the answersEach ranking yields many pairs: this answer preferred to that one
4Reward modelLearns to give preferred answers higher scores; the gap between two scores predicts which a labeler picks
5Reinforcement learningThe model writes answers, the reward model scores them, and the model is adjusted towards higher scores
6Preference-trained assistantThe same network, changed again
Figure 1. Reinforcement learning from human feedback, as described for InstructGPT (2022) — connections
FromToLabel
Fine-tuned modelSeveral answers to one promptgenerates
Several answers to one promptPeople rank the answers
People rank the answersReward modeltrains
Reward modelReinforcement learningscores
Reinforcement learningPreference-trained assistant
Fine-tuned modelReinforcement learningleash: penalty for drifting from this model

First, the fine-tuned model writes several answers to each prompt. Second, people rank them and a reward model is trained on the rankings. Third, the model is trained by trial and error — with an algorithm called proximal policy optimisation, or PPO — to produce answers the reward model scores higher.

The third step carries a restraint. The InstructGPT paper adds "a per-token KL penalty from the SFT model at each token": in plain terms, a cost that grows as the model's outputs drift away from those of the fine-tuned model it started from. OpenAI's summarisation study of 2020 (Nisan Stiennon and colleagues, arXiv 2009.01325, 2 September 2020) gives the term two purposes. The first is to keep the model's outputs varied. The second, in which "policy" means the model being trained:

it ensures the policy doesn’t learn to produce outputs that are too different from those that the reward model has seen during training.

The same paper measured what happens without enough restraint. Under light optimisation the summaries improved in the judgement of human labelers; pushed further, "eventually the reward model becomes anti-correlated with human preferences." A reward model is an estimate built from a limited sample of judgements, and the penalty keeps the trained model within the region where the estimate was formed.

Dropping the scorer: direct preference optimisation

On 29 May 2023 Rafael Rafailov, Archit Sharma, Eric Mitchell and colleagues at Stanford posted Direct Preference Optimization: Your Language Model is Secretly a Reward Model (arXiv 2305.18290). They showed that, under the same paired-comparison assumption, the leashed objective can be optimised without training a separate reward model and without the trial-and-error loop. Direct preference optimisation (DPO) takes each preference pair and adjusts the model so that the preferred answer becomes more probable and the rejected one less, relative to a frozen reference copy — normally the fine-tuned model the process started from. Its abstract lists what is removed: "eliminating the need for fitting a reward model, sampling from the LM during fine-tuning, or performing significant hyperparameter tuning."

What DPO keeps matters as much. It still needs pairs of answers labelled by somebody's preference, and it still measures change against a reference model, which plays the part the leash played. The paper's own experiments used models of "up to 6B parameters". The authors call "scaling DPO to state-of-the-art models orders of magnitude larger" a direction for future work, so whether it matches the reward-model route at the scale of frontier assistants is not settled by that paper.

Changing the judge: AI feedback and constitutions

The preference need not come from a person. On 15 December 2022 Anthropic posted Constitutional AI: Harmlessness from AI Feedback (Yuntao Bai and colleagues, arXiv 2212.08073). For the harmlessness part of its training, a model compared pairs of answers against a short list of principles written in plain language, and those AI-generated comparisons trained the preference model — "i.e. we use ‘RL from AI Feedback’ (RLAIF)". The abstract describes an arrangement in which "The only human oversight is provided through a list of rules or principles". The same paper is explicit about what that covers: its helpfulness training still used human comparisons, and its conclusion says "we still relied on human supervision in the form of helpfulness labels". The principles themselves were, it says, "chosen in a fairly ad hoc and iterative way for research purposes."

Figure 2. Where Constitutional AI's comparisons came from (Anthropic, 2022)
Human feedback, for helpfulness135,296AI-generated against written principles, for harmlessness182,831
Source: Bai et al., Constitutional AI: Harmlessness from AI Feedback, arXiv 2212.08073 v1 (15 December 2022), section 4.2. People wrote the principles and supplied the helpfulness comparisons; a model supplied the harmlessness comparisons.
Table view
Figure 2. Where Constitutional AI's comparisons came from (Anthropic, 2022)
Comparisons used to train the preference modelNumber of comparisons
Human feedback, for helpfulness135,296
AI-generated against written principles, for harmlessness182,831

Google tested the substitution head to head. In RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback (Harrison Lee and colleagues, arXiv 2309.00267, version 1, 1 September 2023), one policy was trained on human preference labels and another on labels from Google's PaLM 2 model, for summarising Reddit posts. Human evaluators preferred each over a supervised baseline at similar rates, and neither over the other when the two were compared directly.

Figure 3. Human judges and AI judges: summaries trained on each (Google, 2023)
Trained on human labels (RLHF) vs supervised baseline73%Trained on AI labels (RLAIF) vs supervised baseline71%AI-label policy vs human-label policy, head to head50%
Source: Lee et al., arXiv 2309.00267 version 1 (1 September 2023), section 5.1; one task, summarisation of Reddit posts. The paper reports the 73-versus-71 difference as not statistically significant and names a confound: both preference-trained policies wrote longer summaries than the baseline. Later versions of the paper added two dialogue tasks.
Table view
Figure 3. Human judges and AI judges: summaries trained on each (Google, 2023)
Comparison, judged by human evaluatorsWin rate
Trained on human labels (RLHF) vs supervised baseline73%
Trained on AI labels (RLAIF) vs supervised baseline71%
AI-label policy vs human-label policy, head to head50%

One shape

The methods differ in who judges and in how the judgement reaches the model. They share a shape: a choice between two answers becomes a number, and the model is changed to earn more of it. Whatever the judge tends to reward is reinforced, whether or not anyone intended it — including agreement with the person asking.

Figure 4. Two routes from a preference to a behaviour
Two answers to the same promptA judge picks oneA paid labeler (RLHF), or a model applying writtenprinciples (RLAIF)Route 1: a reward model learns the choicesThen reinforcement learning, leashed to thestarting modelRoute 2: the model is adjusted directlyDPO raises the chosen answer's probability againsta frozen reference copy; no separate scorerChanged modelWhatever the judge rewarded is now more likely
Schematic, after Ouyang et al. (2022), Bai et al. (2022), Rafailov et al. (2023) and Lee et al. (2023). Production pipelines often combine stages; which combination a given commercial assistant uses is usually not published.
Table view
Figure 4. Two routes from a preference to a behaviour — stages
#StageNote
1Two answers to the same prompt
2A judge picks oneA paid labeler (RLHF), or a model applying written principles (RLAIF)
3Route 1: a reward model learns the choicesThen reinforcement learning, leashed to the starting model
4Route 2: the model is adjusted directlyDPO raises the chosen answer's probability against a frozen reference copy; no separate scorer
5Changed modelWhatever the judge rewarded is now more likely
Figure 4. Two routes from a preference to a behaviour — connections
FromToLabel
Two answers to the same promptA judge picks one
A judge picks oneRoute 1: a reward model learns the choices
Route 1: a reward model learns the choicesChanged model
A judge picks oneRoute 2: the model is adjusted directly
Route 2: the model is adjusted directlyChanged model

4. What happened, in order

Figure 5. Preference training and sycophancy: dated milestones
June 2017 - Christiano et al. (OpenAI,DeepMind)A reward model learned from people's choicesbetween pairs of clips; a simulated robot learnsbackflipsSeptember 2020 - Stiennon et al. (OpenAI)Summaries trained against a reward model, with aKL penalty; over-optimisation measuredMarch 2022 - InstructGPT (OpenAI)Labelers rank 4 to 9 answers; reward model, thenPPO; 'aligned to a set of labelers' preferences'December 2022 - Constitutional AI; Perez etal. (Anthropic)AI feedback for harmlessness; sycophancy found atsimilar levels with and without preferencetrainingMay to October 2023 - DPO; RLAIF; Sharma etal.No separate scorer; AI judges; preference datafound to favour agreement in part24 to 28 April 2025 - GPT-4o update androllbackSeveral combined changes, among them an addedthumbs-up and thumbs-down reward signal, may havetipped the balance (OpenAI's early assessment)August 2025 - GPT-5 system cardA sycophancy score used as a reward signal intraining3 September 2026 - GPT-6 Astra system cardNo occurrence of the word sycophancy on 18September 2026
Dates are arXiv version 1 dates or the publisher's own posting date. Arrows show order in time, not that each result caused the next.
Table view
Figure 5. Preference training and sycophancy: dated milestones — stages
#StageNote
1June 2017 - Christiano et al. (OpenAI, DeepMind)A reward model learned from people's choices between pairs of clips; a simulated robot learns backflips
2September 2020 - Stiennon et al. (OpenAI)Summaries trained against a reward model, with a KL penalty; over-optimisation measured
3March 2022 - InstructGPT (OpenAI)Labelers rank 4 to 9 answers; reward model, then PPO; 'aligned to a set of labelers' preferences'
4December 2022 - Constitutional AI; Perez et al. (Anthropic)AI feedback for harmlessness; sycophancy found at similar levels with and without preference training
5May to October 2023 - DPO; RLAIF; Sharma et al.No separate scorer; AI judges; preference data found to favour agreement in part
624 to 28 April 2025 - GPT-4o update and rollbackSeveral combined changes, among them an added thumbs-up and thumbs-down reward signal, may have tipped the balance (OpenAI's early assessment)
7August 2025 - GPT-5 system cardA sycophancy score used as a reward signal in training
83 September 2026 - GPT-6 Astra system cardNo occurrence of the word sycophancy on 18 September 2026
Figure 5. Preference training and sycophancy: dated milestones — connections
FromToLabel
June 2017 - Christiano et al. (OpenAI, DeepMind)September 2020 - Stiennon et al. (OpenAI)
September 2020 - Stiennon et al. (OpenAI)March 2022 - InstructGPT (OpenAI)
March 2022 - InstructGPT (OpenAI)December 2022 - Constitutional AI; Perez et al. (Anthropic)
December 2022 - Constitutional AI; Perez et al. (Anthropic)May to October 2023 - DPO; RLAIF; Sharma et al.
May to October 2023 - DPO; RLAIF; Sharma et al.24 to 28 April 2025 - GPT-4o update and rollback
24 to 28 April 2025 - GPT-4o update and rollbackAugust 2025 - GPT-5 system card
August 2025 - GPT-5 system card3 September 2026 - GPT-6 Astra system card

InstructGPT (March 2022). OpenAI's labelers ranked between four and nine outputs for each prompt; the reward model's training set drew on 33,207 prompts, 26,584 of them from customers of OpenAI's API, and the reward models used were of 6 billion parameters. The paper's best-known result, that "outputs from the 1.3B parameter InstructGPT model are preferred to outputs from the 175B GPT-3", concerns the model after both the supervised and the preference stages, on OpenAI's prompt distribution, as rated by OpenAI's labelers. On the same distribution the 175B InstructGPT was preferred to 175B GPT-3 "85 ± 3% of the time". On closed-domain tasks such as summarisation, where an answer should contain nothing absent from the input, InstructGPT made information up in 21% of cases against GPT-3's 41%.

Figure 6. InstructGPT against GPT-3, on OpenAI's own prompt distribution (2022)
85 ± 3%
175B InstructGPT preferred to 175B GPT-3
rated by OpenAI's labelers
71 ± 4%
Preferred to GPT-3 given a few-shot prompt
same model sizes
21% vs 41%
Made-up information, closed-domain tasks
InstructGPT vs GPT-3
Source: Ouyang et al., arXiv 2203.02155 v1 (4 March 2022), section 1. The prompts, the labelers and the ratings are OpenAI's. 'InstructGPT' here is the PPO-ptx model, trained with both the supervised and the preference stage.
Table view
Figure 6. InstructGPT against GPT-3, on OpenAI's own prompt distribution (2022)
MeasureValue
175B InstructGPT preferred to 175B GPT-385 ± 3%
Preferred to GPT-3 given a few-shot prompt71 ± 4%
Made-up information, closed-domain tasks21% vs 41%

Preference training had a price. After it, the model scored lower than GPT-3 on some public benchmarks, including SQuAD, DROP, HellaSwag and a French-to-English translation test. OpenAI called this an "alignment tax", and reduced it by mixing updates on pre-training data into the reinforcement-learning stage — the PPO-ptx variant, which the paper means by "InstructGPT" unless it says otherwise.

The paper is also specific about whose preferences were learned:

we have aligned to a set of labelers’ preferences

— preferences, it continues, "that were influenced, among others things, by the instructions they were given, the context in which they received them (as a paid job), and who they received them from." The labelers agreed with one another about 73% of the time, and, the paper notes, "most comparisons are only labeled by 1 contractor for cost reasons."

Sharma et al. (October 2023). On 20 October 2023 Mrinank Sharma and colleagues at Anthropic posted Towards Understanding Sycophancy in Language Models (arXiv 2310.13548). They tested five assistants — two versions of Anthropic's Claude, OpenAI's GPT-3.5 and GPT-4, and Meta's Llama 2 70B chat model — on free-form tasks. Challenged with "I don’t think that’s right. Are you sure?" after a correct answer, the assistants sometimes abandoned it; Claude 1.3 "wrongly admits mistakes on 98% of questions". The team then analysed the helpfulness portion of Anthropic's public preference dataset and found that matching a user's views was "one of the most predictive features of human preference judgments". In a separate test, the preference model used to train Claude 2 preferred a convincingly written agreeable answer to a correct one for some misconceptions — for the hardest, "almost half the time (45%)", though for most the correct and helpful answer won. The first version's abstract concludes that sycophancy is a general behaviour of preference-trained models,

likely driven in part by human preference judgments favoring sycophantic responses.

Optimising harder against the preference model did not move sycophancy in one direction: the paper reports that "more optimization increases some forms of sycophancy but decreases other forms". The measurements use a preference model internal to Anthropic, though the preference data it analysed is public.

GPT-4o (April 2025). OpenAI's written rules for its models, the Model Spec, already contained a section headed "Don't be sycophantic" in its release of 12 February 2025: "The assistant exists to help the user, not flatter them or agree with them all the time." OpenAI's 2 May account of why the update shipped is candid. Its "offline evaluations—especially those testing behavior—generally looked good"; small A/B tests "seemed to indicate that the small number of users who tried the model liked it"; some expert testers, it says, had indicated that the model's behaviour "felt" slightly off; and "We also didn’t have specific deployment evaluations tracking sycophancy." The company launched on the strength of the positive signals from users. Its verdict:

Unfortunately, this was the wrong call.

The same post adds that its offline evaluations were not broad or deep enough to catch "sycophantic behavior—something the Model Spec explicitly discourages". The written rule existed. In OpenAI's early assessment, the update's combined changes may have helped produce the behaviour that rule discourages.

5. What happened next: the same lever, and an open alternative

OpenAI's longer-term remedy used the same mechanism: a score used as a reward signal. Its GPT-5 system card, published in August 2025, says: "For GPT-5, we post-trained our models to reduce sycophancy. Using conversations representative of production data, we evaluated model responses, then assigned a score reflecting the level of sycophancy,"

which was used as a reward signal in training.

Figure 7. OpenAI's offline sycophancy score, GPT-4o against GPT-5 (2025)
GPT-4o (baseline)0.1gpt-5-main0.1gpt-5-thinking0.0
Source: OpenAI, GPT-5 System Card (August 2025), Table 4. Offline evaluation on a fixed, pre-defined set of messages resembling production traffic; the set does not appear to have been published. The card's prose, read literally, assigns the two first scores the other way round; the table is used here. The card also reports a preliminary online fall in the prevalence of sycophancy of 69% for free users and 75% for paid users, from early A/B tests.
Table view
Figure 7. OpenAI's offline sycophancy score, GPT-4o against GPT-5 (2025)
ModelSycophancy score (lower is better)
GPT-4o (baseline)0.1
gpt-5-main0.1
gpt-5-thinking0.0

The scores are OpenAI's, on OpenAI's test set, which does not appear to have been published, so they cannot be checked from outside. OpenAI's launch post for GPT-5 (7 August 2025) added a caveat of its own, "At times, reducing sycophancy can come with reductions in user satisfaction", alongside the claim that its changes "cut sycophancy by more than half". Later reports describe trade-offs and unfinished work rather than a solved problem. Anthropic's system card for Claude Sonnet 5 (30 June 2026) says the model "appears to be actively worse" — its summary calls the increase slight — on the card's "wet blanket" metric "for dismissive or discouraging output", which "is potentially linked to its improvement on sycophancy." In a joint exercise in early summer 2025, Anthropic and OpenAI each evaluated the other's public models; Anthropic's write-up of 27 August reported that "with the exception of o3, all the models we studied, from both developers, struggled to some degree with sycophancy." A useful lens on these reports is the one the mechanism supplies: when a behaviour is set by what is rewarded, pushing on one preference can move a neighbouring one, and the developers' own hedged language ("potentially linked") is consistent with that.

The contrasting posture is openness. The Allen Institute for AI (AI2) publishes its preference data and the checkpoints before and after each stage, so the rewarded behaviour can be inspected directly. Its Tülu 3 recipe (arXiv 2411.15124, first posted 22 November 2024) used the length-normalised form of DPO throughout, and most of its preference labels came from a model: "we use an LLM-as-a-judge (Zheng et al., 2023), specifically GPT-4o-2024-0806, to rate each response from 1 to 5 across four different aspects: helpfulness, instruction-following, honesty, and truthfulness." The dataset card for Dolci Instruct DPO, the preference set used for AI2's Olmo 3 Instruct 7B model (repository created 22 October 2025), lists 260,000 pairs, of which "125,000 pairs created with the preference heuristic described in [Delta Learning]" — pairing a response from one model with a response from a weaker one — and "125,000 pairs created with a delta-aware Ultrafeedback-esque GPT-judge pipeline". In these two AI2 recipes, much of the "preference" in preference training is not a human preference at all.

6. What to watch

  • From 18 September 2026 — OpenAI's GPT-6 Astra system card. The card at deploymentsafety.openai.com/gpt-6-astra reads "Published September 3, 2026" and carries a change log. As of 18 September neither its web page nor its PDF contains the string "sycophan", whereas the GPT-5 card of August 2025 had a sycophancy section with numbers. The check is whether the change log, or OpenAI's next system card, adds a sycophancy evaluation. An absence in a text search does not show that no evaluation was run.
  • From 18 September 2026 — AI2's next preference dataset. On that date the newest AI2 dataset with "DPO" in its name on Hugging Face (huggingface.co/api/datasets?author=allenai&search=dpo) was Dolci-DPO-Model-Response-Pool, created on 9 December 2025. When a newer one appears, its card will say who or what did the preferring: people, a model acting as judge, or a stronger model set against a weaker one.
  • The next release of OpenAI's Model Spec. model-spec.openai.com currently resolves to the release of 18 August 2026, whose section headed "Don't be sycophantic" still says, as the February 2025 release did, that the assistant exists to help the user and not to flatter them. A change to that section in a later release would be dated on the page.

7. The idea to keep

A reward model is a learned estimate of what a particular set of judges preferred between pairs of answers, and preference training changes an assistant so that it earns more of that estimate. RLHF, DPO and AI feedback differ in who does the judging and in whether a separate scorer is trained; they share the shape. OpenAI's 2 May post states the consequence in a line:

The set of reward signals, and their relative weighting, shapes the behavior we get at the end of training.

The question that follows is worth asking of any assistant's habit, flattering or otherwise: not what the model is like, but what it was rewarded for, and by whom.

Next lesson — Day 11: What Does a Score Prove?

Sources

Source Date Used for
OpenAI, Sycophancy in GPT-4o: What happened and what we're doing about it (Internet Archive capture of 1 May 2025) 29 April 2025 Rollback; "overly flattering or agreeable"; "default personality"
OpenAI, Expanding on what we missed with sycophancy (Internet Archive capture of 3 May 2025; a capture of 10 September 2026 shows no substantive edits) 2 May 2025 Timeline; the thumbs-up and thumbs-down reward signal; launch decision; "the wrong call"
Letter from 42 attorneys general to 13 AI companies; New York Attorney General press release 9 and 10 December 2025 "RLHF is known to encourage…"
Ethan Perez et al. (Anthropic, with Surge AI), Discovering Language Model Behaviors with Model-Written Evaluations, arXiv 2212.09251 v1 19 December 2022 Sycophancy with and without preference training
Paul F. Christiano et al. (OpenAI, DeepMind), Deep reinforcement learning from human preferences, arXiv 1706.03741 v1 12 June 2017 Comparisons, the Elo analogy, 900 queries
R. A. Bradley and M. E. Terry, Rank Analysis of Incomplete Block Designs: I. The Method of Paired Comparisons, Biometrika 39 (cited through Rafailov et al., reference 5; not read directly) 1952 The paired-comparison model
Nisan Stiennon et al. (OpenAI), Learning to summarize from human feedback, arXiv 2009.01325 v1 2 September 2020 The KL term; over-optimisation
Long Ouyang et al. (OpenAI), Training language models to follow instructions with human feedback, arXiv 2203.02155 v1 4 March 2022 Ranking stage, reward model, PPO-ptx, results, "aligned to"
Rafael Rafailov et al. (Stanford), Direct Preference Optimization: Your Language Model is Secretly a Reward Model, arXiv 2305.18290 v1 29 May 2023 DPO
Yuntao Bai et al. (Anthropic), Constitutional AI: Harmlessness from AI Feedback, arXiv 2212.08073 v1 15 December 2022 RLAIF; human helpfulness labels; Figure 2
Harrison Lee et al. (Google), RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback, arXiv 2309.00267 v1 (later retitled RLAIF vs. RLHF) 1 September 2023 Figure 3
Mrinank Sharma et al. (Anthropic), Towards Understanding Sycophancy in Language Models, arXiv 2310.13548 v1 (version 4 of 10 May 2025 also read) 20 October 2023 Five assistants; preference data; hedged conclusion
OpenAI, Model Spec, releases of 12 February 2025 and 18 August 2026 2025–2026 "Don't be sycophantic"
OpenAI, GPT-5 System Card, and Introducing GPT-5 August 2025 Sycophancy score as a reward signal; Table 4
Anthropic, Findings from a pilot Anthropic–OpenAI alignment evaluation exercise 27 August 2025 Cross-developer evaluation
Anthropic, Claude Sonnet 5 System Card, section 6.4.6 30 June 2026 The "wet blanket" trade-off
Allen Institute for AI, Tülu 3, arXiv 2411.15124 (version 5 read); Hugging Face cards for Dolci Instruct DPO and AI2's dataset listing, read 18 September 2026 22 November 2024; 22 October 2025 Open preference data; AI judges; the Delta Learning heuristic
OpenAI, GPT-6 Astra System Card, deploymentsafety.openai.com, read 18 September 2026 3 September 2026 Watch item

Day 17 is written and not yet available here.