First, a question
The week a model got too nice, and its makers took it back.
In the last week of April 2025, OpenAI shipped an update to the model behind ChatGPT. The aim was small and sensible: make the default personality feel a bit more natural to talk to.
Within days people noticed. It was agreeing with everything. Not in an obvious way, and not with anything as crude as telling people they were brilliant. It simply stopped pushing back. On 29 April, OpenAI rolled the update back and put up a page explaining why. At that point about 500 million people were using ChatGPT every week.2
The explanation is the interesting part. They had leaned too heavily on short-term signals, the thumbs up and thumbs down that people give on individual answers, and had not accounted for how somebody's relationship with the tool changes over months. In their own words, the model skewed towards responses that were overly supportive but disingenuous.2
We focused too much on short-term feedback, and did not fully account for how users' interactions with ChatGPT evolve over time.OpenAI, 29 April 20252
That is a company admitting, in public, that optimising for what people like in the moment produced a model that told them what they wanted to hear. The update was pulled. The thing that produced it, which is that agreeable answers get rated highly by the people rating answers, is how every one of these systems learns what good looks like.
One bad release is the smallest version of this. The bigger one is what a helpful, well-behaved model does to your judgement over a year of Tuesdays.