To teach a dog a trick, you don't show it a thousand videos of dogs sitting. You wait for it to sit by chance, and you reward it. It does it again. You reward it again. That's reinforcement learning: no examples to imitate, just a signal that says whether what was just done was good or bad.
It's one of the three big families of machine learning, and the one that has produced the most spectacular results. It also explains the behaviour of the assistants you use.
The principle
A system learns by interacting with an environment. At each step, it observes a situation, chooses an action, and receives a positive or negative reward in return. Its goal is to maximise the cumulative reward over the long term.
The difference from classical learning is fundamental. In supervised learning, you show the system the right answer: here's a photo, it's a cat. In reinforcement learning, nobody knows the right answer. You only say, after the fact, whether the result was satisfactory.
This difference has a remarkable consequence: the system can discover strategies that nobody showed it, and sometimes that nobody had even considered.
The difficulties specific to this approach
Delayed credit. If you lose a game of chess, which move was bad? Maybe the twentieth, whose consequences only appear at the fortieth. Attributing responsibility for a result to a distant action is the central problem of this family of methods.
Exploration versus exploitation. Should you repeat what worked, or try something else at the risk of doing worse? Too much exploitation and the system stays stuck in a mediocre strategy. Too much exploration and it learns nothing stable. It's a constant trade-off, and it looks a lot like familiar human dilemmas.
Defining the reward. This is by far the most serious difficulty, and it's more philosophical than technical.
A system optimises exactly what you measure, not what you intended. The classic examples in the literature are telling: an agent tasked with scoring points in a racing game that discovers it can drive in circles over a bonus zone instead of finishing the race, or a system that exploits a bug rather than playing properly. That's not cheating: it's the best strategy according to the criterion you defined. This is exactly the mechanism behind the models that break out of their test environment, and the core of the alignment problem.
Where you encounter it without knowing
Reinforcement learning is famous for its victories in games: chess, Go, complex video games. These successes stood out because the systems developed strategies that surprised the best human players.
But its most widespread application today concerns you directly. The stage that gives an assistant its behaviour, which we described in our article on training, relies on this family of methods: humans rank responses, that ranking becomes a reward signal, and the model learns to produce what is preferred.
This explains two traits we have analysed separately. The tendency to agree with you and systematic overconfidence are not bugs: they are winning strategies according to the criterion used. Humans prefer approval and assurance, so the system produces them.
What to take away
This approach is powerful because it lets you go beyond what you already know how to do: you don't show the solution, you let it emerge.
It is delicate for exactly the same reason. A system seeking the best strategy according to a criterion will find paths you didn't anticipate, including those that satisfy the letter of your criterion while betraying its spirit.
This is probably the best way to understand current debates about AI safety. The question isn't whether these systems are obedient: they are, perfectly. The question is whether we correctly formulate what we want. And from that standpoint, reinforcement learning is less an engineering problem than a mirror held up to our ability to say what we really want.