Imagine a musician asked to improvise. Set to very cautious, they play the most expected notes: it's correct, it's boring, and it's always the same. Set to very bold, they attempt unusual sequences: sometimes brilliant, sometimes frankly dissonant. An AI's temperature is exactly that slider between the predictable and the unexpected.
It's one of the most useful and most misunderstood settings in generative AI. It explains why the same question sometimes yields two different answers, and why certain uses demand you touch it. Let's break it down.
Reminder: how a response is produced
A language model predicts the next word, as we explained in our dedicated article. But it doesn't pick a single word: at each step, it computes a probability for every possible word.
After The sky is, it might assign 40% to blue, 15% to clear, 8% to grey, and tiny probabilities to thousands of others. One then has to be chosen, and that's where temperature comes in.
What the setting does
The temperature modifies how those probabilities are used to draw the word.
At a low temperature, near zero, the model almost systematically takes the most probable word. The gaps between candidates are accentuated: the favourite crushes the others. The result is predictable, coherent, and very close from one response to the next.
At a high temperature, the probabilities are flattened. Less obvious candidates get a real chance of being selected. The result becomes more varied, more surprising, and sometimes incoherent.
Usual values range from about 0 to 1, sometimes up to 2 depending on the model. Many systems run by default around 0.7 to 1, a compromise that gives natural responses without going off in all directions.
A temperature of zero doesn't make an AI more truthful. It makes it more predictable. If the model is wrong, a low temperature simply guarantees it will be wrong the same way every time, consistently. That's precisely what we described in our article on calibration: reliability and displayed confidence are two distinct things. Lowering the temperature improves reproducibility, not accuracy.
How to choose it
Low temperature, from 0 to 0.3. For anything requiring consistency: extracting data from a document, classifying items, producing code, following a strict format, answering a factual question. If you want the same result every run, this is it.
Medium temperature, from 0.5 to 0.8. The everyday setting: writing, explaining, conversation, summarising. Enough naturalness not to feel robotic, enough discipline to stay coherent.
High temperature, from 0.9 to 1.3. For exploration: brainstorming, hunting for unusual angles, creative writing, generating variants. This is the setting that produces ideas you wouldn't have thought of, at the cost of having to sort through them.
Beyond that, coherence degrades quickly and the text often becomes unusable.
A little-known trick
The setting doesn't have to be uniform within a project. An effective practice is to chain two temperatures: generate a dozen ideas at a high temperature, then switch to a low temperature to evaluate them, sort them, and write them up cleanly.
You thus get the best of both regimes: diversity in exploration and rigour in execution. This is, in fact, a logic close to that of reasoning models, where the system is pushed to explore varied avenues before converging.
What to remember
Temperature doesn't change what the model knows; it changes how it draws on what it knows. It's a slider between consistency and surprise, and there's no universal right setting, only a setting suited to what you're doing.
If you only use AI through a consumer-facing interface, you generally don't have access to it, and that's fine: the defaults are reasonable. But if you're building anything on an API, it's one of the first parameters to adjust, and often the one that explains why your results are either too flat or too unpredictable.