Temperature 0.7
Last week, I started using and learning about language models. Here is what temperature taught me.

I have a silver medal from the International Olympiad in Informatics and I confess that last week, for the first time, I typed a sentence into a language model and hit enter.
I want to take a second to reflect on how unusual that is. I spent my adolescence writing code to solve three algorithmic problems in 5 hours; I’m as familiar with segment trees and dynamic programming as most people are with their mobile phone numbers; and still, I hadn’t touched what might be the most interesting output, at least recently, of computers: AI. People I trained with, from camps and on other teams, were already regularly using these tools. Labs that built those models which achieved gold in IOI had people who did Olympiads ten years ago. I sat it all out.
Years of training in Olympiads taught me to depend on my own thinking ability, producing code alone from a blank editor taught me that I don’t need a machine to write code for me. I’m not sure now, as to whether it was integrity or fear. When headlines came of AI solving competitive programming problems at a level comparable to legendary grandmasters, I read the first paragraph, closed the tab and didn’t read about it for another two years. So when I decided to learn how AI works for the first time last week, it was in the way a child goes to the dentist, late braced and expecting to hear what I didn’t want to hear. Instead, what I learned about was, what is for me, a rather interesting number.
If you use a language model through an API - not the chat window that most people see, but the interface underneath - there’s a parameter called temperature. It’s a singular decimal number, typically between 0 and 2, and it controls something that I’ve come to think is one of the more quietly profound ideas I’ve encountered this year.
Every time a model produces a token (a chunk of text, usually a short word or a common part of words), it first computes probabilities of every candidate token. Thousands of them. The odds of most are near zero; a handful are plausible; one is usually the favourite. Temperature controls how you pick from that distribution.
At temperature 0, you always pick the favourite, the most probable token always. The boring option. The output is deterministic, characteristically flat, safe and expected. It never surprises you.
At temperature 0.7, the distribution flattens and the model occasionally starts picking the second or third most likely token. It reads slightly more like a person.
At temperature 1.5, unlikely tokens start to get picked - a strange, sometimes brilliant, sometimes broken output is produced. A complex sentence structure with a brilliant metaphor that wanders off into nothing.
At temperature 2.0, its mostly noise.
It’s a line of maths: a dial with “expected” on one end and “chaos” on the other. I read this, sat down and realised something about my competitive programming career, schooling life and future trajectory.
I won a medal at temperature 0. Let me paint a picture of what kind of competitor I was. In one word, I was reliable. I had a vaster library of standard techniques than most other people: DP optimization patterns, data structures tricks, regret greedy labelling and so much more that was so well practised that I could identify these patterns within an instant. In training and local Olympiads, where the problems were standard and typically involved some common trick, my performance was excellent. In every contest, I read the problems, classified it as some variant of a trick I’d seen, implemented and solved problems. I almost never missed easy problems, but also struggled with weirdly wacky and wonderfully interesting ad hoc problems.
When I came to IOI 2026, I was conflicted. Always at IOI, there’s controversy and that year, there were questions regarding the “heuristic” or unusual nature of the questions. After the contest, I learned that people worse than me, who had done less training, who had smaller libraries performed better - tried something unusual and received more points for it. It didn’t make sense to me. How were they solving things I couldn’t touch? Had I simply not seen enough, not done enough training?
My mentor at the time diagnosed the issue for me: “You’re overfitted on standard problems. You struggle with things you aren’t familiar with.” I nodded, but didn’t quite understand.
I was operating with a temperature set at zero. At every step, I made the highest probability move, the obvious step that problemsetters have anticipated and defeated. Not many new ideas came out. Thus, I struggled on problems with unfamiliar ideas.
My teammates were the opposite: they ran at high temperatures, high variance performances, exploring weird and wacky avenues and often coming up with somewhat cursed unintended solutions. There’s a bitter irony here about randomness beating deterministic solutions within competitive programming.
It’s curious that a number between 0 and 2 can result in something like this, but once I learnt this vocabulary, I can’t stop seeing it, and there’s no place that it fits better than high school.
School is a process generating students. It has a distribution of outcomes, a loss function, feedback loop and reinforcement learning, and a temperature. And most education structurally runs at zero. A marking rubric is a description of the answer each model should fit to, a curriculum is the highest probability path through a subject. A good student becomes defined as one who outputs the expected output, with low variance.
Now, consider this system, and drop a large language model with temperature set to zero into it. It becomes the model student: someone who always outputs the expected answer. Students hand in the learned, expected value and markers recognise the modal answer and award full marks. Nobody learns anything, and every metric says otherwise.
That, to me, is what the panic around AI in education is after stripping away the discussion of cheating. A zero temperature machine can replicate the function of a zero temperature student. What’s the point of education if it isn’t teaching creativity, but rather how to serve as an advanced search engine when chatGPT already exists?
Bans and detectors serve to protect the mode, to make students keep producing zero temperature answers by hand. The alternative is to turn the system’s temperature up. Assess for variance and build assessments that reward high temperatures. This is harder to mark. Rewarding the unexpected has to tolerate the experiments that fail - and I’ve never seen a school that is forgiving of that. The model hasn’t changed that kids with high temperatures survive through gaps in school like Olympiads and not because of it. It only made everything within school life worthless, so that high temperature is really all that’s left.
So we come to the obvious experiment now: if I’ve spent all this time on zero temperature, what happens when I turn it up to 0.7.
Well, in literature, I’ve started exploring wonderful niche alternative readings of texts; in maths, I’ve started learning about AI which prompted this whole post. Indeed, this post itself is an instance of my higher temperature experiment: consider this post to be my version of the zero temperature, experimental version of the generic post “AI is a tool that one should use wisely and not be distracted by”.
This experiment costs, of course. I’ve gone red-faced after making some outrageous, properly wrong claims - in ways that my 16 year old self would’ve been ashamed to see. And unlike a model, I don’t have recovery mechanisms that steers back to coherency; I have to make my own way back.
However, the conversations that followed incorrect, interesting pathways have been some of the most interesting of the past week. A friend I hadn’t talked to in a year appeared in a group chat that was slowly dying as people moved their questions to AI instead of each other, and we talked about two wrong ideas to the same problem. Between our two ideas was the right one.
None of this happens at zero temperature. At zero, the expected thing happens, everyone nods, mechanically, and moves on. It’s efficient - a methodology that works on robots. It’s also, I’ve come to realise, a way of not truly being present yourself; a way of losing yourself to simply emitting the mode and calling it participation. I did that and I’m not sure if I was ever in the room.