Skip to content

How does an AI create an image? The answer is stranger than you think.

She doesn't draw. She starts from a screen of TV static and strips away the noise, again and again, until an image emerges. An explanation of the process that upended visual creation.

Advertisement
The starting analogy 🗿
Michelangelo is credited with the idea that the sculpture already exists within the block of marble, and that all that's needed is to remove the excess to reveal it. AI image generators work strangely in this same way. They don't lay down strokes on a blank page: they start from a block of chaos, a screen of TV static, and progressively remove the disorder until an image emerges. This process is called diffusion, and it's as counter-intuitive as it is effective.

You type a description, and a few seconds later, an image appears. The result is so seamless that you'd naturally imagine a machine drawing, stroke by stroke, like a very fast illustrator. The reality is completely different, and far more surprising. Here's what actually happens.

The principle: learning to remove noise

Most modern image generators rely on diffusion models. To understand how they work, you first need to understand how they're trained, because that's where the trick lies.

During training, millions of real images are taken and progressively destroyed by adding noise—that is, random pixels—step by step, until nothing remains but grey chaos with no information at all. The model is then asked to learn the reverse operation: from a slightly noisy image, guess what it looked like just before. Repeated billions of times, this exercise teaches it a very particular skill: denoising.

Then comes the magic. To generate a new image, you don't start from any photo. You give it pure random noise, a screen of static, and ask it to denoise as it has learned to do. The model removes a bit of disorder, then a bit more, and with each pass a shape comes into focus. After a few dozen steps, a coherent image emerges from the chaos—an image that has never existed before.

Where does your description come in?

At this point, a question arises: if the model only denoises, how does it know it should produce a cat astronaut rather than a mountain landscape?

That's the role of your text, which acts as a guide at every step of the denoising. Your description is converted into a mathematical representation, and at each pass, the model steers its denoising in the direction that best matches that description. Like a sculptor with a precise mental image who removes material accordingly. This ability to link words to visual forms comes from multimodality, the mechanism we explained in our article on multimodal AIs.

Why the same prompt gives different images 🎲
It often puzzles people: why do you get a different image every time with exactly the same text? Because the starting point—that initial random noise—changes each time. This is what's called the seed. A different starting noise leads to a different image, even with the same description. If you fix the same seed and the same description, you get exactly the same image. It's the randomness of the starting point that creates variety, not creativity in any human sense.

Why hands were a problem for so long

You might remember AI images with six-fingered hands, which became a meme. This famous flaw is explained very well by how this works. The model has no conceptual knowledge of the human body; it has never learned that a hand has five fingers. It has only learned what a region of pixels corresponding to a hand looks like statistically. But hands appear in a thousand different positions, often partially hidden, which makes their statistical regularity very fuzzy.

Recent models have largely fixed this flaw, thanks to more data and better architectures. But the anecdote remains illuminating: it reminds us that these systems manipulate visual regularities, not an understanding of the world. This same gap also explains why they sometimes struggle with text inside an image, or with the physical coherence of a complex scene.

What to take away

A generative AI image model doesn't draw; it reveals. It starts from a chaos of pixels and sculpts it progressively, guided by your description, until an image that has never existed emerges. It's a statistical process, not a creative act in the human sense, but its output has become so convincing that the distinction is now invisible to the naked eye, as shown by the professional production models we described with Seedream and Seedance.

Understanding this mechanism changes how you use these tools. It explains why the precision of your description matters so much, why chance plays a role, and why certain things remain hard to achieve. And it raises, implicitly, a question we've explored elsewhere: when a machine brings an image out of noise by following regularities learned from millions of human works, where does creation begin, and where does restitution end?

Advertisement