The oldest painters and the newest machines share one problem: deciding where every mark goes.
Why explain a painting machine with cave paintings?
A machine draws a picture by starting from random static and cleaning it away, little by little, until only a picture is left. Your words steer it at every step. That is the whole idea behind most of today’s image generators, a family of methods called diffusion models, and the rest of this article unpacks it slowly, with help from India’s first painters.
Those painters worked on the sandstone walls of Bhimbetka, in the Raisen district of Madhya Pradesh, about 45 km south-east of Bhopal. UNESCO, which inscribed the rock shelters as a World Heritage Site in 2003, describes paintings that appear to date from the Mesolithic period right through to the historical period. Some of the oldest are thought to be around ten thousand years old, and the archaeologist V. S. Wakankar, who reported the shelters in 1957, argued for far older dates.
Across more than 750 rock shelters on seven hills, they painted animals, dancers and hunts: bison and deer, elephants and peacocks, small figures with bows. Their palette was simple: the paintings are largely in red and white, and the red came mainly from haematite, an iron-rich mineral. Mesolithic figures often carry lines drawn across their bodies, a patterned infill you will see on every animal below.
We chose them as teachers because they make the machine’s problem easy to see. A rock painter decides where each mark goes on a wall; a machine decides what number goes in each square of a grid. Every diagram here is drawn in the Bhimbetka manner with code, not with an image generator, because a diagram has to be exactly right. Even the static in them is computed with the formula from the original research paper.
What does a picture look like to a machine?
To a machine, a picture is a grid of numbers. The squares of the grid are called pixels (short for “picture elements”), and each one stores numbers that describe its colour.
Most screens and image files describe a colour as three amounts: red, green and blue. With 8 bits for each, every amount is a whole number from 0 to 255. The bison below is 30 pixels wide and 20 tall, so to the machine it is 30 × 20 × 3 = 1,800 numbers. A picture 1,024 pixels square is more than three million.
This is the key shift in thinking. A painting machine never handles a brush or meets a bison. Its whole job is to produce a very long list of numbers that, shown on a screen, happens to look like a bison. Researchers usually rescale the numbers before training; Jonathan Ho, Ajay Jain and Pieter Abbeel, for example, mapped 0–255 onto the range −1 to 1. The real question is how a machine could ever know which numbers to choose.

How does a machine learn to paint by ruining pictures?
It learns by watching real pictures being spoiled with random noise, and practising the reverse: looking at a noisy picture and guessing exactly which noise was added. The recipe was set out clearly in 2020 by Jonathan Ho, Ajay Jain and Pieter Abbeel, in their paper on denoising diffusion probabilistic models (DDPM).
Picture a clean image of a bison. Add a sprinkle of random static, nudging every number up or down by a random amount. Then add a little more, and more again. In the DDPM paper this “forward process” runs for 1,000 steps, and by the end nothing of the bison is left: only static, the digital version of sand thrown at a wall.
Now comes the lesson. The network (a very large mathematical function with a huge number of adjustable settings) is shown one of these noisy pictures and told how noisy it is. It has to guess the noise. Its guess is compared with the noise that was really added, its settings are nudged so that the next guess is slightly better, and the whole exercise is repeated with a great many pictures.
One practical detail: during training the machine does not really add noise a thousand times in a row. The mathematics lets it jump straight to any noise level in one go, so each lesson picks a random step and makes that noisy picture directly. The step-by-step story is still the right way to imagine it.
Why guess the noise rather than the picture? Once you can guess the noise, you can subtract it. Ho and his colleagues found that this simple noise-guessing target gave their best-looking results.

How does a machine draw a new picture from pure noise?
It runs the lesson backwards. Drawing starts from a grid of pure random static. At each step the network guesses the noise in it and a portion of that noise is taken away. After enough steps, what remains is a picture that never existed before.
The starting static is nothing special. It is drawn at random, just like the training noise. But the network has only ever learned to strip noise away from real pictures, so when it strips noise from random static, the result drifts toward something that looks like a real picture.
The DDPM authors watched this happen, and reported that “large scale image features appear first and details appear last.” The composition and big shapes settle early; fur, faces and textures arrive at the end. A rock painter works much the same way, blocking in a bison’s body before adding the lines on its flank.

What is a seed?
A seed is the number that starts a computer’s random-number generator. The same seed always produces the same sequence of random numbers, so in an image generator the seed decides the starting static. Keep the seed, the words and every setting the same and, on the same system, you will usually get the same picture again. Change only the seed and you get a new composition of the same idea.
How many steps does it take?
The original DDPM used 1,000 steps. Later methods cut that sharply: in their paper on latent diffusion, Robin Rombach and his colleagues report pictures made with anywhere from 50 to a few hundred steps. Fewer steps are faster, and each method strikes its own balance between speed and quality.
How do your words steer the picture?
Words steer the cleaning. Your prompt is turned into numbers by a text encoder, and the network consults those numbers at every single step, so each bit of noise it removes moves the picture toward what you described.

What is a text encoder?
A text encoder is a language model that turns a sentence into a long list of numbers, called an embedding, which captures its meaning. Text-to-image systems typically use a pretrained language or vision–language model for this job. The numbers mean nothing to a person, but sentences with similar meanings end up with similar numbers.
How does the network “listen” to the words?
Through a mechanism called cross-attention. Rombach and his colleagues added cross-attention layers to their network so that, as it works on each part of the picture, it can look across at the prompt and weigh which words matter there. Roughly speaking, “deer” pulls on the animal and “leaping” pulls on its pose.
What does guidance strength do?
Most systems let you choose how firmly the words steer. The standard method, classifier-free guidance, comes from Jonathan Ho and Tim Salimans. One network is trained both with the prompt and, some of the time, without it; while drawing, its two guesses are compared, and the result is pushed further in the direction the prompt points.
The trade-off is simple. The authors found that increasing guidance strength decreases variety and increases the fidelity of each individual picture. Turn it down and pictures are freer and more varied but follow your words more loosely; turn it up and they follow your words closely, with less variety.
Why do many machines paint a small “latent” picture first?
Because it is far cheaper. Instead of cleaning every pixel number, many systems clean a compressed version of the picture called a latent, and a decoder turns the finished latent into the full-size image at the end.
This is latent diffusion, described by Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer in a paper first posted in December 2021. They pointed out that training diffusion models directly on pixels often consumed hundreds of GPU days (a GPU is the graphics processor that does this arithmetic), and that running them was expensive too.
Their answer was to train an autoencoder first: an encoder that squeezes pictures into a smaller grid, and a decoder that expands them back. The cleaning then happens entirely in the small space. They tried shrinking factors from 1 to 32 on each side and found that factors of about 4 to 16 struck a good balance between efficiency and detail.
In our diagram the latent is 8 × 6 and the picture 64 × 48, a factor of 8 on each side, so there are 64 times fewer places to clean. A real latent holds several numbers in each cell rather than a tiny picture, but the saving works in just this way. Wikipedia’s overview notes that text-to-image systems today are generally latent diffusion models.

Does every machine paint by removing noise?
No. Diffusion is the approach behind most of today’s image generators, but two other families are worth knowing.
- Piece by piece (autoregressive). The picture is cut into small pieces called tokens, and the machine produces them one after another, each chosen in the light of those before it, much as a language model writes a sentence word by word. Jiahui Yu and colleagues (2022) framed text-to-image generation this way, as a sequence-to-sequence problem “akin to machine translation”.
- Straight paths (flow matching and rectified flow). Diffusion’s route from noise to picture winds about; these methods train the network to follow straighter routes. Xingchao Liu, Chengyue Gong and Qiang Liu (2022) note that straight paths are the shortest between two points and can be followed in far fewer steps, in their tests sometimes just one. Yaron Lipman and colleagues (2022) introduced flow matching, a close cousin, and reported faster training and sampling.

What do the three families share?
More than they differ. All of them learn from large collections of pictures, all treat a picture as numbers, and all can be steered by words. What changes is the order in which the numbers are decided, and how many steps it takes to decide them.
How does KamZour’s Studio turn your idea into a picture?
In two stages. First your idea is written out as a detailed scene; then a painter, an AI image model of the kind described above, paints from that written scene.
When you describe something in the Studio, a prompt writer (a large language model) reads what you mean and writes a full description of the picture: who is there, where they are, the light, the mood and the composition. Our Community Standards are part of its instructions every time. The painter sees that scene and nothing else: never your account, and never a reference picture you upload. Our transparency page walks through each step with real examples.
Your chosen style joins the scene. Each style in the Style Library is written down as a description of its medium, palette, line and composition, as we explain in How a Style Is Built. In the language of this article, the style and the scene together are the words that steer every cleaning step.
Why write the scene first? Because the painter’s text encoder only understands what is on the page. “Monsoon love” is a feeling; the painter needs the terrace, the cloud, the shawl and the cranes. The more precise the words, the more precisely they steer. We change the underlying models when better ones arrive, so we describe what each step does rather than naming them.
✦ Made with KamZour
Made with KamZour
Now that you know what happens inside the machine, here is what it does with a good idea. Every picture below was painted in the KamZour Studio: press the button under any of them to open it with the style chosen and the idea already in the box.





Frequently asked questions
Does an AI image really start as random noise?
Yes, in diffusion models, which include most of today’s image generators. Drawing begins with a grid of random numbers, and a trained network removes a little of the noise at each step until a picture is left. Your words steer what the noise becomes.
What is a seed in AI image generation?
A seed is the number that starts the random-number generator, so it decides the starting static. With the same seed, prompt and settings, the same system usually produces the same picture again. Changing the seed gives a new composition of the same idea.
What does the guidance scale do?
It sets how firmly the prompt steers each cleaning step. Higher guidance follows the words more closely but gives less variety; lower guidance gives freer, more varied pictures that follow the words more loosely. The standard method, classifier-free guidance, comes from Jonathan Ho and Tim Salimans.
What is a latent diffusion model?
It is a diffusion model that does its cleaning on a small, compressed version of the picture, called a latent, rather than on every pixel. A decoder then turns the finished latent into the full-size image. The approach was described by Robin Rombach and colleagues, and most text-to-image systems now work this way.
Why does the same prompt give different pictures each time?
Because each run usually starts from different random noise, set by a different seed. The words steer the cleaning, but the starting static shapes many details, such as composition and pose. Fix the seed and the settings, and the result usually repeats.
Does KamZour say which AI model painted my picture?
No. KamZour changes models when better ones arrive, so its transparency page describes what each step does rather than naming the models. Every picture arrives in your Session Library with the scene we wrote beneath it, so you can always read exactly what was painted.
References and further reading
- UNESCO World Heritage Centre: Rock Shelters of Bhimbetka, nomination dossier (Archaeological Survey of India)
- Wikipedia: Bhimbetka rock shelters
- Wikipedia: Diffusion model
- Ho, Jain and Abbeel (2020): Denoising Diffusion Probabilistic Models
- Rombach, Blattmann, Lorenz, Esser and Ommer (2021): High-Resolution Image Synthesis with Latent Diffusion Models
- Ho and Salimans (2022): Classifier-Free Diffusion Guidance
- Yu et al. (2022): Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
- Liu, Gong and Liu (2022): Flow Straight and Fast, Learning to Generate and Transfer Data with Rectified Flow
