KamZour Journal

How Does a Machine Draw a Picture? Explained in Cave Paintings

India’s oldest painters, on the rock walls of Bhimbetka, help explain the newest painting machines: pictures as numbers, noise that is cleaned away, and words that steer every step.

KamZour ExplainsAI artHow it works
A red-ochre cave-painting figure with an hourglass body reaches a brush toward a sandstone wall; a stream of coloured static flows from the brush and settles into a pixelated bison, clean at the head and still noisy at the rump
The oldest painters meet the newest machines: a stream of static settles into a bison. Drawn with code by KamZour in the manner of the Bhimbetka rock shelters.
The oldest painters and the newest machines share one problem: deciding where every mark goes.

Why explain a painting machine with cave paintings?

A machine draws a picture by starting from random static and cleaning it away, little by little, until only a picture is left. Your words steer it at every step. That is the whole idea behind most of today’s image generators, a family of methods called diffusion models, and the rest of this article unpacks it slowly, with help from India’s first painters.

Those painters worked on the sandstone walls of Bhimbetka, in the Raisen district of Madhya Pradesh, about 45 km south-east of Bhopal. UNESCO, which inscribed the rock shelters as a World Heritage Site in 2003, describes paintings that appear to date from the Mesolithic period right through to the historical period. Some of the oldest are thought to be around ten thousand years old, and the archaeologist V. S. Wakankar, who reported the shelters in 1957, argued for far older dates.

Across more than 750 rock shelters on seven hills, they painted animals, dancers and hunts: bison and deer, elephants and peacocks, small figures with bows. Their palette was simple: the paintings are largely in red and white, and the red came mainly from haematite, an iron-rich mineral. Mesolithic figures often carry lines drawn across their bodies, a patterned infill you will see on every animal below.

We chose them as teachers because they make the machine’s problem easy to see. A rock painter decides where each mark goes on a wall; a machine decides what number goes in each square of a grid. Every diagram here is drawn in the Bhimbetka manner with code, not with an image generator, because a diagram has to be exactly right. Even the static in them is computed with the formula from the original research paper.

What does a picture look like to a machine?

To a machine, a picture is a grid of numbers. The squares of the grid are called pixels (short for “picture elements”), and each one stores numbers that describe its colour.

Most screens and image files describe a colour as three amounts: red, green and blue. With 8 bits for each, every amount is a whole number from 0 to 255. The bison below is 30 pixels wide and 20 tall, so to the machine it is 30 × 20 × 3 = 1,800 numbers. A picture 1,024 pixels square is more than three million.

This is the key shift in thinking. A painting machine never handles a brush or meets a bison. Its whole job is to produce a very long list of numbers that, shown on a screen, happens to look like a bison. Researchers usually rescale the numbers before training; Jonathan Ho, Ajay Jain and Pieter Abbeel, for example, mapped 0–255 onto the range −1 to 1. The real question is how a machine could ever know which numbers to choose.

Diagram on a sandstone wall: a bison made of 30 by 20 ochre pixels, with a 3 by 3 block enlarged to show each pixel’s red, green and blue values, such as R 138, G 46, B 27 for dark ochre and R 226, G 201, B 160 for bare rock
A picture is a grid of numbers. The enlarged pixels show the exact values used in this drawing. Drawn with code by KamZour in the manner of the Bhimbetka rock paintings.

How does a machine learn to paint by ruining pictures?

It learns by watching real pictures being spoiled with random noise, and practising the reverse: looking at a noisy picture and guessing exactly which noise was added. The recipe was set out clearly in 2020 by Jonathan Ho, Ajay Jain and Pieter Abbeel, in their paper on denoising diffusion probabilistic models (DDPM).

Picture a clean image of a bison. Add a sprinkle of random static, nudging every number up or down by a random amount. Then add a little more, and more again. In the DDPM paper this “forward process” runs for 1,000 steps, and by the end nothing of the bison is left: only static, the digital version of sand thrown at a wall.

Now comes the lesson. The network (a very large mathematical function with a huge number of adjustable settings) is shown one of these noisy pictures and told how noisy it is. It has to guess the noise. Its guess is compared with the noise that was really added, its settings are nudged so that the next guess is slightly better, and the whole exercise is repeated with a great many pictures.

One practical detail: during training the machine does not really add noise a thousand times in a row. The mathematics lets it jump straight to any noise level in one go, so each lesson picks a random step and makes that noisy picture directly. The step-by-step story is still the right way to imagine it.

Why guess the noise rather than the picture? Once you can guess the noise, you can subtract it. Ho and his colleagues found that this simple noise-guessing target gave their best-looking results.

Diagram: a clean bison picture becomes noisier across five panels until it is pure noise; below, a noisy picture goes to a cave-style figure labelled the network, which produces its guess of the noise, compared with the noise really added
Training: spoil real pictures with noise, then learn to guess the noise. The noisy panels use the same formula as the DDPM paper. Drawn with code by KamZour in the manner of the Bhimbetka rock paintings.

How does a machine draw a new picture from pure noise?

It runs the lesson backwards. Drawing starts from a grid of pure random static. At each step the network guesses the noise in it and a portion of that noise is taken away. After enough steps, what remains is a picture that never existed before.

The starting static is nothing special. It is drawn at random, just like the training noise. But the network has only ever learned to strip noise away from real pictures, so when it strips noise from random static, the result drifts toward something that looks like a real picture.

The DDPM authors watched this happen, and reported that “large scale image features appear first and details appear last.” The composition and big shapes settle early; fur, faces and textures arrive at the end. A rock painter works much the same way, blocking in a bison’s body before adding the lines on its flank.

Diagram: five panels from pure noise to a finished deer, labelled start, step 10 of 50, step 25 of 50, step 40 of 50 and step 50, with a row of fifty tally marks beneath
Drawing: start from noise chosen by a seed, and clean it step by step. Fifty steps are shown as an example. Drawn with code by KamZour in the manner of the Bhimbetka rock paintings.

What is a seed?

A seed is the number that starts a computer’s random-number generator. The same seed always produces the same sequence of random numbers, so in an image generator the seed decides the starting static. Keep the seed, the words and every setting the same and, on the same system, you will usually get the same picture again. Change only the seed and you get a new composition of the same idea.

How many steps does it take?

The original DDPM used 1,000 steps. Later methods cut that sharply: in their paper on latent diffusion, Robin Rombach and his colleagues report pictures made with anywhere from 50 to a few hundred steps. Fewer steps are faster, and each method strikes its own balance between speed and quality.

How do your words steer the picture?

Words steer the cleaning. Your prompt is turned into numbers by a text encoder, and the network consults those numbers at every single step, so each bit of noise it removes moves the picture toward what you described.

Diagram: the words “a red deer, leaping” go to a cave figure labelled text encoder, which produces a list of numbers; dashed arrows carry the numbers to each of four cleaning steps that end in a deer
Words steer the cleaning. The numbers are illustrative; a real embedding is far longer. Drawn with code by KamZour in the manner of the Bhimbetka rock paintings.

What is a text encoder?

A text encoder is a language model that turns a sentence into a long list of numbers, called an embedding, which captures its meaning. Text-to-image systems typically use a pretrained language or vision–language model for this job. The numbers mean nothing to a person, but sentences with similar meanings end up with similar numbers.

How does the network “listen” to the words?

Through a mechanism called cross-attention. Rombach and his colleagues added cross-attention layers to their network so that, as it works on each part of the picture, it can look across at the prompt and weigh which words matter there. Roughly speaking, “deer” pulls on the animal and “leaping” pulls on its pose.

What does guidance strength do?

Most systems let you choose how firmly the words steer. The standard method, classifier-free guidance, comes from Jonathan Ho and Tim Salimans. One network is trained both with the prompt and, some of the time, without it; while drawing, its two guesses are compared, and the result is pushed further in the direction the prompt points.

The trade-off is simple. The authors found that increasing guidance strength decreases variety and increases the fidelity of each individual picture. Turn it down and pictures are freer and more varied but follow your words more loosely; turn it up and they follow your words closely, with less variety.

Why do many machines paint a small “latent” picture first?

Because it is far cheaper. Instead of cleaning every pixel number, many systems clean a compressed version of the picture called a latent, and a decoder turns the finished latent into the full-size image at the end.

This is latent diffusion, described by Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser and Björn Ommer in a paper first posted in December 2021. They pointed out that training diffusion models directly on pixels often consumed hundreds of GPU days (a GPU is the graphics processor that does this arithmetic), and that running them was expensive too.

Their answer was to train an autoencoder first: an encoder that squeezes pictures into a smaller grid, and a decoder that expands them back. The cleaning then happens entirely in the small space. They tried shrinking factors from 1 to 32 on each side and found that factors of about 4 to 16 struck a good balance between efficiency and detail.

In our diagram the latent is 8 × 6 and the picture 64 × 48, a factor of 8 on each side, so there are 64 times fewer places to clean. A real latent holds several numbers in each cell rather than a tiny picture, but the saving works in just this way. Wikipedia’s overview notes that text-to-image systems today are generally latent diffusion models.

Diagram: a noisy 8 by 6 latent grid is cleaned into a blocky clean latent, then a small cave figure labelled decoder turns it into a full 64 by 48 picture of a deer
Working small: clean a compressed latent, then decode it into the full picture. Drawn with code by KamZour in the manner of the Bhimbetka rock paintings.

Does every machine paint by removing noise?

No. Diffusion is the approach behind most of today’s image generators, but two other families are worth knowing.

Diagram with three rows: cleaning, where noise becomes a deer over many steps; piece by piece, where the deer is filled in one patch at a time; and straight path, where a straight arrow with two tick marks runs from noise to the deer beside a dotted winding path
Three ways a machine can paint: clean everything at once, build it piece by piece, or follow a straighter path. Drawn with code by KamZour in the manner of the Bhimbetka rock paintings.

What do the three families share?

More than they differ. All of them learn from large collections of pictures, all treat a picture as numbers, and all can be steered by words. What changes is the order in which the numbers are decided, and how many steps it takes to decide them.

How does KamZour’s Studio turn your idea into a picture?

In two stages. First your idea is written out as a detailed scene; then a painter, an AI image model of the kind described above, paints from that written scene.

When you describe something in the Studio, a prompt writer (a large language model) reads what you mean and writes a full description of the picture: who is there, where they are, the light, the mood and the composition. Our Community Standards are part of its instructions every time. The painter sees that scene and nothing else: never your account, and never a reference picture you upload. Our transparency page walks through each step with real examples.

Your chosen style joins the scene. Each style in the Style Library is written down as a description of its medium, palette, line and composition, as we explain in How a Style Is Built. In the language of this article, the style and the scene together are the words that steer every cleaning step.

Why write the scene first? Because the painter’s text encoder only understands what is on the page. “Monsoon love” is a feeling; the painter needs the terrace, the cloud, the shawl and the cranes. The more precise the words, the more precisely they steer. We change the underlying models when better ones arrive, so we describe what each step does rather than naming them.

✦ Made with KamZour

Made with KamZour

Now that you know what happens inside the machine, here is what it does with a good idea. Every picture below was painted in the KamZour Studio: press the button under any of them to open it with the style chosen and the idea already in the box.

A fluffy ginger cat in a jewelled turban with a heron plume sits on a golden throne, while cat courtiers in robes bow with a platter of fish and a small mouse in a robe steps forward holding a petition
The cat who holds court, in the Deccani Miniature style. Painted with KamZour Imagine in Masterpiece mode.
✦ Make art like this in Imagine · Deccani Miniature
A small steam train with maroon carriages crosses a tall curved stone viaduct beside a thundering waterfall in green monsoon mountains, with a rainbow, white egrets and rice terraces below
The monsoon train, in the Painterly Gouache Animation Storybook style. Painted with KamZour Imagine in Masterpiece mode.
✦ Make art like this in Imagine · Painterly Gouache Animation Storybook
A grown couple stand close together on a white marble palace terrace at dusk, he in a white jama and turban with a saffron shawl, she in a long red dress and veil resting her head on his shoulder, as storm clouds and cranes pass over green hills
The first cloud of the monsoon, in the Kangra Painting style. Painted with KamZour Imagine in Masterpiece mode.
✦ Make art like this in Imagine · Kangra Painting
A great patterned tree on a midnight-blue ground full of bright fish climbing toward a full moon, while parrots float on a lily pad, an owl swims in the river, and a tortoise and a deer look on
The night the fish climbed the tree, in the Gond Painting style. Painted with KamZour Imagine in Masterpiece mode.
✦ Make art like this in Imagine · Gond Painting
By firelight in a rock shelter, a grey-haired woman painter in a woven wrap and cloak paints a large red bison on the wall while a small round robot beside her copies it with its own brush
The new apprentice, in the Midcentury Geometric Fairytale style. Painted with KamZour Imagine in Masterpiece mode.
✦ Make art like this in Imagine · Midcentury Geometric Fairytale

Open the Studio More in Exclusive →

Frequently asked questions

Does an AI image really start as random noise?

Yes, in diffusion models, which include most of today’s image generators. Drawing begins with a grid of random numbers, and a trained network removes a little of the noise at each step until a picture is left. Your words steer what the noise becomes.

What is a seed in AI image generation?

A seed is the number that starts the random-number generator, so it decides the starting static. With the same seed, prompt and settings, the same system usually produces the same picture again. Changing the seed gives a new composition of the same idea.

What does the guidance scale do?

It sets how firmly the prompt steers each cleaning step. Higher guidance follows the words more closely but gives less variety; lower guidance gives freer, more varied pictures that follow the words more loosely. The standard method, classifier-free guidance, comes from Jonathan Ho and Tim Salimans.

What is a latent diffusion model?

It is a diffusion model that does its cleaning on a small, compressed version of the picture, called a latent, rather than on every pixel. A decoder then turns the finished latent into the full-size image. The approach was described by Robin Rombach and colleagues, and most text-to-image systems now work this way.

Why does the same prompt give different pictures each time?

Because each run usually starts from different random noise, set by a different seed. The words steer the cleaning, but the starting static shapes many details, such as composition and pose. Fix the seed and the settings, and the result usually repeats.

Does KamZour say which AI model painted my picture?

No. KamZour changes models when better ones arrive, so its transparency page describes what each step does rather than naming the models. Every picture arrives in your Session Library with the scene we wrote beneath it, so you can always read exactly what was painted.

References and further reading

  1. UNESCO World Heritage Centre: Rock Shelters of Bhimbetka, nomination dossier (Archaeological Survey of India)
  2. Wikipedia: Bhimbetka rock shelters
  3. Wikipedia: Diffusion model
  4. Ho, Jain and Abbeel (2020): Denoising Diffusion Probabilistic Models
  5. Rombach, Blattmann, Lorenz, Esser and Ommer (2021): High-Resolution Image Synthesis with Latent Diffusion Models
  6. Ho and Salimans (2022): Classifier-Free Diffusion Guidance
  7. Yu et al. (2022): Scaling Autoregressive Models for Content-Rich Text-to-Image Generation
  8. Liu, Gong and Liu (2022): Flow Straight and Fast, Learning to Generate and Transfer Data with Rectified Flow
About this article. Written by KamZour Editorial, the team behind KamZour, an art discovery and creation platform from India. Facts are checked against the sources listed above; the pictures are original, made with KamZour Imagine. Read our Community Standards or our philosophy.