Field Notes

Alt text for generated images

When an image came from a model, the prompt is not the alt and the pixels are not the point. The alt carries what the image was made to do on this page.

  • alt text
  • agents
  • images

More and more of the images on the web were never photographed or drawn by hand. Someone typed a sentence, a model produced a picture, and the picture went on a page. The question of what to write in the alt text has a tempting shortcut: use the prompt. The prompt is right there, it describes the image, and it was written by a human.

It is the wrong answer for the same reason that describing the pixels is the wrong answer for a photograph. The prompt says what the author asked for. The alt should say what the image does for the reader.

Why the prompt is not the alt

Prompts are written for the model, not the reader. They are full of instructions that mean nothing to a person listening: "4k, highly detailed, dramatic lighting, trending, in the style of." They often describe things the model did not produce, or produced differently. And they describe the picture in isolation, with no idea which page it will land on.

A prompt for a hero image might be "a diverse group of people reading on different devices in a sunny room, flat vector illustration, warm palette." The alt for that image on this site's home page should be closer to "Six people reading the same page in different ways: by phone in sunlight, by keyboard, by screen reader, zoomed in, in a terminal, on a cracked tablet." Same image. The alt says what the image is arguing, because that is what the sighted reader gets from it.

The same decision as always

Ask what the image is doing here. The five contexts in the studio apply exactly as they do to a photo.

  • Decorative. A generated texture, a mood illustration beside a heading that says it all. alt="".
  • Illustrating a point. Say the point, not the composition.
  • Carrying data. A generated chart or diagram gets a one-sentence summary and the data in text next to it.
  • The content itself. On a page about the image, in a gallery, in a post about the model, the style and composition matter and a longer alt is right.

Two things that are specific to generated images

Say that it is generated when that is the point. If the page is about image generation, or if a reader would reasonably assume a photo, the alt should say "generated illustration" or "AI-generated image of." If the page is a recipe and the image is a generated picture of a kettle, it is just a kettle, and the fact that no kettle was photographed is not something the reader needs.

Do not describe what is not there. Generated images often contain artifacts: text that is not text, hands with the wrong number of fingers, a sign that says nothing. A sighted reader notices and discounts them. Alt text that faithfully describes "a sign with illegible lettering" adds noise. Describe what the image means, and if the artifact is bad enough to change the meaning, the fix is a different image, not a more careful alt.

For the pipeline

If a generator attaches alt text automatically, it has the prompt and the pixels and knows nothing about the page. The best it can do is a neutral description, which is a starting point, not a finished alt. Store it as a draft. When the image is placed, the person or agent placing it decides the real alt in context, and that is the one that ships.

This site stores a sidecar file next to every generated image with two fields: the alt to ship and a longer description for when the image carries detail. The alt is written against the page the image is for, after the image exists, and never copied from the prompt. The decision procedure for agents is in How an agent should decide on alt text.

Open the companion lab

Spotted something this note gets wrong?

Corrections and better examples are the best kind of message.

Send a correction