Field Notes

How an agent should decide on alt text

A five-question decision tree an AI agent can follow to choose alt text, in the order that produces the fewest wrong answers.

  • agents
  • alt text

Agents write a lot of alt text now. Image generators attach it, page builders fill it in, code assistants add it to every img they emit. Most of it is produced the same way: look at the pixels, describe them. That produces alt text that is accurate, consistent, and wrong, because the right alt text depends on the page the image is on, and the pixels do not know what page they are on.

Here is a decision procedure that fixes most of it. Ask the questions in order and stop at the first that applies.

1. Is the image decorative?

It adds mood, spacing, or texture, and a reader who never sees it loses nothing. Backgrounds, dividers, repeated flourishes, a stock photo beside a heading that already says everything.

Use alt="". The empty string is a deliberate instruction to the screen reader: skip this. A missing attribute is the opposite: the screen reader falls back to the file name.

Heuristic for an agent: if the surrounding heading or caption says what the image shows, and the image adds no new fact, treat it as decorative.

2. Is the image the only content of a link or button?

Then the alt is not a description at all. It is the name of the control. Home, not logo. Next page, not arrow pointing right. Open settings, not gear icon.

An empty alt here is a blocker: the link has no name, and a screen reader announces it as "link" with nothing after it.

3. Is the meaning already in the surrounding text?

Do not repeat it. A product page that says "Classic stovetop kettle, 1.7 liters, $49" does not need an alt that says "classic stovetop kettle." It needs the part the text does not say: the color, the finish, the shape of the spout. If the text already says all of it, go back to question 1.

4. Does the image carry data?

Charts, graphs, diagrams, screenshots of tables. The alt is a one-sentence summary of the point: "Monthly signups rose from 120 in January to 410 in June." The data itself goes in a real table or a long description next to the image, where it can be navigated. An alt that tries to hold every number is unreadable in one breath.

5. Otherwise, describe what the image contributes on this page

Not what it contains. The question to answer is: what does a sighted reader learn from this picture, here, that they would miss without it?

The Alt Text Studio runs one image through five contexts. The curated answers there are short in a recipe, specific in a shop, pointed in a safety notice, empty on a banner, and long only in a portfolio. That spread is the goal. An agent that writes the same alt five times has not decided anything.

Rules of form

  • Never start with "image of", "picture of", or "photo of". The screen reader already said "image".
  • Never use the file name, a placeholder, or the word "image" by itself.
  • Keep it under about 150 characters. If you need more, the detail belongs in the text.
  • Write in the language of the page.

These are rules AW-ALT-001 through AW-ALT-005 in the machine-readable ruleset, and the test corpus plants each of them so you can check whether your agent catches them.

Open the companion lab

Spotted something this note gets wrong?

Corrections and better examples are the best kind of message.

Send a correction