Search engines have stopped answering with text alone. Google AI Overviews, ChatGPT, Perplexity, and Gemini weave images, charts, and clips into their answers, and the model decides which ones to show. The visual asset itself is a separate signal here, not just a decoration for the text.
Models are getting better at recognizing images, but in search they rarely analyze every pixel on a page, because at the scale of the whole internet that is too costly and slow. Most often they reach a graphic through the layer of text and metadata: alt text, caption, transcript, structured data, and the surrounding text, and it is that layer that decides whether the material even enters consideration.
Full machine “vision” does happen, especially when an image goes straight to the model, for example uploaded in a query or through Google Lens. That is why, for now, solid attributes and a text layer are the safer route, and classic image SEO, focused on speed and thumbnails, is no longer enough.
Optimizing visual content for AI has two goals: getting the model to show your image or clip in an answer, and getting it to cite what they convey. The same layer of text and data that describes the visual leads to both.