Google patented a way of answering questions where the best article loses to the best photo
7 July 2026 · Content, Technical
Every piece of GEO advice you have read this year is about text. Answer-shaped paragraphs, entity consistency, schema, structure. All of it assumes that when an AI picks a source to cite, it is running some kind of text competition, and your job is to enter the best text.
A Google patent application published in April 2026 describes a growing class of searches where that assumption is wrong. When someone searches with a photo and a question, the system it describes does not go looking for the best article. It goes looking for the best matching image, and whichever page owns that image supplies the answer.
“Google doesn’t go looking for the best article. It goes looking for the matching photo, and whoever owns it supplies the answer.”
The text on your page does not win the retrieval. It gets carried along because it happens to live next to the winning image. Your image is the ticket into the answer; your text is the plus-one.
What Google actually filed
The patent application is called Visual Citations for Information Provided in Response to Multimodal Queries (US20260111481A1). It is assigned to Google and the named inventors include Google Search staff.
Two dates sit on the filing, and they look odd next to each other, so here is what they mean. The priority date is May 2023 (the date Google first filed the invention; the “we thought of this first” timestamp, meaning everything described here existed inside Google by then). The publication date is 23 April 2026 (applications stay private at first and the patent office publishes them on its own schedule, so this is just the paperwork becoming public, not Google doing anything new). And this particular document is a fresh version of the application, filed in December 2025, carrying the original 2023 date. That refiling is the telling part. Google had this idea before AI Overviews had properly rolled out, and was still paying patent lawyers to pursue it two and a half years later. Companies do not do that for ideas they have shelved.
A multimodal query (a search made with more than one type of input, in this case a photo plus a written or spoken question) is the Lens and AI Mode pattern: point your camera at something, ask a question about it, get an AI-written answer with sources.
The abstract opens with the whole mechanism in one line:
“A result image is retrieved based on a similarity between a query image and the result image.”
In English: when you search with a photo (the query image), the thing Google fetches is another image (the result image), chosen because it looks like yours. Not a page. An image.
Blindingly obvious so far, I know. “Search with a picture, Google looks at pictures” is not a revelation. Keep reading, because the interesting part is what never happens next.
Claim 1 spells out the full sequence, and the order of operations is the whole story. Here is what the filing actually says at each step, with my translation after each quote:
- “Obtaining… a query image and a prompt associated with the query image.” Your photo plus your question. The photo of the broken part, plus “what is this and how do I fix it?” (Nothing clever yet.)
- “Generating… an intermediate representation based on the query image.” The photo gets turned into an embedding (a long list of numbers that captures what the image looks like, so two similar images end up with similar numbers).
- “Determining… a result image based on the intermediate representation.” Find an image on the web that matches. Still what you would expect. But this is the retrieval step, and note what is being retrieved: an image, not a document.
- “Determining… a first unit of text, wherein the first unit of text comprises at least a portion of textual content of a source document that includes the result image.” This clause is the entire article in legalese. The text for the answer comes from the page that contains the matching image. Elsewhere the filing describes this text being taken from around the image in the source page, the paragraphs and sentences sitting next to it.
- “Processing… the first unit of text and the prompt with a generative language model to generate a second unit of text.” The AI writes the answer from that text. The abstract adds that the output can be “at least some of the first unit of text, or text derived from the first unit of text” (a direct quote from your page, or a rewrite of it).
- “Providing… the second unit of text for display within an interface.” The answer is shown with the image as a visual citation, a thumbnail with attribution pointing at the source page.
Nowhere in that sequence is there a second competition where the best article about the topic gets a chance to supply the answer. The document that owns the matching image is the source.

Why this breaks the mental model
Previously, I had assumed a two-step model of visual search:
- The image identifies the subject. “That is a monstera with root rot.”
- The answer gets assembled the usual way. From whichever pages best answer “monstera root rot treatment”.
Image for identification, text competition for the answer.
Under that model your images do not affect whether you get cited, so nobody optimises them beyond alt text and file size. Under the patent’s model, the consequence is blunt:
- A mediocre article with the best-matching photo beats the definitive article with a stock photo.
- The best content on the topic can lose the citation to whoever owns the image that looks most like what the user snapped.
- Stock photography cannot win this game for you, because the same image matches to every site that licensed it. A photo only you have is a retrieval key only you own.
This is a genuinely different ranking competition with different winners, and as far as I can tell nobody is deliberately competing in it.
The scale is not niche either. By Google’s own numbers, Lens handles nearly 20 billion visual searches a month, and 20 percent of them are shopping-related (that is roughly 4 billion monthly searches where someone photographs a thing they might spend money on). Those are the highest-intent queries there are: the user is standing in front of the problem, holding the camera.

How a store turns this into money
The mechanism is interesting. The money is in what you do with it. Here is the worked example, then the pattern.
The replacement parts store. A man’s washing machine stops draining. He does not know what the broken part is called, so he cannot type a useful query. What he can do is point his phone at it: “what is this part and where do I get a new one?”
Right now, the store selling that part has the manufacturer’s catalogue photo on a white background. So do the forty other stores pulling from the same supplier feed. That image can never make one store the source; it matches everywhere at once.
So the owner re-shoots. For his 200 best sellers, he photographs each part the way a customer actually meets it:
- still installed in the machine, grubby, at a wonky angle
- the common failure states: the snapped impeller, the perished seal, the kinked hose
- the model badge and the part in the same frame
And directly under each photo, he writes the text the AI will lift: what the part is called, which models it fits, the symptom it causes when it fails, the price, the buy button.
Now the chain runs: customer snaps his broken pump, the store’s “pump still installed in a Bosch drum, slightly corroded” photo is the closest visual match on the web because nobody else photographed the real-world state, the answer is generated from the text sitting next to that photo, and the customer taps through to a page selling exactly the thing in his hand. That is not traffic. That is a customer arriving pre-qualified at the till.
The same play works anywhere the customer can see their problem but cannot name it:
- The plant shop: photos of the diseased leaf, the pest damage, the scorched tips, each next to the treatment product it sells.
- The plumber: photos of the weeping joint, the furred-up valve, the boiler error screen, next to “this is what it is, this is what it costs to fix, here is my number.”
- The skincare brand: photos of the actual skin condition its product addresses, next to the routine that addresses it.
- The car parts seller: the worn brake disc photographed on the hub, not on a white background.
The rule underneath all of these: everyone competes on text for typed queries, where the biggest sites usually win. Almost nobody competes on owning the photograph of the actual problem. Original photos of the messy, real-world version of what people point cameras at are cheap to produce, impossible for competitors to duplicate, and under this mechanism they are the retrieval key for the highest-intent query format in existence.
The catch worth understanding
Four honest caveats before you brief a photographer.
- This is a patent application, not a confirmed production system. Google files thousands of these and plenty never ship as described. What a filing tells you is how Google is thinking about a problem, and this one dates back to May 2023, which is plenty of time for the thinking to have made it into products.
- It describes one pathway, not the whole system. Nothing stops Google blending this image-first path with ordinary text retrieval, and in production it almost certainly does. The claim describes the mechanism, not its market share.
- It covers image-plus-question queries specifically. Typed AI Mode queries and standard ChatGPT prompts are a different pipeline. Everything I wrote about what actually gets cited by AI still applies there.
- Proximity is described loosely. The filing says the lifted text can come from around the image, but it does not commit to a distance. Treat “answer text next to the image” as sensible engineering, not a guarantee.
The test, which matters more than the patent, is whether real citations track image ownership. That is checkable: run image-plus-question queries, look at what gets cited, and see whether the cited pages contain a near-duplicate of the matched image or are simply the best text results. I am setting that experiment up now, and the results will be the follow-up piece whichever way they land.
So what do you actually do
In priority order, by budget.
- Gold: run a problem-photography programme. Photograph your products and, more importantly, the problems they solve, the way a customer with a phone camera meets them: in situ, imperfect, recognisable. Put the answering text (name, fitment, symptom, price, action) directly beside each image. Start with your 50 highest-margin pages, not your 50 highest-traffic ones; this mechanism serves buyers, not browsers. (Building content around the questions and problems buyers actually have is the job my agency does day in day out, so weigh my bias accordingly.)
- Middle: fix your best pages only. Pick the 20 pages that make you the most money and give each one a single original, in-situ photograph placed inside the section that answers the main question, not in a decorative banner at the top. Caption it with the answer.
- Cheap: audit what you already have. Reverse-search your key product images in Google Lens (upload your own image and see what it matches to). If your image matches forty other sites, it is stock or a supplier feed photo and it cannot win you a citation. Check you are not blocking Google-Extended or image crawlers in robots.txt, and check your images are not rendered in ways bots cannot fetch. Alt text and descriptive file names still help the system connect image and meaning.
And regardless of budget: stop paying for stock photography on pages you expect to earn citations. It is camouflage. It makes your page look like everyone else’s to the one system that is choosing sources by looking.
Where I land
The GEO industry has spent two years arguing about text while Google filed paperwork describing a retrieval path where text comes second. I would not bet the house on any single patent shipping as written, but I would also not keep optimising exactly one half of every page when the other half is 20 billion queries a month of uncontested ground.
The next piece will have the experiment data. If you find this sort of thing useful (patents and research nobody else is reading, translated into things you can actually do), the newsletter box is in the footer below, and the experiment results will land there first. And if you want your visibility strategy looked at before your competitors buy a camera, get in touch.