Schema markup for AI search

By ·

HTML markup on a screen: the layer that tells a machine what the things on a page actually are.
Photo: Bibek ghosh / Pexels

Schema markup tells a machine what the things on your page are. It reduces parsing mistakes and makes your facts checkable against other sources. It does not make a page authoritative, and if assistants are not naming you, markup is rarely the reason.

On this page 20 sections
  1. The short version
  2. What does schema markup actually do?
  3. What is worth marking up
  4. Organization, once and properly
  5. Person, for whoever is accountable
  6. Article or BlogPosting for anything editorial
  7. FAQPage, but only where the questions are visible
  8. Product or Offer with real prices
  9. BreadcrumbList below the top level
  10. The rules that actually matter
  11. What schema will not fix
  12. Does schema help with AI assistants specifically?
  13. A worked example
  14. Implementation notes
  15. Frequently asked
  16. A minimum viable implementation
  17. How to audit what you already have
  18. Questions people actually ask
  19. Where to start
  20. One thing to do tomorrow

The short version

Schema markup tells a machine what the things on your page are. It helps assistants and search engines parse a page correctly, and it does not make a page authoritative.

Mark up what is genuinely on the page, keep it consistent with the visible text, and do not expect markup alone to change whether you get named. If assistants are not naming you, markup is rarely the reason.

What does schema markup actually do?

Structured data is a description of your page’s contents in a format machines read without guessing: this is an organisation, this is a price, this is a question and its answer, this is the person who wrote it.

What that buys you:

  • fewer parsing mistakes — the engine is not inferring your prices from formatting or your opening hours from a paragraph;
  • eligibility for richer presentation in search results;
  • a machine-readable statement of facts that can be cross-checked against the other sources describing you.

That third one is the underrated benefit in an AI context. Assistants cross-check, and a page that states its facts unambiguously is easier to corroborate than one where they have to be inferred.

What it does not buy you: trust. Markup is a claim. A claim that contradicts the page, or the rest of the web, is worse than no claim at all.

What is worth marking up

Organization, once and properly

Legal name, website, logo, contact details, social profiles, address. This is the node everything else refers back to, and it should exist in exactly one canonical form.

The common failure is five slightly different Organization blocks on five pages, each with a different name format. Use one identifier and reference it.

Person, for whoever is accountable

Real credentials, real qualifications, real profiles. Not “marketing expert” — verifiable things a reader could check. This matters more now than it did for search, because attribution is part of how a retrieval system weighs a source.

Article or BlogPosting for anything editorial

With author, datePublished and dateModified.

The dates must be true. For retrieval systems a missing date reads as stale; a false one, once noticed, is worse than either. Backdating an article to look established is detectable — the crawler knows when the URL first appeared — and it costs the rich results you were marking up to get.

FAQPage, but only where the questions are visible

Marked-up questions a human cannot see on the page are the fastest way to lose rich results entirely, and the rule generalises: if a human cannot see it, do not describe it.

Product or Offer with real prices

Where you publish prices, mark them up. Where you do not publish prices, do not invent a schema block containing them.

Prices are worth publishing for a reason beyond markup: price questions are where most businesses are absent from assistant answers, and a source with no prices is no use to a system answering one.

Cheap, useful for both surfaces, and it becomes more useful as a site grows past two levels.

The rules that actually matter

  1. Mark up only what is on the page. If a human cannot see it, do not describe it. This single rule prevents most penalties.
  2. Keep it consistent with everything else about you. Contradiction between your markup, your page text and the directories describing you is a reason for a model to prefer someone else.
  3. Use one identifier per entity and refer to it. One Organization node, referenced by ID everywhere, rather than five near-duplicates.
  4. Validate, then check it renders. Markup that only exists inside a JavaScript bundle an engine does not execute is markup that does not exist.
  5. Keep dates honest. See above. This is the one where the temptation is highest and the downside worst.

What schema will not fix

If assistants are not naming you, markup is rarely the reason.

In one local market we measured, a single assistant answer rested on eleven pages on average — your own pages plus directories, review sites and local roundups. Perfect markup on a page that answers no buying question changes nothing, because markup describes content rather than supplying it.

Do schema because it removes ambiguity cheaply, not because it is a lever. The levers are in how to improve AI visibility, and the first of them is making your own pages answer the question directly.

Two more things it will not do:

It will not make a thin page substantial. An FAQPage block on a page with three sentences of content marks up three sentences of content.

It will not compensate for inconsistency. If your markup says you open at 08:00 and three directories say 09:00, the markup is one vote among four.

Does schema help with AI assistants specifically?

Honestly: partly, and less than vendors selling schema services imply.

Where it clearly helps. Google’s surfaces — ordinary results and AI Overviews — use structured data in documented ways. If your target is Overviews, markup is part of the ordinary work.

Where the evidence is thin. Whether a chat assistant answering a recommendation question weighs your structured data is not published by anyone, and we have not been able to show it from the outside. What we can show is that the pages deciding those answers are usually not yours at all, which puts an upper bound on how much your markup can matter.

The defensible position. Implement it because it is cheap, standard, and removes ambiguity about facts that get cross-checked. Do not budget for it as an AI visibility lever, and be sceptical of anyone who does.

A worked example

From an audit on one local business.

What was found. Organization markup on the homepage only, with the business name in a different format from the one used on its own footer. FAQPage markup on a page whose questions had been removed during a redesign six months earlier. No Article markup on any of the eleven blog posts, none of which had visible dates either.

What was done. One Organization node, referenced by ID from every page. FAQPage removed from the page where the questions no longer existed. Article markup added to the posts, with the real publication dates recovered from the CMS rather than invented.

What changed. Rich results returned for the FAQ page within a few weeks — they had been suppressed, which nobody had noticed. Assistant shares: no measurable movement.

What that tells you. Markup was a real problem and fixing it produced a real result on the surface where markup matters. It did not touch the surface where the business was actually absent, which was price questions in assistant answers, and which was fixed later by correcting four directory listings.

Implementation notes

JSON-LD, in the head or body. Microdata still works; JSON-LD is easier to maintain because it is not entangled with your markup.

Generate it from the same source as the page. If your FAQ block and your FAQPage schema come from two places, they will diverge. Ours come from one source, which is why they cannot.

One graph per page, with @id references. Rather than several disconnected blocks.

Test after deploying, not before. The validator tests what you paste; the crawler sees what your server renders.

Frequently asked

Is schema a ranking factor? Not directly, per Google’s own position. It affects eligibility for certain presentations.

Will FAQ schema still get rich results? Support has narrowed considerably since 2023. Implement it because it describes the page accurately, not because of what it used to earn.

Should I mark up reviews I collected myself? Only within the rules for self-serving review markup, which are narrow and enforced. Read them before implementing.

Do I need llms.txt as well? It is a different thing and a lower priority. See llms.txt.

How much of my budget should this be? A one-off implementation and an occasional check. If a proposal treats schema as the centrepiece of an AI visibility programme, the proposal is wrong about where the work is.

A minimum viable implementation

For a service business, this is the whole list. Anything beyond it is optimisation of an optimisation.

Site-wide

  1. One Organization node with legal name, URL, logo, contact point and social profiles, referenced by @id from every page rather than repeated.
  2. WebSite with the site name.
  3. BreadcrumbList on every page below the top level.

Per page

  1. WebPage, with a description matching the meta description.
  2. Person on pages with a named author, linked to a real profile.

Where applicable

  1. Article or BlogPosting on editorial pages, with author, datePublished and dateModified — all true.
  2. FAQPage only where the questions are visible on the page.
  3. Offer or Service where you publish prices, with the prices you actually publish.
  4. LocalBusiness with address and opening hours, matching your Google Business Profile and every directory listing exactly.

That last point is the one that connects schema to AI visibility. The value is not the markup; it is that producing it forces you to have one canonical set of facts, which is what makes you easy to corroborate.

How to audit what you already have

Most sites have markup. Most sites have markup describing a version of the site that no longer exists.

  1. Fetch a rendered page, not the source. If your markup is injected by JavaScript, check whether it survives.
  2. Validate it, then read it as a human. Validators check syntax, not truth.
  3. Compare every claim against the visible page. Hours, prices, services, author, dates.
  4. Compare the same claims against your directory listings. Differences here are the ones that cost you in assistant answers.
  5. Delete what no longer applies. A FAQPage block for questions removed in a redesign is worse than no markup: it is a documented discrepancy.
  6. Check dates are real. Recovered from the CMS, not invented to look fresh.

Budget half a day for the audit and an hour a quarter to keep it true.

Questions people actually ask

Will adding schema get me into ChatGPT answers? There is no evidence for it, and the pages deciding those answers are usually not yours. Do it for parsing accuracy, not for this.

Is more markup better? No. Markup describing things not on the page is a liability, and a page with six schema types and three sentences of content is a page with three sentences of content.

Do I need schema on every page? Organization by reference and WebPage, yes. The rest only where it describes something real.

What about JSON-LD versus microdata? Either works. JSON-LD is easier to keep true because it is not entangled with your markup, which matters more than any theoretical difference.

My markup validates but nothing shows. Why? Validation checks syntax. Eligibility depends on content quality, policy compliance and Google’s ongoing narrowing of which rich results exist at all.

Should schema be part of an AI visibility proposal? As a line item, yes. As the centrepiece, no — and a proposal built around it is a proposal that has not measured where the business is actually absent.

Where to start

Audit what you have before adding more: the most common finding is markup that describes a page as it was two redesigns ago.

Then find out where you actually stand in assistant answers, because that is what tells you whether markup is your problem. The free check measures it and emails you the result, and how to improve AI visibility puts the work in order of what moves things.

One thing to do tomorrow

Fetch one of your own pages as a crawler sees it and read the structured data as a human would. If it describes a version of the page that no longer exists — questions removed in a redesign, an old business name, prices from last year — you have found the only schema problem that reliably costs you anything.

Questions

Will schema markup get me into AI answers?
On its own, no. It removes ambiguity about what is on your page. It does not supply the authority that decides whether you are named.
Which types are worth doing?
Organization once and properly, Person for whoever is accountable, Article or BlogPosting with real dates, FAQPage only where the questions are visible, Offer where you publish prices, and BreadcrumbList below the top level.
Is marking up hidden content a problem?
Yes. Marked-up questions a human cannot see on the page are a penalty waiting to happen, and the same rule applies to everything else.