llms.txt, schema.org, sameAs: the technical signals AIs actually read

· 6 min read

AI engines don't read your site like a human does: they look for technical signals that make your content understandable, trustworthy and citable. The four main ones are schema.org structured data (which describes who you are and what you do in a format the machine understands), the llms.txt file (a page that guides AI bots toward your important content), entities and the sameAs field (which link your name to reference sources like Google, Wikidata or your directories), and citable content (direct, sourced answers that are easy to extract). None of these signals alone guarantees a citation, but their absence often makes you invisible. A GEO agency like ReplySeal puts these elements in place in a coordinated way, starting by checking that AI bots are actually allowed to read your site.

Why an AI doesn't read your site like a human

When a customer types a question into ChatGPT, Perplexity or Google's AI search, the machine doesn't go through your site page by page the way a visitor would. It relies on what it learned during its training and, for engines that retrieve information in real time like Perplexity, on a quick reading of pages deemed reliable.

The problem: an AI needs clear reference points to understand who you are, what you offer and whether it can cite you without risk of getting it wrong. A visually beautiful site can remain unreadable to it if the information isn't structured and explicit.

That's where technical signals come in. They aren't hidden tricks, but standardized ways of telling the machine: here is my name, my profession, my city, my answers to common questions, and here are the sources that confirm my existence. The four signals that follow are the most useful for a professional practice or a business in France.

Schema.org structured data: describing your business for the machine

Structured data, also called schema.org, is a standard vocabulary that Google and the engines understand. Concretely, it's a small block of code invisible to your visitors, added to your pages, that labels each piece of information: this is a business name, this is an address, this is an opening time, this is a customer review, this is a frequently asked question.

For a practice or a business, the most useful types describe the local organization (LocalBusiness), the specific service provided, and sometimes questions and answers. These tags feed Google's Knowledge Graph, the knowledge base that then powers AI Overviews and Gemini.

An honest nuance: the evidence is mixed on the direct effect of schema on citation by ChatGPT or Perplexity. Some 2026 analyses suggest these engines don't always read the markup as structured data. The benefit therefore comes mainly from the clear content that the schema accompanies, and from consistency on Google's side. Practical verdict: schema is put in place because it can't hurt and helps on Google's side, but it isn't enough on its own.

The llms.txt file: a guide at the entrance to your site

The llms.txt file is a recent signal, designed specifically for language models. It's a simple text page placed at the root of your site that plays the role of a table of contents for AI bots: it indicates which content is important, where to find it, and how to describe you in a few lines.

The idea is close to the robots.txt file that SEO specialists know, but geared toward AIs. Rather than letting a model guess what matters on your site, you present it with a clear, prioritized version of your business, your services and your reference pages.

You have to stay clear-eyed: llms.txt is an emerging convention, not yet a standard universally respected by all engines. Putting it in place costs little and presents no risk. It's a reasonable bet, provided you don't oversell it as a miracle solution. It's part of a set, alongside schema and a point that's often forgotten: allowing AI bots to read your site.

Entities and sameAs: linking your name to reliable sources

An AI trusts an entity more when it finds it consistently in several places. In plain terms, if your practice appears under the same name, the same address and the same phone number on your site, your Google listing, your profession's directories and possibly Wikidata, the machine considers that you are a real and well-identified entity.

The sameAs field, built into schema.org, serves exactly this purpose: it links your page to your other official presences, such as your Google Business listing, your professional profiles or your Wikidata entry. It's a way of telling the machine: these different profiles all refer to the same entity.

Two concrete levers stand out for the French market: NAP consistency (identical Name, Address, Phone everywhere) and presence in recognized sector directories, for example avocat.fr for lawyers. These have become hygiene factors: their absence penalizes you, their presence makes you identifiable.

Citable content: giving the AI an answer it can reuse

The most decisive signal remains the content itself. An AI cites what it can easily extract: a direct answer to a specific question, clearly worded, ideally backed by a verifiable fact or source.

Concretely, a page that starts by answering the question asked, before elaborating, has a better chance of being reused than a text that circles around the subject. That's the answer-first principle: the answer first, the context afterward. Frequently asked questions handled in a question-and-answer format work well, because they match the way people query AIs.

A reference point from research on the subject, notably the Princeton work on generative engine optimization: adding verifiable citations and statistics to content can significantly increase its probability of being included in an AI answer. Conversely, vague content, without sources and without a clear answer, gives the machine little to grab onto.

What a GEO agency actually puts in place

Put together, these signals don't require understanding everything technically: they require being laid down in the right order and maintained. That's the role of an AI visibility agency like ReplySeal, in done-for-you mode, without you having to touch any code.

The sequence is logical. First, check that AI bots (GPTBot, ClaudeBot, PerplexityBot, Google-Extended) aren't blocked by the robots.txt file, because a block makes everything else useless. Next, put in place schema.org structured data and the sameAs field pointing to your official profiles. Then publish a clear llms.txt file. Finally, produce citable content, in answer-first format, with answers to your customers' real questions.

  • Technical audit: robots.txt, NAP consistency, directory presence
  • Schema.org markup and sameAs links to your reference sources
  • Publishing an llms.txt file and citable content
  • Monthly monitoring of your presence in AI answers

None of these elements guarantees a citation. But their absence is, itself, a guarantee of invisibility.

Frequently asked questions

Is the llms.txt file really read by AIs?

It's a recent convention, designed to guide language models toward your important content. Not all engines respect it systematically yet. Putting it in place costs little and presents no risk, so it's a reasonable bet, provided you see it as a complement and not as a solution that's sufficient on its own.

Is schema.org enough to be cited by ChatGPT or Perplexity?

No. The evidence is mixed: some 2026 analyses indicate that these engines don't always read the markup as structured data. Schema mainly helps on Google's side, via the Knowledge Graph that feeds AI Overviews. The real factor remains the clear, citable content that this markup accompanies.

What does the sameAs field actually do?

It links your page to your other official presences (Google Business listing, professional directories, Wikidata) to signal to the machine that it's indeed the same entity. Combined with strict consistency of your name, address and phone everywhere, it helps engines identify you reliably.

Do you need technical skills to put all this in place?

No, if you go through an agency in done-for-you mode. ReplySeal installs and maintains these signals for you, without you having to touch your site's code. The work mainly consists of laying down the elements in the right order and keeping them up to date.

What matters most among these signals?

Citable content, meaning direct, sourced answers to real questions. Next comes the consistency of your entity (NAP, sameAs, directories) that makes you identifiable. Schema and llms.txt are useful supports, but secondary compared to clear content and a well-linked identity.

Do these signals guarantee appearing in AI answers?

No, and you should be wary of any promise of a guaranteed result. These signals increase your chances of being understood and cited, but no engine commits to a citation. What is certain is that a site without these signals, or whose AI bots are blocked, is almost always invisible.

Want to know which signals are missing from your site today and whether AIs can even read you? Launch the free ReplySeal audit: we check your robots.txt, your schema, your directory presence and your visibility in AI answers, then tell you what to fix first.

Try ReplySeal
llms.txt & schema: the signals AIs read · ReplySeal