Recently, I have been spending a lot of time researching GEO. At first, the question seemed straightforward:

How can a brand become more visible and more likely to be recommended by AI systems such as ChatGPT, DeepSeek, and Gemini? But the deeper I went, the more I realized that this question may come too early.

Before asking how to get recommended, there is a more fundamental question: Can AI actually understand your website? That is why I recently built a free Website AI Readiness Check in Toket.

It is not a tool that promises to get your website recommended by ChatGPT. I wanted to start one layer earlier: From a technical and information-structure perspective, is your website ready to be understood by machines?

⸻ Websites were traditionally built with search engines in mind Anyone who has worked with websites is probably familiar with the basics of SEO.

Does the page have a title? Is the description configured? Is robots.txt correct? Is there a sitemap? Is the canonical URL correct? Can search engines crawl the page?

These concepts have existed for years, and relatively mature practices have developed around them. But now there is another information entry point.

More users are directly asking AI: “What products exist in this category?” “Which companies solve this problem?” “Should I choose A or B?” “Is there a tool that can do this?”

Instead of opening websites one by one from the first page of search results, users may begin directly with an AI-generated answer. The question therefore changes slightly.

We used to ask: Can search engines find me? Now we also need to ask: Can AI understand me? ⸻ Being searchable does not mean being understandable

This has become increasingly obvious during my recent research. A website may look perfectly clear to a human. There is a logo. A brand statement.

Product pages. An About page. Contact information. After browsing for a few minutes, a person can probably understand what the website represents.

Machines may see something different. They may need to determine: Is this a company, an individual project, or a product? What is the relationship between the brand name and the product name?

Which domain is the official website? Which page should be treated as an authoritative source? Is a statement describing an actual product capability or simply marketing copy?

Are descriptions of the same brand consistent across different pages? Humans naturally fill in many of these gaps using context. Machines may not.

Especially when an AI system needs to decide across many sources: Who is who, what is true, and which source should be trusted? The information structure of the website starts to matter.

⸻ I did not want to begin with a magical GEO score Building a score is not particularly difficult. Scan a website. Collect dozens of signals.

Assign weights. Then display: “Your GEO Score is 72.” It looks like a complete product. But the important question is: What does 72 actually mean?

Does it prove that ChatGPT is more likely to recommend the website? No. Does it prove that Gemini is more likely to cite it? No.

If a number cannot ultimately be connected to explainable evidence, then no matter how polished the interface looks, the number may simply become another layer of packaging.

So I wanted to start with things that can be directly inspected. Can important pages be accessed? Does robots configuration create obvious barriers?

Are canonical signals consistent? Does a valid sitemap exist? Is structured data available to machines? Are relationships between Organization, Person, SoftwareApplication, and other entities clearly expressed?

Are the brand, product, and organization described consistently? Are there obvious technical problems that make machine understanding harder?

These may sound like technical questions. But they all point to something much simpler:

If a machine that knows nothing about you enters the website for the first time, can it quickly understand who you are, what you do, and which information should be trusted?

⸻ What do robots.txt, canonical, and structured data actually mean? If you are not a web developer, these terms can sound unnecessarily technical.

They are easier to understand in plain language. robots.txt roughly tells automated crawlers: “You can access these areas, but not those.” A sitemap is closer to a map of the website:

“These are the important pages you should know about.” A canonical URL tells systems: “If you discover several similar URLs, this is the version I want you to treat as the primary one.”

And structured data is closer to a machine-readable description of what the website represents. A normal page might say: “Toket is an independent AI product lab.”

Structured data can make the relationships more explicit: This is an Organization. This is its name. This is its official website. This is the Founder.

This is a SoftwareApplication. These are the relationships between those entities. Most users will never see this information. Machines can.

The goal is not to trick AI into ranking a website. The goal is simpler: Reduce how much the machine has to guess. ⸻ What about llms.txt? This is one of the files frequently discussed in GEO right now.

At a high level, it is an attempt to provide LLMs with a clearer description and navigation layer for a website. But I am increasingly uncomfortable saying:

“Adding llms.txt improves AI rankings.” I do not think there is enough evidence to make such a direct causal claim. It can be a low-cost addition to a machine-readable website structure.

It can help organize important content and pages more clearly. But it is not a magic file. That reflects one of the principles I have become increasingly strict about while researching GEO:

Do not turn “this may help” into “this definitely works.” ⸻ A free website check cannot tell you why ChatGPT does not recommend you This boundary matters.

Imagine that a website passes every technical check. That still does not mean: ChatGPT will recommend it. DeepSeek will cite it. Gemini will consider it the best product in its category.

AI-generated answers may depend on many factors beyond the website itself. Existing model knowledge. Search results. Third-party content. Brand history.

External references. Community discussions. Content quality. Platform-specific retrieval systems. Even the same question may produce different answers at different times.

So Website AI Readiness Check is designed to answer: Does your website have obvious foundational barriers? It is not: An AI recommendation probability predictor.

Those are two different questions. ⸻ Why I still think a free readiness check is useful GEO discussions tend to move quickly toward more complicated questions.

How should content be written? How do we increase brand mentions? Should we publish on Xiaohongshu? What about WeChat Official Accounts? How do we get cited by AI?

How do we increase recommendation probability? But if a brand’s own website does not clearly communicate its basic identity, everything that follows becomes harder.

The homepage may use one brand name. Structured data may use another. The About page may describe the company differently again. Canonical URLs may be wrong.

English pages may contain Chinese metadata. The relationship between the product and the organization may be unclear. Important pages may not even be properly exposed to machines.

These problems are not particularly exciting. They do not sound as attractive as “ChatGPT ranking strategies.” But they have one important property:

They can be inspected. ⸻ Researching GEO in China has made me care much more about what can actually be proven

Recently, I have also been studying Xiaohongshu, WeChat Official Accounts, and other parts of the Chinese content ecosystem. The biggest impression so far is simple:

There are a lot of black boxes. Why was something recommended? Why was it not recommended? Where did an AI system learn about a brand? Did a particular piece of content actually influence the final answer?

Often, we can observe the outcome. But we cannot observe the complete causal chain. AI makes this even more difficult because it is extremely good at generating explanations that sound reasonable.

But reasonable does not mean true. That is why I increasingly prefer to begin with higher-confidence questions. Can the website be crawled? Is the page structure correct?

Is the brand entity clear? Is public information consistent? At least these questions can leave evidence behind. ⸻ My understanding of GEO is changing

Initially, I thought of GEO as: How do we optimize a brand so AI recommends it? Now I increasingly think about it as: How do we systematically reduce uncertainty when AI tries to understand a brand?

That includes technical infrastructure. Content. Brand identity. Third-party signals. Communities. External references. And many black-box systems that we still cannot fully explain.

So perhaps the first step should not be “optimization.” The first step may simply be: Understand your current state. ⸻ That is why I built the Website AI Readiness Check in Toket.

It will not tell you: “Fix these three things and ChatGPT will recommend you next week.” I do not want it to become that kind of product. It is closer to a basic website health check for the AI era.

First, determine: Can AI systems access the site? Does the website provide machines with sufficiently clear information? Are there obvious ambiguities between the brand and its products?

Which technical issues should be addressed first? Once that foundation is in place, we can move to the next question: How does AI actually understand, describe, and recommend your brand?

That question is much more complicated. And much more of a black box. The more I work on GEO, the more I believe that what the field needs right now may not be more “secrets.”

It may need more evidence that can actually be verified.

Estimate task cost in the AI Cost Analysis or refine prompts in the Prompt Optimizer.