Today was another day of cleaning up after AI hallucinations. Recently, I have been researching GEO more deeply across Chinese platforms.

Beyond websites and AI search, I have been looking at Xiaohongshu, WeChat Official Accounts, and the possible relationships between content platforms, AI systems, recommendations, citations, and brand discovery.

At first, I thought the main questions would be: What kind of content is easier for AI to discover? Which platforms are more useful for brand visibility?

Can certain patterns be reused across different channels? But after spending more time on the problem, I realized the hardest question is much more basic:

What actually happened? Many Chinese platforms are effectively black boxes from the outside. I can see whether content gets exposure. I can see whether a brand is mentioned.

I can see whether an AI system eventually recommends something. I can sometimes observe that content from one source appears more often than content from another.

But explaining the mechanism behind the outcome is extremely difficult. I usually cannot tell whether a recommendation came from: The content itself.

Account authority. User behavior. Historical data. Internal platform labels. Search systems. External websites. Or some signal I cannot observe at all.

I can see the result. The middle of the process is mostly hidden. This feels very different from traditional SEO. SEO is not fully transparent either.

But at least there is a relatively mature observation framework: Crawling. Indexing. Keywords. Links. Rankings. Page structures. Site authority.

Not every relationship can be proven directly, but there are established ways to investigate them. GEO is different.

Especially when the research moves from public websites into content platforms, recommendation systems, and closed ecosystems. Sometimes I cannot even confirm what the system actually saw.

That makes the field extremely vulnerable to speculation. The dangerous part is that AI is very good at filling these information gaps with stories.

If I ask: “Why might this content have been recommended?” AI can quickly produce ten plausible explanations. Better semantic structure. Higher account authority.

Stronger engagement. Clearer entity relationships. Better crawlability. More consistent publishing frequency. Every explanation sounds reasonable.

But: Reasonable does not mean true. Without real data, official documentation, or repeatable experiments, these are hypotheses at best. AI often presents hypotheses with the tone of established rules.

That has become one of the most frustrating parts of the research. I am no longer only researching GEO. I also have to research whether the GEO conclusions generated by AI are hallucinations.

Sometimes I use AI to save time. It produces a complete-looking framework. Then I spend a large amount of time checking it. Eventually I discover:

Some parts have evidence. Some parts are industry assumptions. Some parts were simply inferred from context. The tool that was supposed to accelerate the research creates another layer of verification work.

This is forcing me to redefine how Toket GEO should work. Previously, I focused more on: How can I identify the factors that influence AI recommendations?

Now I think the first step should be: Define the evidence level. I am increasingly separating information into three layers. The first layer:

What did I observe? For example: A platform produced a specific result. An AI system cited a specific source. A brand was mentioned for a specific query.

These are observable facts. The second layer: What do I think caused it? Maybe content structure matters. Maybe platform authority matters. Maybe entity consistency matters.

These are useful experimental directions, but they are not conclusions. The third layer: What can I actually prove?

Only when a pattern survives repeated tests, controlled comparisons, public documentation, or stable reproduction should it become a stronger conclusion.

This distinction sounds simple. But it matters a lot in GEO. The field is extremely vulnerable to: Turning correlation into causation. Turning experience into platform rules.

Turning one experiment into a permanent law. Turning AI inference into research findings. All of these can make a report look professional. But when a client eventually asks:

“How do you know?” The weakness becomes obvious. That is why I am becoming less interested in universal “GEO playbooks.” Especially for Chinese platforms.

Many systems simply do not expose enough information for strong conclusions. If I do not know, I should say I do not know. If something cannot yet be proven, I should say that too.

That does not make the research useless. It means the research method needs to change. I increasingly want Toket GEO to become an evidence system.

Not a system that tells users: “This is how the platform algorithm works.” But one that says: This is what we observed. These findings are confirmed.

These are high-probability hypotheses. These areas still need validation. This is what the next experiment should change. That may be less attractive than publishing “10 GEO secrets.”

But it is closer to real research.

While researching Xiaohongshu, WeChat Official Accounts, and similar platforms, this difference becomes especially clear.

After content is published, the way a platform understands it, distributes it, and possibly influences external AI systems is rarely a simple linear path.

Sometimes I cannot even confirm whether an AI system mentioned a brand because it directly saw the original content, or because it learned about it through another website, repost, database, or existing knowledge.

If the source path itself is uncertain, saying: “We did X, therefore AI recommended your brand” is extremely risky. This is also changing how I think about GEO reports.

AI loves certainty. “The core cause is…” “The main factors include…” “The priority recommendation is…” These sentences sound professional. But in a black-box system, more accurate language may be:

“Current evidence suggests…” “We observed…” “This factor may be related, but causation is not yet confirmed…” “The next controlled test should validate…”

The language sounds less confident. But the report becomes more trustworthy. I used to worry that this kind of writing would sound less professional.

Now I think acknowledging uncertainty is part of professional work. The real problem is not uncertainty. The real problem is sounding certain when the evidence is weak.

After spending more time on GEO, my understanding of the field is changing. It may not be a problem where we quickly discover “the platform rules.”

It may be closer to continuous experimentation across multiple black-box systems. Create a hypothesis. Design questions. Collect samples. Observe changes.

Control variables. Test again. Slowly reduce uncertainty. This is much slower than I originally expected. It is also much more frustrating. But at least the process is real.

Even though these platforms have been giving me a headache lately, this may also be the most interesting part of Toket GEO. Not because I already know the answer.

But because many answers may not actually be known yet. If Toket eventually creates real value here, I do not want it to come from producing the most impressive GEO theory.

I want it to come from being able to say clearly: This is fact. This is hypothesis. This is still unknown. This is what we should test next.

AI can generate answers very easily. Black-box platforms can also create strong illusions of causality. Put the two together, and it becomes very easy to produce conclusions that sound correct but are not actually proven.

So I am adding a principle to Toket GEO that I increasingly believe in: Do not turn correlation into causation, and do not present speculation as platform rules.

That may make the research slower. But at least it avoids pretending that a black box has already been understood.

Estimate task cost in the AI Cost Analysis or refine prompts in the Prompt Optimizer.