What makes an AI comparison fair and reproducible.
A comparison is only worth reading if someone else, running the same two sites on the same day, would reach the same result. Here is the standard AIOStory holds itself to.
Most claims about which business "wins in AI" fall apart the moment you ask a second question: would anyone else get the same answer? Run the check again tomorrow, from a different machine, and watch the winner change. When a comparison moves depending on who ran it, it was never a measurement. It was a story with a chart attached. Comparison is the concept AIOStory owns, so the burden is on us to say plainly what separates a comparison you can stand behind from one you cannot.
The same inputs must produce the same output
Determinism is the first requirement, and it is not a nicety. If two sites can trade places between two runs with nothing about either site having changed, the comparison is measuring noise. AIOStory scores each site on the AIOTruth evaluation engine, the same scorer that powers the single-site check at AIOInsights, and that engine takes no LLM into the scoring and no randomness into the result. Same page, same score, word for word. That is a constraint we accept even when it is inconvenient, because the alternative is a number that flatters whoever happens to be looking.
The practical test is simple. A fair comparison should be something you could hand to a skeptic with the two URLs and the date, and they could reproduce it themselves. If the only way to get our result is to trust us, we have not measured anything. We have asserted it.
A comparison must read what a machine reads
The second requirement is that the evidence has to be the evidence a machine actually sees. AIOStory reads public, on-page signals: the homepage-level content, structured data, robots.txt, sitemap.xml, and llms.txt, fetched the way an AI crawler would fetch them. It does not read a company's analytics, its private CRM, or anything a visitor could not also read. That boundary matters because the thing being compared is legibility to a machine, and a machine cannot see what is not on the page.
It also means being honest about what we are not doing. AIOStory does not query ChatGPT, Claude, Gemini, Perplexity, or AI Overviews and then present the answer as theirs. A deterministic scan of public signals is a proxy for how readable a business is, and it is a good one, but it is not the same act as asking a model live. Confusing the two would be the easiest way to make a comparison feel more authoritative than it is, which is exactly why we refuse to do it.
The method cannot bend to pick a winner
The third requirement is the one that separates measurement from marketing. The rubric is fixed before the two sites are chosen, not after. Six pillars, the same six every time: clarity, identity, substance, reach, reputation, and locality. No pillar is weighted up because it happens to favour one side. There is no owned-domain floor, no flattering exception for a site we like, and no adjustment once we can see who is ahead. If you can change the method to change the outcome, the outcome was decided by the person holding the method, not by the sites.
This is why AIOStory does not attack anyone. A comparison that exists to make a chosen competitor look bad has already abandoned the fixed rubric, because its result was the input. The point of a real comparison is that neither side is protected and neither is targeted. A site that reads poorly reads poorly, including ours.
Coverage has to be stated, not implied
The fourth requirement is disclosure of scope. A brand with forty locations is not fully described by four URLs, and pretending otherwise is its own kind of dishonesty. So every AIOStory comparison discloses how much it sampled: N of M, stated on the result, so the reader knows whether they are seeing the whole brand or a slice of it. A conclusion drawn from a sample is only as strong as the sample is representative, and hiding the sample size is how a thin read gets dressed up as a full audit.
The same discipline applies when a site cannot be read at all. A location behind a bot challenge, or one that simply does not respond, is reported as unreadable, not scored as if everything were fine and not silently dropped so the average looks better. An unreadable location is a finding, often the most important one, because when AI cannot read a location it recommends one it can.
The honest objection: an on-page scan is not the model's live answer
Here is the strongest case against everything above. A reproducible on-page score is not what a customer actually hears when they ask an assistant for a recommendation. Models pull from far more than one homepage: third-party coverage, reviews scattered across the web, training data none of us can inspect. So is a deterministic page scan even measuring the right thing?
It is measuring the part we can measure honestly, and that is the trade worth making. The alternative, sampling live model answers, is seductive and unreproducible: the same prompt returns different text on different days, across accounts, across regions, with no way for a reader to verify what we saw. That is precisely the kind of number this whole note argues against. On-page legibility is not the entire story an AI can tell, and we say so directly. But it is the layer a business controls, the layer that changes when the page changes, and the layer a second person can check. A smaller claim that holds is worth more than a larger one that evaporates when someone tries to reproduce it.
A comparison you cannot reproduce is an opinion with a chart on it.
None of this makes a comparison the end of the work. It makes it a reliable starting point. When a fair, reproducible comparison shows that AI reads a competitor more clearly across a market, that gap is real, and closing it is strategy and implementation, which is where Digilu takes over. The comparison earns its authority by being checkable. What a business does with the result is the part that changes the outcome.
Further reading
- The full AIOStory method, pillar by pillar, with the sampling rule and the no-fabrication guarantees aiostory.com/methodology
- AIOTruth, the deterministic evaluation engine each site is scored on before AIOStory compares them aiotruth.com
- AIOInsights, the front door where a single site is evaluated on the same engine aioinsights.com
These links are the ecosystem's own documentation, offered as navigation and not as independent proof. AIOStory owns comparison; AIOInsights is the front door; every qualified next step is Digilu.