Skip to content

Measurement, opened up

The number is useful only when you can challenge it.

Aelo keeps the prompt, every returned answer, failures, model, region, sample count and source evidence close enough to inspect.

measurement.v2aelo-brand-scorer.v2brand 2026-08-29.1

Counting rules

Mechanical where possible. Labelled where judgment remains.

The score cannot be edited by the sentiment analyser, a missing provider, or a URL copied from answer prose.

01

Mention rate

Deterministic

Successful answers containing an exact brand or known alias count as mentions. Visibility is mentions divided by successful samples. Failed samples do not enter the denominator.

02

Ranked position

Deterministic

When an answer contains a numbered, bulleted or headed list, Aelo records the first list item containing the brand. Position is averaged only where a valid position exists.

03

Source provenance

Deterministic

A structured provider URL is a provider citation. A URL found only in generated prose is a mentioned link. Invalid, local and private-network URLs cannot become citation evidence.

04

Sentiment

Labelled analysis

An analyser may classify the surrounding context. Its schema-checked output cannot change the deterministic mention count. When unavailable, Aelo labels the fallback and records zero analyser confidence.

A real run, including the mistakes

One question pattern. 100 calls. The winner changed 45% of the time.

This independent ChatGPT and Gemini experiment is why Aelo treats one answer as a receipt, not a visibility score.

01 · Question

What are the best [category] brands in India?

The same question pattern was repeated 10 times per category on each engine.

02 · Findings

45% top-answer volatility

Across 100 calls, the #1 recommendation changed on 45% of repeated checks.

03 · Gap

Two data-quality failures

Blue Tokai and Blue Tokai Coffee Roasters split one entity. An initial Gemini run also exhausted its response budget and produced unreliable output.

04 · Action

Fix the measurement before interpreting it

The invalid Gemini run was discarded, the response budget was corrected, brand aliases were normalized, and the identical questions were run again.

05 · Movement

A corrected 45% result—no invented lift

The published figure comes from the corrected dataset. The invalid run is not shown as a baseline, and no brand-improvement claim is made without a matched post-action measurement.

Case boundary: this experiment measured recommendation volatility, not market share or the causal effect of an SEO change. Aelo applies the same rule in product: compare only matched prompts, engines, models, regions, modes and scoring versions.

Confidence

A range, not fake precision.

Aelo uses a 95% Wilson interval around the observed mention rate. Fewer samples create a wider interval, so the interface should sound less certain.

Illustrative interval / 12 samples

Observed 58% · range 31–81%

31 lower58 observed81 upper

Fewer than 8 successful samples stays low confidence. Medium needs at least 8 and an interval no wider than 50 points. High needs at least 20 and an interval no wider than 30 points.

Before versus after

Aelo refuses the easy comparison.

Both periods need compatible conditions and enough successful samples. Even then, an observed change is not proof that one action caused it.

01Buyer prompt
02AI engine
03Provider model
04Region
05Measurement mode
06Search mode
07Scorer + contract
08Sentiment analyser

If a condition fails, the verdict is inconclusive. When all conditions match, improvement or regression still requires non-overlapping 95% intervals.

Test the method

Open the answer. Check what Aelo counted.

Run one real answer