|
|
|
🚀 The AI Trend Oracle Report
The Number Was Stable. The Recommendation Wasn't. |
|
Last week I ran the test I had promised, published a clean result, and then spent five days being corrected on it in public by four different people, every one of them right.
The corrected version is more useful than the one I published. So that is where we start. |
|
 |
|
|
|
|
|
🔍 Featured Insight |
|
The label held. The answer underneath it moved.
The test. Three prompts, two engines, twenty runs each. 120 responses. One clean transactional question, one clean explanatory question, one deliberately ambiguous one. Method posted publicly before the first run, including the threshold and what I would conclude if the answer turned out to be boring.
I had argued in several threads that run-to-run instability in how an engine interprets a prompt might undermine the diagnostic tables a few of us were designing. The pre-registration said plainly: if the result comes back stable, my objection was wrong and I say so.
The result was a clean null.
All six cells came back 20 out of 20 on a single executed task. Every cell 83 to 100 percent on an exact binomial interval. One distinct task observed per cell across twenty runs.
The objection was wrong. I published that.
Then somebody asked a question I had not thought to ask.
A commenter suggested separating two things I had been treating as one thing: task-selection stability, and response-content stability. His point was that a cell can be 20 out of 20 on a label while what it actually says varies underneath, cited sources, named entities, and recommended brands.
So I went back through the raw responses with that second axis.
On the transactional prompt, ChatGPT executed "Recommendation" twenty times out of twenty. The tool it recommended did not hold.
Run 1: one product. Runs 2 through 15: a different product. Runs 16 through 20: back to the first.
Fourteen out of twenty for the leading brand. The exact interval on that is 45.7 to 88.1 percent, which under the threshold I had pre-registered is not a stability result at all. It is Inconclusive.
A perfectly stable label, sitting directly on top of the only number a brand in that category actually cares about, which was not stable.
On Perplexity the same prompt broke differently. The top pick held every run. But the alternatives named alongside it churned, products appearing in one run and absent from the next and the pricing cited for the same products differed run to run.
Two engines. Identical stable task label. Two different instabilities underneath. A report recording only the task would have called both of them clean.
Then it got worse, and this is the part worth your time
The block structure in that ChatGPT cell is not what randomness looks like. Fourteen consecutive, then five consecutive, is three runs in twenty draws.
And three of those responses opened with these phrases:
"I'd slightly revise my earlier answer" "rather than relying on my previous answer" "I checked the current landscape again"
Those runs were in a temporary chat with memory off. There is no earlier answer.
Something carried state across runs that should have been isolated. Until I know what, those twenty responses are not twenty independent observations, and my clean 20 out of 20 is partly an artifact of that.
Four corrections in one week, none of them mine
-
I reported task stability. What I had evidence for was task-label stability.
-
The runs inside a single sitting were not independent, and I had not checked that before publishing.
-
I compressed two of the cells into a summary when recording them instead of keeping the twenty individual responses, so that part of the record cannot answer questions that came up afterwards.
-
I proposed an analysis to separate the two possible explanations, and was shown it cannot work inside one continuous sitting, run order and clock time are the same variable.
Every one of those came from somebody in a comment thread who owed me nothing.
Here is what I want to pull out of it, because it generalises well past me:
Pre-registration protects the analysis you planned. It does nothing for the analysis the data turns out to demand.
You cannot know in advance which question the results will raise, and it is usually the second question that changes the answer. That is not an argument against pre-registering. It is an argument for keeping far more raw material than your planned analysis needs.
Which is exactly where I fell short, and I only found out because somebody asked. |
|
|
🎯 This Week's Strategic Lens |
|
The measurement is most reliable exactly where it matters least
This one is not mine. It came from somebody in the thread who read the results and said the thing I had missed, and it is the most commercially consequential sentence anybody handed me this month.
On the clean transactional prompt, both engines executed the same task, twenty out of twenty each. Perfect cross-engine agreement.
On the deliberately ambiguous prompt, they did not agree at all. ChatGPT read it as a request to explain a concept and did that twenty times out of twenty. Perplexity read it as a request to diagnose a situation and did that twenty times out of twenty. Both were perfectly consistent. Both were answering a different question.
Now the part that matters commercially.
Real buyers do not write measurement-clean prompts. They type what is in their head. Half a sentence, personal, vague about what they actually want. Buyer-voice questions are ambiguous close to by definition, that is what makes them buyer-voice rather than keyword.
So the prompts every visibility tool in this category is running are structurally the ones most likely to be read differently by different engines.
Which gives you this:
A clean prompt like "best tool for X" is the easiest case to measure and the least representative of how anyone actually asks.
An ambiguous prompt that looks like real purchase behaviour is the hard case to measure and the one that decides purchases.
Take five engines and ask them all: what should a mid-sized SaaS company use to improve AI visibility?
One reads that as recommend a platform. One reads it as explain a strategy. One reads it as diagnose what to fix first.
You cannot put those three outputs in a recommendation comparison table and call the differences visibility. They are not three answers to one question. They are three answers to three questions, and nothing in the output tells you that happened.
The fix is cheap, which is the good news. Before comparing two engines on a prompt, check whether they executed the same task on it. That is one run per engine. A precondition, not a study.
If they read the prompt the same way, compare the answers.
If they did not, the comparison is not noisy. It is empty, and the honest output is not comparable rather than a number.
I would rather a report told me it could not compare two things than handed me a figure that averaged an explanation against a diagnosis. |
|
|
🚨 Market Intelligence |
|
Measurement just became actionable, and that changes what measurement has to do
On August 25, Featured launched GEO Audit. It identifies which AI prompts mention your brand, which competitors appear instead, which publications the engines are actually citing, and it ranks those publications as PR targets. It will draft the pitches too.
Their launch data covers 22,881 AI citations across 405 audits run between June 2 and August 21. The figure that stopped me:
34.5 percent of citations came from sites with Domain Authority below 40.
That is a genuinely useful finding and it cuts against the received wisdom in this category. A third of citations are not coming from the large publications everyone is competing over. They are coming from smaller sites nobody is bidding on.
Pricing starts at $29 a month.
I want to be fair about this, because the reflex in my part of the industry is to treat any commercialization as a threat. Identifying which outlets influence AI answers and then pitching them is ordinary public relations done with better targeting than guesswork. "Focus on the five sources that matter instead of guessing" is honest advice and better advice than most of what is being sold.
But something does change once those five sources have names.
Until now, improving your position in AI answers has been diffuse. Publish more, get mentioned more, hope. When a tool can hand you a ranked list of the specific domains that move your specific score, the work stops being diffuse and starts being a target list. Marketers will pursue those domains. Of course they will. That is not corruption. That is marketing responding rationally to a distribution channel that has become measurable.
The consequence lands on the measurement, not on the marketers.
Once source influence is actionable, the measurement has to get better at explaining how a source came to exist because the population of sources is about to stop being a natural sample and start being an optimized one. |
|
|
📊 What's Actually Happening |
|
Provenance: the thing a scan structurally cannot see
Two things landed in the same few weeks and they turn out to be the same problem wearing different clothes.
The first arrived in my inbox on Wednesday.
A media researcher at a pay-on-results PR firm had read my LinkedIn article about two AI visibility tools disagreeing, and thought the argument could become a feature in a major business publication. No charge unless it runs. Would I like to talk.
The firm is real. It is publicly listed. It discloses its model plainly on its own website, which is more than a lot of this industry manages.
So I went and looked at where their own coverage lives.
And I could not tell, from the outside, which of it was editorial and which was placed.
That is the entire problem, stated in one line, and I am not going to pretend to a certainty I do not have. Some of it may well have been written by journalists who found the story. Some of it appears in sections that publications use for partner content. From where I was standing, with the same view an AI engine has, the two are not distinguishable.
There is a real spectrum here and it is worth naming properly, because collapsing it is exactly the counting error this newsletter has spent a month on. Earned. Editorially selected after a pitch. Contributed. Placed. Sponsored. Paid. Those are six different provenance states. A placed article can still go through genuine editorial review. A sponsored piece can be accurate and useful. None of them is automatically illegitimate.
But they are not the same evidence, and:
The scan sees the publication. It does not see how the coverage got there.
I have not yet seen a visibility report that records how a source came to exist. If one exists I would like to be shown it, in the same spirit as everything else in this newsletter and I would rather be corrected than keep repeating something that has stopped being true.
The second thing landed on August 2, and it is the same hinge.
Article 50 of the EU AI Act became applicable that day. Providers of generative systems must now mark synthetic outputs in a machine-readable format so they are detectable as artificially generated. Systems already on the market before that date have until December 2 to comply.
There are carve-outs, and one is worth reading twice. AI-generated text published to inform the public on matters of public interest does not require disclosure if it has been subject to human review and editorial control, with somebody identifiable holding editorial responsibility for it.
So in the disclosure regime, the line between labelled and unlabelled is editorial responsibility.
And in the provenance question, the line between earned and placed is also editorial responsibility.
Same hinge, two regimes, and neither one is visible to a system that counts domains. A machine-readable mark tells you a machine wrote it. It does not tell you who decided it should exist.
I did not reply to the PR email, and not out of any principle about public relations, which is a legitimate trade older than most of the things I sell. I passed because of what a placed article would be inside the specific systems I measure. Two weeks ago I published that four of my own sources resolved to two independent judgments. Adding a placement I had arranged, and then selling other people a tool for spotting exactly that, is not a position I could hold with a straight face. |
|
|
|
👨💻 Founder's Note |
|
The distortion in my own business model
Somebody in one of these threads made an observation about why nobody in this category says the honest thing out loud. His version: the honest version has no billing shape. A snapshot invoices once, a real instrument only works as a retainer, so the pressure runs toward describing one as the other.
He named two billing shapes and left out the third. Mine.
I am going to name it, because it would be strange to spend a month writing about undisclosed incentives and quietly leave my own out of it.
A subscription creates pressure for the number to visibly move.
Not once. Every month, in the renewal conversation, in the moment somebody looks at an invoice and asks what it bought. A one-off report is judged on whether it was useful. A subscription gets judged on whether something has happened since last time.
That pressure selects, quietly and without anybody deciding it, for a metric that is sensitive over one that is valid.
A flat month may be the correct result. It may be the best result and the position held, nothing was lost, no competitor displaced you. But "nothing changed" is much harder to make feel valuable than "your score improved six points," and a metric that wobbles a little looks more like a working product than one that sits still. Nobody has to be dishonest for that to shape a roadmap. It only has to go unexamined.
I do not have a clean answer. The partial one, and I think it is real, is to agree with a client what a good outcome looks like before the first scan runs. Set the prompt set and the threshold up front, so a flat month reads as the position held, which is what we agreed success looks like, rather than nothing happened, why am I paying. The identical data point reads as failure or success depending entirely on whether the expectation existed before the number did.
That helps and it does not solve it, because by month four a client who agreed that stability is success starts wanting movement again. We are all bad at valuing the absence of change. The best structural suggestion I have heard is to pre-register a review date alongside the threshold, month six, we reopen what we are measuring so the restlessness has somewhere scheduled to go instead of arriving as dissatisfaction.
And here is the part I did not see coming when I started writing this section.
Last week I spent 120 runs measuring whether the movement is even real. A meaningful share of what I found was an engine rereading the same prompt, or a session carrying state it should not have carried. Movement with no brand cause underneath it at all.
Which is uncomfortable, because movement is the thing my business model quietly rewards.
Three billing shapes, three ways to be wrong, and none of them fixed by wanting to be honest. Only by writing the incentive down somewhere somebody can hold you to it.
So it is written down.
— Dana Billingsley | Founder, Axis Suite |
|
|
🛠 THE STABILITY CHECK |
|
Fifteen minutes, on any AI visibility report including one of mine. Last month's audit asked whether your sources were independent. This one asks a different question: is your number stable, or does it just look stable?
1. Run one prompt five times and record the recommendation, not the appearance. Most people record whether the brand appeared. Record which brand was actually advanced as the answer. Those are different measurements and only one of them is what a buyer acts on.
2. Check whether the runs were independent. Read each response for phrases like as I mentioned, my earlier answer, revising what I said. If those appear in a fresh session with memory off, something carried state and your five runs are not five observations. This caught me and it will catch you.
3. Watch what changes underneath a stable answer. If the top recommendation holds across runs, look at what moved anyway — which competitors were named alongside it, what claims were made about pricing or features, which sources were cited. A stable headline over a churning body is a real finding and no dashboard reports it.
4. Before comparing two engines, check they answered the same question. Label what task each one performed. Explain, compare, shortlist, recommend, diagnose. If the labels differ, you do not have an engine comparison, you have two different questions and one table.
5. Ask what a flat result would look like in your reporting. If your tool has no way to express held position, nothing to act on, then it will find something to report every month whether or not anything happened. Which tells you something about the tool, and something about who is paying for it. |
|
|
☕AI Caffeinated Wisdom |
|
|
Your Weekly Dose of Caffeinated Wisdom
You have a friend who always answers the same way.
Ask them where to eat and you do not get a list, you do not get hedging, you do not get well, it depends. You get one restaurant, named with total confidence, with a reason attached.
Every time. Same shape, same certainty, same delivery.
After a while you stop thinking of them as opinionated and start thinking of them as reliable. That is a small and reasonable slide and nearly everybody makes it, because consistency of manner reads as consistency of judgment. In people, it usually is.
Then it comes up that four different friends asked them the same question in the same week and got four different restaurants. All named with total confidence. All with a reason attached.
Nothing they said was wrong. Every individual answer was fine. What was not true was the thing you had quietly concluded from the form of the answers, which was that a settled view sat underneath them.
The manner was stable. The answer was not. And the manner is the part you can see.
This is precisely what my own test did to me. The engine performed the same kind of task, twenty times out of twenty, with total consistency. Underneath that, the recommendation moved.
If I had recorded only the shape of the answer, I would have written down completely stable and believed it.
The question to ask of a confident, consistent source is not whether it answers the same way.
It is whether it answers the same thing.
Stay steady. ☕
|
|
|
🔔 CLOSING SIGNAL |
|
Three things happened close together and they are one story.
My own test showed that a number can hold perfectly still while the thing it is meant to be measuring moves underneath it.
A tool shipped that names the exact publications the engines cite for your category and ranks them as outreach targets.
And a firm offered to place an article for me in that kind of publication, in a way I would not have been able to distinguish from the outside afterwards.
Put them together and you get where I think this category actually is. We have built measurements that count appearances, in a market that is rapidly learning to optimize for those appearances, using instruments that cannot reliably tell a stable answer from a stable label.
None of that makes the measurement worthless. It means the number was never the product. What the number is standing on is the product.
Two questions I would put to any evidence panel, including mine:
How many of these would still be here if one of them went away?
And how did each of them come to exist?
Neither appears on a dashboard. Both are answerable in an afternoon. And the brands that come through the next model update intact will be the ones who could answer both without going to look first. |
|
|
|
|
🛠 Coming Next Week
First, this study is finished.
Not paused, not continuing quietly. Finished. I said before I started that I would publish what came back either way, including if it was boring, and I have, the null, the four corrections, and the two questions it raised that it cannot answer.
Those two stay open and I am going to leave them open rather than close them with numbers I do not trust. Whether the blocked runs were a session carrying state or something moving underneath in the index, I do not know. The data cannot separate the two, and the design that would is written down in a public thread for anyone who wants to run it: two batches the same day, one as a continuous sitting, one with a fresh isolated session per run. If the blocking shows up only in the continuous batch, the collection workflow was carrying state. If it shows up in both, something moved underneath.
I am not going to be the one who runs it. Saying that plainly is better than letting it hang around in a "coming soon" that quietly never arrives, which is how most of these end.
What the study was for, it did. It answered the question I asked, it turned out I had asked a slightly wrong question, and four people I have never met made it better in public. That is a good outcome for 120 queries.
Second, the thing this issue argues for, I now have to try to build.
The claim above is that no report I have seen records how a source came to exist. It is easy to write that in a newsletter. It is harder to put a field in a product.
I have a head start on it that I nearly talked myself out of. Axis Suite already ships a field that regularly comes back unknown: whether a cited source is owned by the brand or independent of it, alongside a record of how that call was made, domain match, name inference, or not determined at all. On real scans it lands on inferred or not determined more often than I would like. We ship it anyway, because unknown is a real state and concealing it would be worse than reporting it.
So the question is not whether a mostly-unknown field is sellable. I already know the answer to that one.
The question is whether provenance can be determined at all, and it is a harder problem than independence for a specific reason.
Independence is mostly knowable. You can look up who owns a domain.
Provenance mostly is not. Earned, editorially selected after a pitch, contributed, placed, sponsored, paid you cannot tell most of those from a URL, and for several of them nobody outside the publication and the client ever can.
So next week I will write up what a provenance field would actually look like, including where it breaks. My expectation going in is that it splits three ways: a small set of states that can be determined, a larger set that can only be flagged as likely, and a residue that is honestly unknowable from the outside. Whether what remains is useful enough to put in front of somebody is the part I do not know yet.
That is where I actually am, so that is what you will get.
Which is the whole method, restated one more time: say what you measured, say what you did not, and be specific about which is which. |
 |
|
|
|
|
🚀🤖✨📊🎨
The Axis Suite
AI Recommendation Intelligence. AI Narrative Defense. Agentic Visibility Infrastructure. |
|
"Axis Suite is the independent intelligence layer that explains what AI believes about your brand, why it believes it, and what decision that belief ultimately drives."
👉 Axis Suite | Read the Proof Center
📬 Thank you for being part of this.
Here's to staying visible, staying understood, and becoming the obvious choice.
Stay caffeinated, stay inspired. See you next week! ☕
From the trenches,
The Axis Suite Team 💪
|
|
|
 |
|
© 2026 TrendAxis, LLC™. All rights reserved.
|
|
|
|
|