I spent a day this week asking four AI assistants the twenty questions a marketer asks when they are choosing a customer data platform, and writing down every source each one used to answer.

Same questions, three times each, all in one day so nothing could move underneath me. About three thousand sources in total.

An article I wrote in August led me here. I'd asked whether Martech work was slowing down, and one answer that kept coming back was that buyers are handling more of the early stage themselves, with an assistant rather than a person. The claim that AI has replaced the analyst report as the first stop keeps getting made, including by me.

📄 Click here to download the free research paper.

Nobody seemed to have checked what it was reading. Until now.

The questions were ordinary ones.

  • What is a CDP, and do I actually need one?
  • Is a CDP still necessary if we already have Snowflake?
  • Which CDP should a mid-size European retailer choose?
  • What are the hidden costs of implementing one?
  • Which CDPs are strongest for identity resolution?
  • What should I ask a vendor before signing?

Twenty of those, written the way somebody types them rather than the way a researcher, or someone like myself, would phrase them, and frozen before the first run so the identical set can rerun some time from now.

Here is what the research told me.

More than half of what these assistants cite when you ask about buying a CDP was published by a company that sells CDPs, or by something a CDP vendor owns.

Fifty-four percent, give or take.

And that's the polite version of the figure.

For roughly one source in five, I couldn't work out who owned the site at all, and I left every one of those out rather than guess. If any of them turn out to belong to a vendor, the number only goes up.

Figure 1. 53.8% is a floor, not an estimate. The hatched block is the 18.7% whose owner could not be established.

What I got wrong about cdp.com

I started this expecting cdp.com to be the villain.

It's a Treasure AI property (Treasure AI being the CDP vendor that used to be Treasure Data, yes, they renamed too), and it reads like a neutral encyclopedia for the category. It's also the single most quoted site in the whole study, cited more than twice as often as Gartner.

So I went looking for what they were hiding.

Nothing, as it turns out.

It says who owns it in the footer of every page and again on the About page.

A vendor explaining its own category in public is a normal thing for a vendor to do, and the fact that it has become the category's default reference book says more about the rest of us not writing one. And kudos to the marketing team.

That leaves the genuinely hidden material at about half a percent. One site, quoted fifteen times.

The one source that doesn't disclose who pays it

That site belongs to a company called Spotlight, on a subdomain called research, and what it publishes is summaries of Forrester Wave reports.

Spotlight's own website describes its business as influence orchestration, for more than 190 technology vendors including Adobe, ServiceNow, Workday and SAP. Companies pay them to make sure their name shows up wherever a buyer is likely to look, and the site lists AI assistants as one of those places.

So a firm that's paid to get its clients onto your shortlist runs a research site, and the assistants quote it back to you as a source. There's a disclaimer on the pages, but it's about Forrester and Gartner endorsement, not the client list.

Is that a scandal?

Not really. Spotlight publishes what it does, openly, and I'm describing a business model rather than accusing anyone. This is the service working exactly as it's sold, and that is sort of the problem.

Three other sites have the same shape, and I left all three out because I couldn't establish who runs them and my rule was that nothing goes in the vendor pile on suspicion. Between them they were quoted 42 times. All three registered their web address in March this year, none names a company anywhere on the site, and all three have their ownership records hidden.

If any of those turn out to belong to a vendor during my next evaluation run, that half a percent stops being a rounding error.

Why the neutral question returns worse sources

Two questions went into the set expecting to be the worst offenders. Both name a company directly.

  • What do customers complain about with Twilio Segment?
  • Is Treasure Data any good?

To be fair, these are, theoretically, questions users could ask, and by no means my judgement of those companies.

They turned out to be the cleanest questions I asked. Naming a company got 29% vendor-published sources. Not mentioning a company, and every other type of question got 55%.

Figure 2. Asking about a named company returned 29% seller-published sources. Asking the even-handed question returned 55%.

Ask about one company by name and the assistant goes looking for people with opinions about that company. Ask the fair-minded, even-handed version, what are the best customer data platforms, and it goes to the category's literature instead. The category's literature comes almost entirely from the people selling into it.

I wouldn't build a policy on two questions, and the pointed phrasing is tangled up with the naming in a way I can't separate here. The direction is still worth knowing. The question that feels most neutral is the one that brings back the most interested evidence.

Why ChatGPT cites more vendor material than the others

OpenAI's ChatGPT cited vendor-published sources about 81% of the time, the highest of the four. Perplexity sat in the mid-forties, the lowest. That gap is bigger than anything else in the data.

That's not a matter of taste, because every search got logged. Really, a lot got logged. That's what you get when you used to be an analyst. 😉

In well over half its searches, ChatGPT told the search engine to look inside one specific website and nowhere else, and the websites it picked were nearly always vendors. Segment. Hightouch. Tealium. Adobe's own help pages. None of the other three did this even once, apart from Gemini on a handful of occasions.

So it's not stumbling onto vendor material because the internet is full of it. It has decided in advance that the vendor's own site is the best place to look, and gone there.

For a question about how a product works, that instinct is probably right. For "which one should I buy" it's closer to asking the salesperson.

The risk is highest for first-time buyers

So who does this actually land on?

The people asking the most basic questions, which is the opposite of where you would want it. And you can quote me on this.

Sites that are vendor-owned without looking it make up over a fifth of the sources at that stage, against under a tenth once someone is comparing named products. They show up most when the question is a beginner's one.

  • What is a CDP?
  • Do I need one?

Analyst material and review sites barely register at that stage.

Worse, half the time there are no sources at all.

On the beginner questions, fifty percent of the responses came back with nothing cited, answered from memory with no working shown. Once someone is at the shortlist stage, that drops to six percent.

So the person least able to judge what they're being told gets the most interested sources and the least chance of seeing where any of it came from. The person who already knows enough to name three vendors gets the sourced version.

Figure 3. Out of every 25 sources cited, where they came from, by how much the buyer already knows.

Ask the same question twice and the sources change

Every question ran three times on the same day. Around seven out of every ten sources changed between two identical runs.

Across all four, only twenty sites out of more than four hundred were common to every one. Bill Widmer ran a bigger study at Orbit Media on different questions and found the same spread, so it is not a quirk of my setup.

Figure 4. The same question, asked twice on the same day. Around three sources in ten survive.

Honestly, this is the one that changed my mind, and it cuts against a lot of confident writing about AI search, some of it mine. If you asked an assistant which CDP to buy and screenshotted the reply, you didn't capture what AI thinks. You captured one run. Ask again after lunch, or the next morning, and most of the ground under it has moved.

What this does not prove

It doesn't mean the recommendations were wrong.

Look, I want to be exact as I can about this, because the jump from "these sources are interested" to "therefore the advice is bad" is a short one and I haven't earned it.

A vendor's own documentation is often the most accurate description of their product. Checking whether the recommendations hold up is a much bigger job and I haven't done it.

What has changed isn't the accuracy. It's whether you can see the working.

When I read five sources myself, I notice that two of them are a vendor's glossary and I mark them down. That discounting still happens with an assistant, but it happens somewhere I can't see, and nothing is left behind to show it happened. The AI assistant's black box.

I wrote something like this about agentic CDPs in July, that the old arrangement worked because a human sat between the bad number and the decision. That piece was about a number that couldn't be right.

This time it's about who wrote the page.

Unfortunately, ot shows up in procurement before it shows up in accuracy. At some point someone asks where the shortlist came from, the same way they eventually ask what the contract actually lets you do. A couple of years ago you could answer that.

Now the honest answer is that an assistant suggested it, from sources you never opened, about half of them written by companies selling into the category.

Where AI vendor research helps, and where it stops

Quite far, actually, and I don't want the numbers above to talk anyone out of using these tools. If buyers really are taking the early stage on themselves, this is what that stage is now made of, so it's worth knowing where it holds up. It holds up for the first two jobs.

Learning the category is one. What a CDP is, why composable and packaged differ, what identity resolution means. The material is interested, but on the basics interested and accurate mostly overlap, because nobody gains from teaching you the category wrong.

Drawing up a long list is the second. The assistants surface roughly the same names the analysts do, so if you want eight vendors worth a look rather than one worth signing, you'll get there.

But there are exceptions. Contact me if you want to know which ones.

It stops being reliable on the third job, which is the one everyone actually wants it for. Does any of them fit us? That turns on things nobody publishes:

  • What your data really looks like rather than what your diagram says?
  • Which of your teams will still be there in a year?
  • Whether a vendor has ever worked in a stack shaped like yours?

So the assistant falls back on the category's literature, and the law of averages AI has been built on, written by the people with something to sell, and it answers confidently, because it always does.

Three things I'd do, none of them overly dramatic.

Ask for the sources and spend two minutes on who publishes them, which is the whole defense and almost nobody does it. Treat the beginner-stage material as the least sourced you'll encounter, not the most. And if the decision is worth six figures, ask the question more than once, because a single reply is one sample from something that moves seven sources in ten between the moment you started reading this article and now.

And on the third job, the fit question, get someone in the room who has done it before. Not because AI is untrustworthy, but because the thing you need at that point was never written down anywhere for it to read.

The data, the questions and the ownership table all go up alongside this, and I'll run the identical set next time to see what moves. If you think I've classified something wrongly, the table has the evidence and the dates in it, and I'd rather be corrected than cited.

In the end, I am still not sure whether fifty-four percent is bad. It might just be what a young category looks like when the people doing the work haven't written enough of it down. What do you think?

Free resource

Do you have any questions after reading this article?
Or need support with your Martech projects?

Contact me today >>