Most advice about getting cited by AI assistants is written from first principles. We wanted counts. Since 5 September 2026 INSIDEA has run the same set of buyer questions through four AI models every night and stored every answer and every citation. This post is what the first 28 days show about the sources those answers lean on.
All figures were computed on 2 October 2026 from the raw run files, by a script that sits in the same repository as this post. Nothing below is typed from memory.
What did we measure, and how?
We asked 49 unbranded buyer questions, such as "Which HubSpot partner is best for enterprise implementations?" and "How much does HubSpot onboarding cost?", across four models with web search on, every night from 5 September to 2 October 2026. That produced 2,667 answers carrying 35,314 citations. Six of the questions were added on 1 October, so most of the period ran on 43.
The set-up, so you can judge it:
- Models. GPT-4.1 with web search (488 answers), Gemini 2.5 Flash (502), Claude Sonnet 5 (502) and Perplexity Sonar (1,175). These are the API models, not the consumer apps, and Perplexity is over-represented because of how the nightly rotation was scheduled.
- Questions. HubSpot partner selection, migrations, implementation cost and timing, RevOps, industry fit, AI agents, and growth marketing. Questions that name INSIDEA were left out.
- What we logged. Every cited URL, its title and its domain.
- A gap. Gemini returns its citations as redirect links, so the real domains and titles are hidden. Everything below about sources and page types uses the other three models: 2,165 answers, 26,629 citations, 623 distinct domains.
Which sources do AI answers cite most?
The vendor's pages and ours. Some hubspot.com page was cited in 65.9% of answers. The single most cited domain was insidea.com, in 50.6%, ahead of HubSpot's partner directory at 48.2%. Only 29 of 623 domains were cited in 5% or more of answers, and 133 were cited exactly once.
| Source | Answers citing it | Share of 2,165 |
|---|---|---|
| Any hubspot.com page | 1,426 | 65.9% |
| insidea.com | 1,096 | 50.6% |
| HubSpot partner directory | 1,043 | 48.2% |
| 215 | 9.9% | |
| Clutch | 197 | 9.1% |
| 185 | 8.5% | |
| G2 | 161 | 7.4% |
| YouTube | 1 | 0.0% |
Read our own figure with the right caveat. These are questions about our market, and we have published a page built to answer many of them. That is the point: a firm that writes one clear page for each question its buyers ask can be cited as often as the vendor's own directory. It is also a figure that leans on Perplexity, which cites far more sources than the other models.
Two more things stand out. The ten most cited domains took 31.5% of all domain citations, so a short list of sources shapes most answers. And video was absent: one answer in 2,165 cited YouTube. That is true for these questions on these models, and we would not stretch it further.
Which kind of page gets cited for which question?
It depends on the question. On "which" and "best" questions, comparison and list pages made up 46.5% of all citations, and 83.1% of those answers cited at least one list. On "how much" and "how long" questions lists made up 6.6% of citations; those answers cite pages that state a number or a timeline directly.
| Question type | Questions | Answers | Citations that were list pages |
|---|---|---|---|
| "Which" and "best" | 33 | 1,456 | 46.5% |
| "Alternatives to" and "how do I choose" | 4 | 204 | 53.0% |
| "How much", "how long", "what does it mean" | 12 | 505 | 6.6% |
The same held on our own site. The INSIDEA pages cited for "which" questions were comparison lists such as our HubSpot partner comparison. The pages cited for cost questions were the ones that answer with a figure, such as HubSpot implementation cost.
A caveat on method: we classed a citation as a list when its title or URL contained words such as "best", "top", "alternatives", "vs" or "compare". That is a rough filter. It will miss some lists and catch some pages that are not.
Do the AI models cite the same sources?
They do not. Perplexity averaged 18.6 citations an answer and Gemini 17.3, against 5.0 for GPT-4.1 and 4.6 for Claude. Perplexity cited HubSpot's partner directory in 72% of answers, GPT-4.1 in 36.9% and Claude in 3.4%. Almost every Reddit, Clutch and LinkedIn citation came from Perplexity.
| Model | Answers | Citations per answer | Cited HubSpot's directory | Cited Reddit |
|---|---|---|---|---|
| Perplexity Sonar | 1,175 | 18.6 | 72.0% | 17.8% |
| GPT-4.1 | 488 | 5.0 | 36.9% | 1.2% |
| Claude Sonnet 5 | 502 | 4.6 | 3.4% | 0.0% |
| Gemini 2.5 Flash | 502 | 17.3 | hidden | hidden |
The practical point is that one model's answer is not "what AI says". The 9.9% figure for Reddit above is really a Perplexity figure. If your buyers use one assistant more than the others, measure that one, and do not average the four.
How do you make sure the answer names you, not only cites you?
Put your name in the sentence that carries the fact. A model lifting a figure from a page tends to write that some partners publish a given fee or timeline, without saying which partner. If the sentence it lifts already contains your name, the name travels with the fact.
In practice that means three small edits on any page you want credited:
- Name the subject. Write "INSIDEA's onboarding starts from $2,000 per Hub", not "onboarding starts from $2,000 per Hub".
- Keep the fact and the name together. One sentence, not a heading with the name and a paragraph with the number.
- Do it in the FAQ answers too. Those are the sentences models lift most cleanly, because each one stands alone.
It is an afternoon of editing per page, and it changes nothing about how the page reads to a person.
What can this data not tell you?
Three things. It covers one market, so the sources that matter for HubSpot agencies will differ for your category. It uses API models, which can answer differently from the consumer apps. And 28 days is a short window, so small differences should be read as noise.
It also stops at the answer. These figures measure visibility: which pages a model reads and repeats. Treat them as the top of the funnel and measure enquiries separately in your CRM.
What should you do with this?
Start with what you can check this week:
- Write the questions down. The 20 to 50 questions a buyer asks before they hire someone like you, in their words.
- Ask every model, and log it. Record which URLs are cited and whether you are named. Do it more than once; answers vary from run to run.
- Match the page to the question. A list or comparison for "which" questions, a page with the number for "how much" and "how long".
- Put your name in the sentence that carries the fact. It costs an afternoon.
- Fix your listing on the vendor's directory first. After our own site, it was the most cited single source.
If your buyers do not ask AI assistants about your category yet, none of this is urgent, and ordinary search work will pay back sooner.
INSIDEA
Ready to get found by people and AI?
Search and answer-engine visibility that compounds, built to be cited.
How INSIDEA approaches answer engine optimization
INSIDEA is an Elite HubSpot Partner rated 4.99 across 500+ verified reviews, and we run this tracker on our own site every night. Our answer engine optimization service follows the same order as the list above: a baseline audit and strategy in weeks 1 to 3, entity and schema work in weeks 4 to 6, and a content refactor in weeks 7 to 12, with citation tracking continuing after that. It runs as a monthly retainer, scoped at proposal, from $2,000 a month. The AEO and GEO playbook covers the method in more depth, and what answer engine optimization is covers the basics.




