
Much of the current advice around generative engine optimization (GEO) is based on theory, isolated screenshots, or a single campaign. I wanted measurable results, so I ran two structured experiments back to back, tracked the results manually, and documented what worked, what failed, and what changed between the two tests.
The first experiment, for an existing brand we consulted on, cost thousands of dollars and ran over several months. The second was a 30-day cold-start experiment a SaaS link building agency with no measurable AI presence when the test began.
Each test tracked 15 commercial-intent keywords. The first covered four AI platforms, while the second covered six. I logged 775 citation events across both experiments, and one of my original conclusions didn’t survive the second test. Let’s dig into how I set these up, what I observed, and how you can apply it to your work.
Experiment 1: The consultant test
We tracked 15 commercial-intent keywords across four platforms: ChatGPT, Claude, Gemini, and Perplexity. Every query was run manually, with and without a VPN, to account for possible location-based differences.
Our strategy was to place listicles on sources that were commonly surfaced by LLMs for our keywords, as well as on the brand’s site. We supported these through PR, guest posts, and organic LinkedIn activity.
By the end, the brand appeared for roughly 10 to 12 of the 15 keywords. Peak keyword presence reached 37.01% on April 29. The overall trend was upward, with meaningful week-to-week variation. The citations by platform were as follows: ChatGPT 148, Claude 96, Gemini 87, and Perplexity 64.
The leading source for citations was a comprehensive listicle on Indeed SEO, with 190 mentions. MEXC and a GlobeNewswire release followed, but neither came close to the same volume. In all, listicles accounted for 72.4% of citations and PR 24.1%. Guest posts, the owned site, and LinkedIn split what was left.
We also learned four major lessons you can apply to your next project.
See where your brand appears in AI search, where competitors are winning, and what it takes to become the answer AI recommends.
Listicle placement, PR, and guest posts reinforce each other
In this experiment, listicles and PR appeared to work as one system rather than as separate channels. The listicles that earned citations were usually the same ones being amplified through PR. The listicle introduced the claim, the press release reinforced it, and guest posts referenced both. Each layer appeared to perform better when supported by the others.
We found securing a placement on a source the model already cites can compound visibility quickly. Here, Indeed SEO appeared around 10 times in ChatGPT’s answers before we contacted the publication.
After our placement went live, it remained our largest citation source for months. That placement generated more citations than all the others combined.
Guest posts appeared to extend the effect of the PR placements. Claude, in particular, cited guest posts that referenced features in Yahoo and Business Insider, even when it didn’t cite those publications directly.
Source decay is rapid and authority matters
Roughly half of the sources stopped being cited within 30 days. One MEXC placement fell from 29 mentions to 11 week over week, while a Triple Review listicle that had been cited consistently declined without any direct intervention. The result suggests that one-time publication is unlikely to sustain visibility on its own.
In our tests, source authority and relevance appeared to matter more than content quality alone. A practical hierarchy was:
- Government and education sources.
- News publications.
- Industry-relevant sites.
- General sites.
In Phase 1, industry placements supported by PR produced strong results. In Phase 2, better-written listicles on general sites produced little visibility. PR appeared capable of amplifying a strong placement, but not compensating for a weak source.
Depth appeared to matter more than frequency. Our first five listicles covered the topic only briefly, and most generated little visibility. The one comprehensive piece continued to earn citations.
The peer set appeared to influence visibility. One listicle placed us alongside Lily Ray and Aleyda Solis. After those names were removed, performance declined within days. This suggests that the models may evaluate the surrounding entities, not only the individual mention. A similar pattern appeared in PR, where being named alongside recognized experts outperformed a standalone feature on a stronger outlet.
SERP visibility still appeared to influence LLM visibility, especially in ChatGPT. When the model relied on web search to resolve a query, brands absent from the retrieved results were also absent from the answer.
Capitalization and query type matter
The capitalization of queries appeared to affect which sources were retrieved. Capitalized and lowercase versions returned different citations in three repeated tests, although this finding requires further validation.
Query type appeared to influence the source types selected by the models. Software and tool queries favored high-authority review sites, while service queries more often returned listicles. Matching the placement type to the query appeared to improve the likelihood of being cited.
Exact-match keywords and answer placement are important
Exact-match keyword targeting still appeared to matter. We ranked for “Best LLM SEO Consultant” but barely appeared for “Best AI SEO Consultant,” despite the similar intent. The first phrase had a dedicated listicle, while the second didn’t. In this test, broader semantic coverage didn’t bridge the gap, suggesting that high-value commercial keywords may require dedicated assets.
In our tests, pages performed better when the answer appeared within the first 100 words. A key takeaway block near the top of the page produced a larger improvement than any other on-page change we made.
Other content findings include:
- FAQ content performed better when it was visible by default rather than hidden behind expandable sections.
- Self-contained sections appeared to perform better. For example, “What to look for when hiring an LLM SEO expert” and “Where to hire one” worked better as separate sections than as a single combined section.
- Question-based headings also performed better in our tests. For example, “How is AI SEO different from traditional SEO?” outperformed “AI SEO vs. traditional SEO.”
- Freshness also appeared to matter. Recent data, current references, and visible publication dates were associated with stronger citation performance.
We learned a lot from this first test.
The original plan for round two was a list of things to test: Person schema, LinkedIn cadence, and a YouTube push. Instead, I got the chance to run the whole playbook from zero on a different brand in a different vertical, which is a far better test of whether any of the findings are generalizable.
Dig deeper: 3 GEO experiments you should try this year
Experiment 2: The cold start test
The test involved a SaaS link building agency with no measurable AI presence when the experiment began. During the baseline window from April 30 to May 29, the brand had no measurable presence on any tracked platform.
The experiment ran from May 30 to June 28, and covered six platforms, adding Google AI Mode and Grok to the list from the first experiment. We tracked 15 commercial-intent keywords that a SaaS buyer might use while evaluating an agency. Every platform and keyword was checked manually.
The experiment resulted in 298 appearances in 30 days from a standing start. The appearances by platform were: Gemini 104, Google AI Mode 95, Claude 59, ChatGPT 32, Grok 4, and Perplexity 4.
Gemini and Google AI Mode accounted for roughly two-thirds of the platform totals listed above. This differed sharply from experiment one, where ChatGPT led. The comparison should be treated cautiously because AI Mode wasn’t tracked in the first experiment, and the two niches weren’t directly comparable.
Business impact: During the 30-day window, 18.5% of new users arrived through referral traffic, and another 3.25% through GA4’s AI Assistant channel. Together, these channels accounted for just over one-fifth of all new users. One Perplexity referral led to a prospect who later became a paying customer, even though Perplexity was the lowest-volume platform in the experiment.
Here’s what this experiment taught me.
Focus on observed citations and earned placements
I learned you want to build the target list from observed citations, not from DR alone. Before beginning outreach, I ran all 15 keywords through every platform, logged the sources that appeared, and ranked them by citation frequency. That ranked list became the outreach list.
Five of eight targets appeared in the final citation mix. Indie Hackers increased from 44 mentions during prospecting to 146 after our placement went live, a 232% improvement. Bruce Jones SEO increased from 26 to 69, up 165%, while TechBullion rose from 15 to 37, a 147% jump.
The approach didn’t work in every case. RankTracker and HR.com showed fewer citations after placement than during prospecting, and two shortlisted sites hadn’t appeared at all by the end of the measurement window. The method appeared to improve the hit rate, but it didn’t guarantee citations.
Placement concentration was extreme. Three sources — Indie Hackers, Bruce Jones SEO, and our own listicle — accounted for 342 of 437 total source mentions, or approximately 78%. Seven other live placements shared the remaining mentions. In this experiment, a small number of sources drove most of the visibility.
Within this experiment, earned placements outperformed owned content by a wide margin. Of the same 437 source mentions, third-party listicles generated 85.8%, our self-published listicle generated 14.0%, and PR generated 0.2%.
Measure citations over time
Time to citation ranged from one to 18 days. Two placements were cited the day after publication, while others took 10 or 11 days. Four placements were live but hadn’t been cited by the end of the measurement period. The variation suggests that checking only once, one week after publication, isn’t a reliable measurement approach.
The slowest source to be cited was the one we controlled. Our own listicle took 18 days, longer than every third-party placement, including two that were picked up overnight.
Owned listicles are a foundation, not a growth lever
This is where I had to revise my original conclusion.
After experiment one, I recommended publishing listicles on an owned site because many top-ranking brands appeared to benefit from their own content. Experiment two suggested that an owned listicle is more of a foundation than a primary growth engine. It was the slowest source to be cited and contributed 14% of mentions, while earned placements generated most of the visibility.
Dig deeper: How to know if your GEO is working
Comparative content is effective
The owned listicle improved when it became less promotional and more comparative. On June 23, we updated it to include our leading competitors instead of presenting the brand alone. Mentions of that source rose from four to 49, a 12.25-fold increase in the final count. The daily visibility curve also increased during the same period, from 22 on June 23 to a peak of 95 on June 27.
This mirrored the peer-set effect observed in experiment one, but from the opposite direction. Removing recognized names was followed by a decline in round one, while adding recognized competitors was followed by a substantial increase in round two. The same pattern appeared across two brands and two verticals, making it one of the findings I’d prioritize for further testing.
Comparative coverage outperformed advocacy in this test. The 12.25-fold increase followed an update that made the page more useful as a category resource rather than as a page focused primarily on our own company. The domain, author, and keyword target remained the same. The main change was the scope of the content.
Citations don’t equal clicks
The most-cited and most-clicked sources weren’t the same.
Indie Hackers generated more citation volume than any other source, but its referral traffic remained flat. TechBullion produced fewer citations but increased sessions from one to 64. Claude.ai referral sessions tripled, while ChatGPT referral sessions increased by 166%.
These results suggest that citation volume alone isn’t sufficient for deciding which sources deserve further investment.
Intent varies by model
In this experiment, Claude concentrated more heavily on high-intent terms. Gemini led in total appearances, 104 to 59, but Claude led or tied for first on five of the 10 best-performing keywords.
Gemini’s volume was distributed more broadly, while Claude’s was more concentrated on commercial queries. For a business evaluating platform value, that distribution may matter more than the headline total.
Exact-phrase assets also performed well
Our two strongest keywords were “Best SaaS Link Building Agency in USA” and the same phrase with “2026” added, with 18 appearances each. This repeated the pattern from experiment one in a different niche and suggests that dedicated assets may still be necessary for high-value commercial keywords.
Track your visibility across AI search, uncover missed opportunities, and grow your presence where customers are asking questions.
What I’d tell someone starting today
These experiments provided me with a solid list of lessons that can help you prioritize your efforts.
- Run your target keywords through the relevant AI platforms before investing in placements. The sources already being cited should inform the outreach list, and that list may differ significantly from a traditional prospecting spreadsheet.
- Budget for maintenance, not only initial publication. In our first experiment, roughly half of the sources stopped being cited within 30 days.
- Pay attention to entity associations. Across both experiments, performance changed when recognized companies or experts were added to or removed from listicles.
- Don’t treat citation volume as the final business outcome. The most frequently cited sources weren’t always the strongest referral sources, and one low-volume Perplexity referral resulted in a paying customer.
Traditional SEO focuses largely on ranking pages. Both experiments indicated that LLM visibility depends more heavily on associations: which sources mention you, who appears alongside you, and how recently those relationships were reinforced. The main disagreement was the role of owned content, which the second experiment showed was less powerful than I initially believed.
LLM visibility continues to evolve, but it’s never too early to build on these experiments and see what your data tells you.
Dig deeper: GEO for people who have to hit revenue targets

