What AI Actually Sees

Four organizations. One AI visibility study. An agency, a nonprofit, an art gallery, and a SaaS startup.

Generative Engine Optimization  ·  Answer Engine Optimization  ·  AI Search Strategy
emblitstein.com  ·  Client identities, brand names, and URLs anonymized. Bot traffic data from real monitored logs.

Client A

Regional marketing agency
Southeast US

Client B

Regional nonprofit
Southeast US

Client C

Contemporary art gallery
Southeast US

Client D

B2B SaaS startup
West Coast US

About the Pilot Testing Program

Earlier in 2026 as part of an introductory pilot program for my consulting business, I tracked how four organizations showed up in AI answers. The four organizations do not intersect other than three of them existing in the same region. They work in different fields. They serve different people. They produce different content.

All four had the same problems. Citations came from one or two topics only. Other topics got zero citations. Perplexity bots were blocked for most every client. Claude's live assistant was partly blocked for all four. Almost all citations pointed to just one or two pages.

AI visibility challenges are multi-facted. It is about your site's structure and content, but also about the popularity of your industry. Crawler access and content format decide whether AI cites you and having more content does not automatically fix the problem.

36,469
Total queries monitored
Across all 4 clients, 87-day window
1,471
Total citations earned
Combined across all platforms
2 of 5
Avg. verticals producing citations
3 zero-citation verticals per client
23
Total Perplexity citations
Client D only. A, B, C: zero
ClientTypeRegion QueriesCitationsMention rate
ARegional marketing agencySoutheast US12,7644563.57%
BRegional nonprofitSoutheast US8,8473624.09%
CContemporary art gallerySoutheast US5,218901.72%
DB2B SaaS startupWest Coast US9,6405635.84%
Combined 36,4691,4714.03%

Where are citations actually coming from?

ChatGPT had 74.1% of all 1,471 citations. That is not because ChatGPT is the best platform. It is because ChatGPT was the only platform all four clients could be read by. Perplexity was blocked. Claude's live assistant was partly blocked. Three of four clients had nearly all citations go to one platform.

Platform Client AClient B Client CClient D TotalShare
ChatGPT339271823981,09074.1%
Gemini9874714232121.8%
Claude191710372.5%
Perplexity00023231.6%
Total 456362 90563 1,471100%

Claude's 2.5% share is a result of access failures regardless of model. Gemini has 21.8% of citations.Client D was the only client to get any Perplexity citations off the bat, the other three companies had Perplexity blocked entirely as a result of normal bot filtration parameters on their servers.

Why the Perplexity window matters and what we can learn from it

AI crawlers do not all work the same way. Training crawlers index your content for future model updates while live assistant fetchers read your content in real-time when a user asks a question.. People are starting to understand that difference. What gets less attention is that these crawlers also return on very different schedules. That matters a lot when access is either blocked, broken, or limited to protect server load.

CrawlerTypeRe-index cycleError recovery time
ChatGPT-User / Claude-UserLive assistant fetcherReal-time, triggered by user queryImmediate once resolved
GPTBot / ClaudeBotTraining crawler7–21 days1–2 cycles (~14–42 days)
PerplexityBotLive + training hybrid45–90 days (observed from logs)1 to 3 cycles, up to 270 days
GooglebotTraining / index3–30 days (authority-dependent)1–2 cycles

Client A's logs showed this clearly. PerplexityBot got access for a short window in late February 2026. It read a few pages before it was blocked again on a schedule. From March 1 through the end of the study, every Perplexity visit went only to robots.txt. Each visit got a 403 error and stopped. The crawler came back every 3 to 5 days. Every time, it found the same block.

Fixing a Perplexity block today will not produce citations tomorrow. The re-index cycle appears to be somewhere between 45 to 90 days.

How AI agents see your site and what you should let them read

Most sites add block rules when a crawler causes a problem and may not be as nuanced when block as to which bots to allow from a souce and which to bar. Protecting yourself from massive bot traffic is standard, but AI creates a new problem of how to open access, when, and to which bots. Client A's bot logs showed 36 different AI crawlers during the study. Several large crawler categories were blocked across the whole site the entire time.

CrawlerTypeVisitsSuccess rateStatus
ChatGPT-UserLive assistant fetcher4,26098.2%Healthy
OAI-SearchBotAI indexing crawler6,17582.7%Healthy
DuckAssistBotAI assistant fetcher13680.9%Healthy
ApplebotAI indexing + assistant2,07358.9%Partial
ClaudeBotTraining crawler17,42335.4%Partial (training only)
Claude-UserLive assistant fetcher4942.9%Partially blocked
DeepSeekAI assistant1267.1%Mostly blocked
PerplexityBotLive + training hybrid1016.9%93% blocked, robots.txt only
GPTBotTraining crawler (OpenAI)2,2290.0%Fully blocked
meta-externalagentTraining crawler (Meta)30,9950.0%Fully blocked
AmazonbotTraining + indexing24,1900.0%Fully blocked

Source: Client A bot traffic logs, monitoring period 2025–2026. Success rate = HTTP 200 responses only.

ChatGPT-User is the live assistant. When a real user asks a question, it reads your site. It got in 98.2% of the time. GPTBot is OpenAI's training crawler. It builds the model's base knowledge of the world. It made 2,229 visits. It was blocked every time. The site was open to live queries but invisible to training data. These are two different bots with two different rules on the same server.

Crawler access framework

CategoryExamplesRecommendation
Live assistant fetchers ChatGPT-User, Claude-User, PerplexityBot, DuckAssistBot, DeepSeek, Gemini-Fetch Always allow. These generate real-time citations. Permit them in robots.txt before any general block rule.
Training crawlers GPTBot, ClaudeBot, meta-externalagent, Amazonbot, TikTokSpider Allow selectively. Blocking removes you from future training data. It does not affect live citations directly.
AI indexing crawlers OAI-SearchBot, bingbot, Googlebot, Applebot, YouBot Allow. These feed the retrieval indexes AI platforms draw from at answer time.
Unknown / catch-all User-agent: * and unrecognized agents Review before blocking. A blanket disallow catches live fetchers too.
What to keep away from AI: Keep pricing negotiation pages private. Remove old press releases with outdated claims. AI reads every page it can reach.

How much traffic is coming from AI, and how fast is that changing?

Traffic source Start of period End of period Change Notes
Google organic (avg.)61.4%57.8%−3.6ppShare declining; absolute volume stable
Direct / dark traffic18.2%17.9%−0.3ppLargely stable
Social / paid / email14.1%13.6%−0.5ppFlat to slightly declining
AI bot traffic (all crawlers)3.1%6.8%+3.7ppMore than doubled across the period
AI-referred user traffic0.4%1.1%+0.7ppFastest-growing referral source; Client D led at 2.3%

Percentage point changes reflect share of total traffic, not absolute volume. Estimates across all four clients.

Your AI referral traffic is an industry signal, not just an individual site success signal

Most brands see low AI referral numbers and think the site is the problem. It is not always the site. Your AI referral traffic shows how ready your whole industry is for AI citation.

Low numbers mean the whole category lacks AI-ready content. AI has nothing good to cite in your space. As more brands build structured content, citation rates rise for all of them.

The 50% threshold

As of March 2026, 48% of all Google queries trigger AI Overviews (BrightEdge, Digital Applied). Above that line, most discovery queries end with an AI summary. Users do not click through. Brands that fix access and structure now will get cited when that line is crossed.

Client A: Regional Marketing Agency

Southeast US  ·  Moderate regional recognition  ·  3.57% overall mention rate

The agency showed up in 456 of 12,764 queries. 86.6% of citations came from two verticals. The rate fell from 5.2% to 3.2% over the study because the query set grew to include verticals where the client had no content.

VerticalQueriesCitationsRate
Outdoor / Active Lifestyle~1,20020143.9%
Creative Agency Services~1,5001948.8%
AEC~3,20000.0%
Nonprofit~2,80000.0%
Rebranding~2,89100.0%

Every query that produced a citation had a location term in it. A content audit tested 43 topic pieces, and twenty-two scored 0 for AI visibility. Only one piece scored 4, a year-in-review article that named the agency directly.

Client B: Regional Nonprofit

Southeast US  ·  Moderate regional recognition  ·  4.09% overall mention rate

The nonprofit showed up in 362 of 8,847 queries. 80.1% of citations came from two topics. Fundraising and Advocacy got zero citations across 7,579 queries. Fundraising content is written for donors. It tells stories. AI needs facts it can extract. A mission statement does not produce a citation.

VerticalQueriesCitationsRate
Environmental Education / Coastal Conservation~1,10018047.2%
Volunteer & Community Engagement~1,16811012.4%
Major Gifts & Fundraising~2,10000.0%
Corporate Partnership / CSR~2,80000.0%
Advocacy & Policy~2,67900.0%

Client C: Contemporary Art Gallery

Southeast US  ·  Local / emerging recognition  ·  1.72% overall mention rate

The gallery showed up in 90 of 5,218 queries. That is the lowest rate of all four clients. 87 of those 90 citations came from named entity queries. The gallery's name or a specific artist's name had to be in the query. 97% of all citations pointed to the homepage alone.

Geography had to be very specific. Broad regional terms did not work. A city name or neighborhood name was needed. Smaller brands need tighter geographic anchors than well-known brands.

VerticalQueriesCitationsRate
Gallery Discovery / Exhibition Listings~8205814.1%
Artist & Work Research (named artists in inventory)~650297.1%
Art Investment & Secondary Market Guidance~1,10020.2%
Art Fair & Event Coverage~1,20010.1%
Collector Services & Acquisition Consulting~1,44800.0%

Client D: B2B SaaS Startup

West Coast US  ·  Early-stage / growing recognition  ·  5.84% overall mention rate

The SaaS client had the highest mention rate. Product pages have structure AI can read. Feature lists and pricing tables are easy to extract. Competitor comparison queries still got zero citations. AI cited the client's competitors instead. The client had no content that compared itself to anyone.

VerticalQueriesCitationsRate
Core product category / primary use case~1,80030433.8%
SMB / team-scale use cases~1,50015310.2%
Integration ecosystem & API use cases~2,100743.5%
Enterprise / procurement queries~2,200321.5%
Competitor comparison queries~2,04000.0%

Client D was the only client to get any Perplexity citations. It earned 23. Every one included a third-party source. Perplexity requires external proof before citing a brand. ChatGPT will cite your own pages directly. Perplexity will not.

VerticalPerplexity citationsAvg. positionPrimary source cited
Core product category147th–9thThird-party review platforms
SMB / team use cases69th–11thIntegration partner docs and blog coverage
Integration ecosystem311th+API documentation pages
Enterprise / procurement0No third-party procurement coverage indexed
Competitor comparison0Competitors cited instead

Takeaways

Content structure impacted mention rate better than brand size or industry. Client D's product pages not only have built-in structure but the structure is a universal standard across SaaS products. Client C's content is mostly storytelling with little data. That is why it scored lowest. Citation rates fell consistantly for all companies throughout the course of the 90 day test as more verticals were revealed through fanout queries.

PatternClient AClient BClient CClient D
Overall mention rate3.57%4.09%1.72%5.84%
Rate drift (start → end)5.2% → 3.2%6.3% → 3.7%2.8% → 1.4%7.1% → 5.2%
Citation-producing verticals2 of 52 of 52 of 54 of 5
Zero-citation verticals3 of 53 of 53 of 51 of 5
Perplexity access93% blocked94% blocked95% blocked66% blocked / 23 citations
Claude-User effective rate42.9%~39%~42%~61%
ChatGPT-User success rate98.2%97.2%96.3%99.1%
GPTBot (training crawler)100% blocked~blocked~blocked~blocked
Citation page concentration~95% on 2 pages~92% on 2 pages~97% on homepage~71% on 3 pages
Geographic requirement100% regional100% regional100% neighborhood-level60% geo / 40% category
Most commercially significant zeroAEC + RebrandingFundraising + AdvocacyArt investment + Collector servicesCompetitor comparison

Content gap framework

Zero-citation verticals do not all have the same problem or the same solution.

Gap typeSignalFix
Missing Competitors may get cited on queries where you never appear. If AI already has a source for the topic then recency and region area a way to overtake competitors in citations. Write new content for these topics. Put the direct answer first. Add a FAQ block. Include location terms.
Weak You rank in search but AI does not cite you. The answer may be buried, i.e. there is no direct answer in the first 150 words. Rewrite the intro as a direct answer. Add numbered steps. Change every H2 into a question.
Outdated AI cites a newer version of the same topic from a competitor. Your version is more than 18 months old. Add a last-verified date. Refresh every stat with its source and year.

Content structure that earns citations particularly for physical or SaaS products

ElementWhy it earns citationsWhat to avoid
Direct answer within first 150 wordsAI takes the first clean answer it finds. Everything after that is secondary."In today's digital landscape..." and any opener that delays the answer.
Question-phrased H2 headingsMatches the question the user asked. "How does X work?" works. "Background" does not.Vague section titles, noun phrases, internal jargon.
Single-sentence declarative claimsLong sentences are hard for AI to pull from. Short factual statements get cited more.Hedged language: "many studies suggest that, in most cases..."
Inline source attribution"Source: Gartner, 2025 (n=1,200)" raises AI credibility scoring.Footnotes, "research shows", and undated stats.

Human vs. bot query divergence

This was probably the most interesting observation during the course of testing. For the project 3 researchers were hired and they used Scribe to record search conversations and interactions with AI. They were instructed to test in their regular preferred browser, with ingognito, and using a VPN at various locations. A human using a local browser got drastically different results than automated tools like Profound, or even custom listening tools using API calls via AppScript. How you phrase a query matters but where you are matters more. Bot traffic that runs through these AI tools is also very different than human user traffic. Because AI is highly subjected to individual user environments and geo-locations, data gathered via bot queries was different. Automated tools give you a starting point, and in a future of agentic search that data is important, but user data should be gathered separately for an accurate picture.

QueryTypeResultKey observation
"best marketing agencies for environmental nonprofits" Evaluator, no geo National / generic agency list No location term returns national category leaders. Regional fit does not matter to AI without it.
"best marketing agencies for environmental nonprofits in the Carolinas" Evaluator, geo-anchored Localized agency list The location term changed the entire result set. Client A was absent even though it operates in that market.
"nonprofit digital marketing strategy" Naive, no geo General strategy content No agency citations at all. Broad queries produce educational content, not vendor lists.
"What content helps conservation nonprofits in the Carolinas appear in AI answers?" Long-tail, geo-anchored Specific local keyword recommendations AI returned very specific local terms like estuary names and local species. Those are the terms your content needs to contain.

Source: Human-researcher query session, Client A engagement. Queries run from regional browser, not neutral datacenter IPs.

What location terms do: A location term in a query tells AI to filter toward local content. Content that names the region gets cited. Content that assumes local reputation does not.

Adding a location term changed the entire result set. National names dropped out. Regional providers appeared. Client A was not on that list. Its content does not name its region clearly or often enough.

When asked what helps a regional nonprofit get cited, AI named very specific local terms. Named waterways. Local species. Event names. Your content needs those exact terms at that level of detail.

Agentic commerce

AI agents are already completing purchases for users. The standard marketing funnel has multiple steps which an agent collapses into a single prompt or task.

In the 2025 holiday season, AI drove 20% of all retail sales, $262 billion (Salesforce Commerce Cloud, US Chamber of Commerce, February 2026). That is double what AI drove in 2024. Gartner projects 33% of all web content will be built for AI search by 2026.

In B2B, agents are starting to buy from other agents. Google's Universal Commerce Protocol (UCP) is a system where brands pay to be in AI agents' shortlists. Etsy, Wayfair, Shopify, Target, and Walmart were early adopters.

This means a few things, particularly when we look at the citation and ranking differences between bot and human-user queriesa across regions. If agentic search and agentic purchasing overtake the market, then monitoring AI bot citaiton levvels becomes more important. However there are other oncoming regional avenues as well where the human-user data is more valuable, (car navigation ads, etc.)

FormatHow it worksWhat brands must do
Google UCP Brands pay to appear in agents' consideration sets when AI completes purchases on users' behalf. Match UCP feed standards. Keep product and service data structured and current.
B2B agent-to-agent procurement Buyer-side AI agents evaluate vendors and start purchases with no human needed for routine orders. Write clear service descriptions. Use structured pricing. Make your content readable by machines.
Conversational commerce agents A user asks an agent to find and buy something. The agent handles the whole process in one session. Keep brand data consistent across all feeds. Reviews and availability need to be indexed and accurate.
The brands that perform in agentic commerce have already done the AI visibility work. Clean crawler access and structured content. The setup that gets you cited today is what gets you recommended by an agent tomorrow.

Recommendations

In order of impact.

  • 1
    Fix crawler access first

    Let all live assistant fetchers through. That includes Claude-User and PerplexityBot. Check that blanket bot-block rules are not also catching GPTBot. Perplexity takes 45 to 90 days to re-index.

  • 2
    Audit robots.txt, CDN settings, and WAF rules

    A robots.txt audit alone misses CDN and WAF rules. Whitelist live fetcher user agents in robots.txt before any general block rule. Then check CDN allow/deny settings. Then check WAF rate limits for crawler traffic. All three can block crawlers independently.

  • 3
    Map your zero-citation verticals

    Every client had at least one commercially important vertical with zero citations. Find yours. Run queries your buyers actually use. Note which verticals never produce your brand. Those are your priority content gaps.

  • 4
    Reformat for AI extraction

    Put a direct one-sentence answer in the first 150 words of every key page. Change H2 headings to questions. Replace hedged claims with single declarative sentences. Cite sources inline.

  • 5
    Add geographic specificity

    Regional brands need explicit location terms in their content. City names, neighborhood names, named local landmarks. Every citation in this study that came from a location query required the location to be named in the content.

  • 6
    Build external citation presence for Perplexity

    Perplexity does not cite brand pages directly. It cites third-party sources. Get coverage in review platforms, industry directories, and partner documentation. Client D's 23 citations all came through third-party sources.

  • 7
    Start tracking AI referral traffic now

    Set up UTM source tracking for AI platforms in your analytics. Segment bot traffic from user traffic. Establish a baseline before you make changes. You cannot measure improvement without a starting number.