All Research →
AEOGeoAI Research Report · July 2026

Category Substitution and Retrieval-Layer Fragmentation in Generative AI Systems.An Empirical Comparison of Base LLM APIs and Consumer AI Search Products.

A structured prompt study of 360 base LLM API responses across three models and five query types, with comparative consumer-product data from ChatGPT, Perplexity, Claude.ai, and Gemini. The study documents how AI systems behave when asked to recommend providers in a service category that has not yet accumulated sufficient training-data coverage to support confident naming.

Author: Kevin H Wilde ORCID: 0009-0002-3753-5482 Type: Empirical research paper Data collected: July 2026 Analysis: July 2026 Report version: 1.0 Published: 21 July 2026 License: CC BY 4.0
AI citation research
Research summary
Paper: Category Substitution and Retrieval-Layer Fragmentation in Generative AI Systems
Sample: 40 queries in 5 groups × 3 LLMs × 3 repetitions (360 API responses); ~20 consumer-product responses across 4 products
Models (base API): Claude Haiku 4.5, GPT-4o-mini, Gemini 2.5 Flash Lite
Consumer products: ChatGPT (web, incognito), Perplexity (web), Claude.ai (web), Gemini (web)
Type: Empirical — direct measurement of LLM response behavior across structured query conditions
Central observations:
  • Base LLM APIs exhibit three concurrent failure modes in emerging B2B service categories: category refusal, category substitution, and firm hallucination.
  • Consumer products with retrieval augmentation name real specialist providers, but no single provider was named by all four consumer products tested.
  • URL citation is essentially absent from base API responses (357 of 360 contained no external URLs), regardless of query type.
  • Response length correlates inversely with model willingness to name providers, providing a quantitative confidence proxy across query types.
  • A category-age effect is observed: adjacent categories established 3–5 years prior produce correct specialist naming in base APIs; the AI-search-agency category (established <2 years) does not.
  • 100% consistency across all three repetitions per cell — categorical response behavior in LLMs at this temperature setting is highly stable.
360
base API responses across 3 models, 40 queries, 3 repetitions
0/357
base API responses containing any external URL citation
0/4
consumer AI search products that agreed on the same set of named providers
100%
consistency across 3 repetitions per query × model cell
Central observation

Base LLM APIs, when queried about providers in emerging B2B service categories without retrieval augmentation, respond with incorrect-category substitutions rather than either correct naming or honest refusal. The same underlying model families — in their consumer-product configurations with retrieval layers enabled — name real specialists with source citations, but disagree substantially with each other. The industry's "AI citation" discourse conflates these two categorically different retrieval modes.

This report is an empirical companion to the Directories Were Declared Dead in 2012 analytical report and builds on methodology established in the Miami AI Search Visibility Study 2026. Unlike those reports, which used consumer-facing AI products, this study isolates and compares base API behavior from consumer-product behavior to identify where naming and citation behavior originates.

Abstract

Abstract

We conducted a structured prompt study across three base LLM APIs — Claude Haiku 4.5, GPT-4o-mini, and Gemini 2.5 Flash Lite — using 40 queries in five groups: ten emerging B2B commercial queries targeting AI-search-agency services, ten mature commercial controls across diverse industries, six informational controls, five adjacent-category calibration queries, and nine keyword-phrase queries replicating real SEO target terms. Three repetitions were collected per query-model cell, yielding 360 total API responses. We additionally collected approximately twenty responses from four consumer AI products — ChatGPT, Perplexity, Claude.ai, and Gemini — for the same buyer-intent queries.

Base LLM APIs exhibited three concurrent failure modes when asked to recommend providers in the AI-search-agency category: category refusal (declining to name specific providers), category substitution (naming tools, PR firms, and management consultancies as replacements), and firm hallucination (generating plausible-sounding but unverifiable local provider names). These failure modes did not appear in the mature commercial control group, where all three models freely named specific firms across diverse industries. Consumer products with retrieval augmentation named real specialist providers for the same buyer queries, but exhibited substantial fragmentation: the sets of named providers showed minimal overlap across the four consumer products tested, and no provider was named by all four.

URL citation was essentially absent from base API responses (357 of 360 contained no external URLs), regardless of query type. A category-age pattern was observed, consistent with the hypothesis that the accumulation of publicly documented specialist coverage affects base-model provider naming: adjacent categories established approximately three to five years prior to the study (Web3 marketing agencies, AI voice-agent firms, ESG consultancies) produced correct specialist naming in base APIs; the AI-search-agency category — established less than two years prior — did not. Response length correlated inversely with model willingness to name providers, providing a quantitative proxy for model confidence across query types. All 120 query-model cells were fully consistent across three repetitions, indicating that categorical response behavior in LLMs is highly stable at standard temperature settings.

These findings suggest that the industry discourse on AI citation optimisation frequently conflates two categorically different retrieval modes — base LLM training-data response and retrieval-augmented consumer-product response — which produce different outcomes by different mechanisms and are not interchangeable targets for a single optimisation strategy.

Research questions

Research questions

This paper investigates:

The paper is empirical: it reports measured response behavior across a structured prompt battery. It does not test causal interventions and does not claim to explain the internal mechanisms of any specific AI system.

Section 1

Context: the AI citation industry and its measurement gap

An industry has emerged around the premise that businesses can be "optimised" to appear in the responses of AI systems such as ChatGPT, Claude, Gemini, and Perplexity. Practitioners variously describe this as Answer Engine Optimisation (AEO), Generative Engine Optimisation (GEO), AI Search Optimisation (AIO), or Large Language Model Optimisation (LLMO). The services offered typically include entity-chain building, schema markup, structured data optimisation, and editorial link acquisition.

Several published studies have examined AI citation patterns. SE Ranking analysed 75,550 AI Overview responses and found that only 20.85% cited any news source, with the top three publishers (BBC, The New York Times, CNN) accounting for 31% of all media mentions. Ahrefs studied 75,000 brands and found that branded web mentions showed the strongest correlation with AI Overview brand visibility (Spearman's r = 0.664), significantly outperforming backlinks (r = 0.218). Pew Research Center found that 58% of US adults encountered an AI-generated summary in at least one search during March 2025. The Columbia Journalism Review documented consistent patterns of incorrect citation, fabricated URLs, and confident wrong answers across eight AI search tools.

However, the measurement approach used across these studies has a structural limitation that this paper attempts to address: they measure AI Overview behavior (Google's retrieval-augmented search summary layer) or consumer-product behavior (ChatGPT with browsing enabled, Perplexity with its crawler active), rather than base LLM API behavior. The two modes are substantially different in architecture. An AI Overview is a retrieval-augmented generation (RAG) system that fetches real-time web content before generating an answer. A base LLM API call with no tools enabled generates from training weights alone. These are not interchangeable systems, and it does not follow that findings in one mode generalise to the other.

This study was designed specifically to expose and quantify that gap, using a five-group query design that allows direct comparison of base API behavior with consumer-product behavior for the same buyer queries.

Section 2

Important limitations

Several methodological constraints apply to this study's findings.

Model versions and configuration

The base API study used Claude Haiku 4.5, GPT-4o-mini, and Gemini 2.5 Flash Lite — the lightweight tiers of each model family, selected for cost efficiency in a 360-response batch study. Consumer-product testing used the default web configurations of each product, which may include larger model variants, tool integrations, and personalization layers not present in the API calls. Differences in observed behavior may therefore reflect model size, tool availability, or both, and cannot be cleanly attributed to either factor alone.

Consumer-product sample size

Consumer-product responses were collected manually across approximately five queries per product, using incognito or fresh-account sessions where possible. This sample is sufficient to establish qualitative behavioral patterns but is too small for statistical inference. The consumer-product data is presented as illustrative comparison rather than as a parallel controlled study.

Geographic and personalization variance

Consumer AI product responses vary by user location, prior session context, account history, and time of query. All consumer-product testing was conducted from the researcher's location during a single session window in July 2026. Results may not generalise to other user contexts.

The firm-name extraction results

Firm names were extracted from API responses using pattern-matching against bold-formatted names and capitalized multi-word entities followed by common firm-type suffixes. This method captures a subset of named firms and introduces classification noise: some extracted strings are not firm names, and some firm names in responses were not captured. The extraction results are used to illustrate qualitative patterns rather than to generate precise counts.

This study measures response behavior — what AI systems say — rather than the underlying mechanisms that produce that behavior. The internal retrieval logic, training data provenance, and ranking signals of ChatGPT, Claude, Gemini, and Perplexity are proprietary and were not observable. All mechanism-level claims in this paper are inferences from behavioral observation, not direct measurement.

Section 3

Methodology

Infrastructure

Base API calls were made via a Cloudflare Worker endpoint routing through Cloudflare AI Gateway. Each query was sent to each model independently with no shared context (no system prompt other than Gemini's standard helpfulness instruction), a maximum token limit of 2,000 tokens, and default temperature settings. Results were collected as structured JSONL with timestamp, model identifier, query identifier, response text, response length, and URL extraction from response text.

Query set design

Forty queries were distributed across five groups:

Repetitions and consistency

Each query was submitted to each model three times independently. This yielded 360 total responses (40 queries × 3 models × 3 repetitions). All 120 query-model cells were fully consistent across repetitions in their watchlist-mention outcomes — no query-model combination that named a watched firm in one repetition failed to name it in another, and vice versa.

Consumer-product comparison

Five buyer-intent queries from Groups A and E were submitted to ChatGPT (web, incognito session), Perplexity (web), Claude.ai (new account, fresh sessions), and Gemini (web, incognito) in July 2026. Responses were captured in full including citation panels and source links where present. Consumer-product testing was not repeated across sessions and is presented as qualitative comparison rather than controlled parallel data.

Watchlist

Nine agency names were monitored for mention across all 360 base API responses: AEOGeoAI, FuelOnline, ProCloser, SkySEO Digital, Hyperlink Miami, Pivotal Consulting, Scalz, SEOSmooth, and Goodie AI. None appeared in any base API response across any query type.

Conflict-of-interest disclosure

This study was designed and conducted by the founder of AEOGeoAI, a Miami-based AI search optimization service. AEOGeoAI was included in the monitored provider set and appeared in two of four consumer-product comparisons. This creates a potential researcher-conflict and self-observation issue, particularly in the consumer-product comparison section. To reduce interpretive bias, all consumer-product responses were retained in their entirety and are included in the supplementary archive on Zenodo. Readers are encouraged to verify the consumer-product results independently using the queries documented in Section 3. No base API result was altered or excluded based on its outcome. The base API watchlist produced zero mentions of AEOGeoAI across all 360 responses, which is the finding the researcher had the greatest incentive to soften; it is reported as observed.

Section 4

Results

4.1 Response-length gradient by query group

Average response length varied systematically across query groups, with mature commercial queries producing the shortest responses and informational queries the longest:

Query groupAvg response length (chars)URL cite rateHedging rate
B — Mature commercial5770%4%
A — B2B emerging8430%9%
E — Keyword phrases1,0490%4%
D — Adjacent category1,2784%20%
C — Informational1,7330%2%

Mature commercial queries produced responses approximately 35% shorter than B2B emerging queries and 66% shorter than informational queries. This gradient was consistent across all three models. The pattern suggests that response length in these models correlates inversely with confidence in providing a direct named recommendation: mature categories produce short, decisive answers; emerging categories produce longer categorical explanations in place of names; informational queries produce extended teaching-mode responses.

Response length functions as a quantitative proxy for model confidence in provider-naming. The shorter the response, the more directly the model named specific firms. This gradient was consistent across all three models and all three repetitions.

4.2 URL citation

Of 360 base API responses, 357 contained zero external URLs. Three URL citations appeared in Group D (adjacent category) responses; all three were classified as "other" domain types rather than editorial or directory sources. No news citations, no directory citations, and no agency website citations appeared in any base API response across any query group.

This stands in direct contrast to published industry research on AI Overview citation behavior, which documents news source citation rates of 20.85% across 75,550 responses (SE Ranking, 2025). The discrepancy is methodological: AI Overviews are retrieval-augmented systems that fetch web content before generating responses; base LLM API calls do not. The observed behavior is therefore consistent with architecture, not a deviation from expected behavior in consumer products.

4.3 Firm naming in base API responses

Using pattern-matching extraction of bold-formatted and capitalised multi-word entities from all 360 responses, the following firm-naming patterns were observed by group:

Query groupNamed firms (representative sample)Category accuracy
B — Mature commercialThrive Internet Marketing, Neil Patel Digital, WebFX, Ignite Visibility, CBRE, Colliers International, Greenberg Traurig, Akerman, Deloitte, KPMGCorrect
D — Adjacent categoryCoinbound, NinjaPromo (Web3); Voiceflow, Nuance, Rasa (AI voice); Sustainalytics, EcoAct (ESG); Shopify Plus/Partners (Shopify)Correct
A — B2B emergingSemrush, HubSpot, Copy.ai, Jasper (tools); Edelman, Weber Shandwick (PR); Boston Consulting, Accenture, Deloitte (management consulting)Wrong category
E — Keyword phrasesFragments: "AI Search," "Miami AI Search," "Digital," "Social" — no complete firm namesNone identified
C — InformationalNo firm names — teaching-mode responses onlyN/A

The mature commercial and adjacent-category groups produced correct specialist naming consistently. The B2B emerging group produced systematic wrong-category substitution: marketing software tools, PR agencies, and management consulting firms were named as substitutes when no AI-search-agency specialists could be identified from training data.

4.4 Category-age pattern

The contrast between Groups B, D, and A reveals a pattern consistent with a training-data accumulation hypothesis. The study does not establish category age as the causal variable — the sample covers only three age brackets, and other co-varying factors (category volume in public documentation, number of named specialists in trade press, platform coverage) may be the operative mechanism. With that caveat, the observed gradient is:

This gradient is consistent with a training-data accumulation account. The study does not establish category age as the sole causal variable; the operative factor may be the volume of publicly documented specialist coverage in a category, which tends to correlate with category age but is not identical to it. Categories that had not yet accumulated substantial coverage in editorial publications, agency directories, review platforms, and service-category reference pages — for whatever reason — appear to produce substitution behavior rather than correct naming or honest refusal.

4.5 Three concurrent failure modes in emerging B2B queries

Analysis of individual Group A and E responses reveals three distinct failure modes operating simultaneously:

Mode 1: Category refusal. Models explicitly declined to name specific providers. Claude Haiku's response to "Which agency should I hire to help my business rank in ChatGPT?" began: "I'd gently push back on the premise here. There's no meaningful way to 'rank in ChatGPT' like you would in Google search results." This is not mere uncertainty — it is an active challenge to the question's framing, informed by the model's training-data understanding of what AI systems do.

Mode 2: Category substitution. When models did name firms, they named firms from adjacent categories rather than AI-search-agency specialists. Across Group A responses, the most frequently named entities were Semrush (12 mentions), Edelman (4), HubSpot (4), Copy.ai (4), Weber Shandwick (3), Boston Consulting (3), and Accenture (2). None of these are AI-search agencies.

Mode 3: Firm hallucination. Some responses generated plausible-sounding local firm names — for example, "AI Search Solutions Miami" — that do not correspond to any documented business. These names follow naming conventions for the category (tool-adjacent, AI-forward branding, local suffix) and would be indistinguishable from real firms to a buyer with no prior knowledge of the market.

The three failure modes are not mutually exclusive. A single response may contain all three: a premise challenge (Mode 1) followed by tool-firm suggestions (Mode 2) and a plausibly-named local firm that does not exist (Mode 3). Buyers reading these responses have no indication that the content is unreliable.

4.6 Consumer-product comparison

Consumer AI products with retrieval augmentation produced qualitatively different responses to the same buyer queries. The following products were tested for five buyer-intent queries from Groups A and E. All four named real specialist firms and cited external sources — behavior absent from base API responses.

The headline finding in the consumer-product comparison is not which specific agencies were named. It is that no provider was consistently named across all four consumer products. The sets of named providers showed minimal overlap, confirming that "AI visibility" in consumer products is not a single unified outcome but a product-specific one.

ProductRetrieval architectureSources cited?Provider count (approx.)Notable named providers (sample)
ChatGPT (web)Bing search integrationYes12+AEOGeoAI, Elevate AI Consulting, Sky SEO Digital, Anderson Collaborative, Miami GEO Pro
Perplexity (web)Own crawler + open webYes3AEOGeoAI, Hyperlink, IRPR Agency
Gemini (web)Google Knowledge Graph + SearchYes15+Onely, iPullRank, First Page Sage, Omniscient Digital, Republica Havas, Anderson Collaborative
Claude.ai (web)Undisclosed (Brave Search likely)Yes10+White Shark Media, Fuel Online, Sky SEO Digital, Fisher Agency, Doc Digital SEM

Cross-product overlap was minimal. Sky SEO Digital appeared in both ChatGPT and Claude.ai responses. Anderson Collaborative appeared in both ChatGPT and Gemini responses. No provider was named by all four products. The four products effectively returned four different shortlists for the same buyer intent.

4.6a Note on AEOGeoAI's appearance in the consumer-product comparison

AEOGeoAI — the organisation that conducted this study — was named in two of the four consumer products tested (ChatGPT and Perplexity). This result is reported here in full in the interests of transparency, consistent with the conflict-of-interest disclosure in Section 3.

This result should not be interpreted as a ranking claim or a measure of market position. The consumer-product comparison was an exploratory sample of approximately five queries per product, not a statistically representative survey. The queries were limited to Miami-specific buyer intent. AEOGeoAI's appearance in these results reflects its current presence in the open-web retrieval indices used by ChatGPT and Perplexity; it does not predict visibility in Gemini or Claude.ai, or in different query contexts.

The scientifically relevant observation is the fragmentation finding — that four consumer products returned four different provider shortlists for the same buyer intent — of which AEOGeoAI's results are one data point. The fragmentation finding holds regardless of whether AEOGeoAI appeared in any product's responses.

4.7 Circular citation pattern

In Claude.ai responses, the source most frequently cited for Miami AI-search-agency recommendations was content published by Fuel Online — a competitor in the same category. Fuel Online had published a "Top AI SEO Agencies in Miami" article that ranked Fuel Online prominently. Claude.ai retrieved and cited this article, and Fuel Online appeared prominently in the resulting recommendation. This creates a potentially circular discovery pattern: a provider's own promotional content may become a source used by an AI system to support recommendations of that same provider.

This observation is reported as a behavioral pattern, not as an established mechanism. The study does not establish whether Claude.ai treated the article as independent editorial validation or simply retrieved the highest-ranked available source for the query. This pattern was predicted in the AEOGeoAI directories report (July 2026), which identified the absence of authoritative third-party editorial content as a structural gap in the sampled providers' profiles.

4.8 Repetition consistency

All 120 query-model cells (40 queries × 3 models) were fully consistent across three repetitions. No cell produced a different watchlist-mention outcome in any repetition. This finding challenges the common assumption that LLM responses are highly variable, at least for categorical naming behavior at standard temperature settings. The consistency finding has methodological implications: a single repetition may be sufficient for categorical studies of this type, reducing study cost without sacrificing reliability.

Scope boundary

What this study does not show

This study does not establish that publishing on a particular website, acquiring backlinks, creating a Wikidata item, building schema markup, or performing any other external optimization activity causes a business to appear in AI-generated recommendations. It does not establish that any single retrieval source is a universal ranking factor. It does not establish that the category-age pattern observed is causal, or that category age specifically (as opposed to the volume of publicly documented specialist coverage, which correlates with age) is the operative variable. It does not establish that the consumer-product citation results for AEOGeoAI are attributable to any specific prior action. It establishes that different AI products exhibit materially different provider-naming and citation behavior for the same buyer queries, and that current base LLM APIs and retrieval-enabled consumer AI search products should not be treated as interchangeable systems for the purpose of "AI visibility" optimization.

Section 5

Discussion

5.1 Two retrieval modes, not one

The most significant finding of this study is not any specific pattern within base API behavior or consumer-product behavior separately, but the categorical difference between them. The same query — "Which agency should I hire to help my business rank in ChatGPT?" — produces:

These responses differ not in degree but in kind. The gap arises from retrieval architecture: the consumer product fetches current web content before composing its answer; the base API does not. The same model family, configured differently, produces categorically different outputs for a buyer with real commercial intent.

The industry's discourse on "AI citation optimization" does not consistently distinguish between these modes. Services marketed as "get your business cited by ChatGPT" may be targeting the consumer product (where real-time retrieval determines what gets named), the base API (where training data determines the ceiling of what can be named), or both — with different mechanisms, different timelines, and different levers for each. The present study does not evaluate whether either mode can be reliably influenced by external optimization activities; it establishes only that they are different problems.

5.2 The category-age hypothesis

The gradient from mature commercial naming (correct, confident, short) through adjacent-category naming (correct but hedged) to emerging-category substitution (incorrect, longer, hedged) is consistent with a training-data accumulation account: base LLMs reflect the state of the web at their training cutoff, and categories that had not yet accumulated specialist coverage — in editorial publications, agency directories, review platforms, and service-category reference pages — cannot be named by those LLMs.

Adjacent categories that were established as named verticals in 2021–2023 (Web3 marketing, AI voice agents, Shopify launch agencies, ESG consulting) produced correct naming. The AI-search-agency category — which first appeared in significant public coverage in approximately 2024–2025 — did not. No external optimization activity by any individual provider can materially change this picture until the category itself accumulates sufficient training-data coverage for the models to have specialists to name.

This observation is practically relevant for buyers evaluating AI visibility services. The study does not establish whether external optimization activity can alter base-model provider naming in emerging categories. It does establish that the absence of category-level specialist coverage in training data creates a substantial constraint on what current base APIs can reliably name. Services promising to "get your business cited by ChatGPT" in a category for which ChatGPT's base API has no trained specialists to name should be asked to show evidence of the mechanism by which they propose to achieve that outcome.

5.3 Fragmentation in consumer-product responses

Consumer products with retrieval produce real specialist naming, but the specific firms they name vary substantially across products. No provider was named by all four consumer products tested. This fragmentation has practical implications: buyers using different AI products to research providers receive different recommendations, and there is no single "AI citation" outcome that a provider can optimize toward. Visibility in ChatGPT and Perplexity does not predict visibility in Gemini or Claude.ai, and vice versa.

The fragmentation appears to correlate with retrieval architecture. Products that use open-web crawl retrieval — ChatGPT via Bing integration and Perplexity via its own crawler — returned overlapping but not identical shortlists. Gemini, which uses Google's Knowledge Graph and Search index, and Claude.ai, which appears to use Brave Search or a similar crawler, returned different shortlists again. Different consumer products are drawing on structurally different source sets, and visibility in one product is not transferable to another without separate work on the relevant retrieval infrastructure.

AEOGeoAI appeared in the ChatGPT and Perplexity responses for Miami-specific buyer queries, and not in the Gemini or Claude.ai responses. This is noted as a finding consistent with the retrieval-architecture pattern, and reported in full per the conflict-of-interest disclosure in Section 3. It is not the primary finding of the consumer-product comparison; the primary finding is the fragmentation itself.

5.4 Implications for buyers

Buyers evaluating AI visibility services should apply a small number of practical distinctions that the industry's current marketing copy does not consistently make:

Section 6

Future work

Several directions for further investigation follow from this study's findings. AEOGeoAI intends to pursue several of the following, and openly invites other researchers to pursue any of them independently.

Section 7

Data availability

The full prompt set (40 queries across five groups), the JSONL dataset of 360 base API responses, the watchlist of monitored firm names, and the firm-name extraction output are archived on Zenodo as supplementary material to this report, under a Creative Commons Attribution 4.0 International licence (DOI to be added on publication).

The consumer-product response data (approximately 20 responses across four products) was collected manually and is included in the supplementary archive as plain-text documents. It is not suitable for automated analysis due to sample size and collection variance.

The Cloudflare Worker endpoint code used to batch the base API queries is available on request. It is not publicly released as a general-purpose tool due to the API key management requirements of the Cloudflare AI Gateway integration.

Correspondence regarding the study design, data, or replication should be directed to [email protected].

Reproducibility. The base API component of this study is fully reproducible by any researcher with API access to Claude, GPT-4o-mini, and Gemini. The 40 queries are documented in Section 3 and in the supplementary archive. Results may vary from the present study if model versions or configurations change; the structural pattern — mature categories produce short correct naming, emerging categories produce substitution behavior — is expected to be stable until category-level training-data coverage shifts.

Suggested citation

Suggested citation

APA Wilde, K. H. (2026). Category Substitution and Retrieval-Layer Fragmentation in Generative AI Systems: An Empirical Study of Base API and Consumer Product Responses to Emerging B2B Service Queries. AEOGeoAI Research. https://aeogeoai.net/ai-citation-agency-study-2026
MLA Wilde, Kevin H. "Category Substitution and Retrieval-Layer Fragmentation in Generative AI Systems." AEOGeoAI Research, 21 Jul. 2026, aeogeoai.net/ai-citation-agency-study-2026.
BibTeX @techreport{wilde2026categorysub, author = {Wilde, Kevin H.}, title = {Category Substitution and Retrieval-Layer Fragmentation in Generative AI Systems}, institution = {AEOGeoAI Research}, year = {2026}, month = {July}, url = {https://aeogeoai.net/ai-citation-agency-study-2026}, note = {ORCID: 0009-0002-3753-5482} }
CC BY 4.0 This work is licensed under a Creative Commons Attribution 4.0 International License. You may share and adapt with attribution.

A permanent DOI-registered version of this report is available on Zenodo (DOI to be added on publication). This site version is the canonical reference; the Zenodo record is a mirror.

Related research

This report is part of the AEOGeoAI research programme on AI search visibility. See also the Miami AI Search Visibility Study 2026 (515 businesses tested across ChatGPT, Claude and Gemini), the Directories Were Declared Dead in 2012 analytical report, and all published research.

Browse all research →

Open access · DOI-registered · CC BY 4.0