An analytical audit of directory presence across seven Miami-area AI-search providers, drawing on Ahrefs and Semrush backlink data collected 14 July 2026. This paper distinguishes between free auto-approve directories, vetted business directories, vertical-specific directories, and PR-sourcing tools often conflated as "citations."
.shop domains across six of seven providersWhen the SEO industry retired "directory building" as a practice in 2012, at least three distinct activities appear to have been abandoned simultaneously. Only one of those may have been the intended target. The type of directory that most plausibly matters for AI citation — via Google Knowledge Graph inheritance — is the type most consistently absent from the sampled providers' profiles.
This report is an analytical companion to the Miami AI Search Visibility Study and the earlier sector observation report on backlink patterns across Miami AI-search providers. It presents a testable hypothesis rather than an empirical claim. See Section 2 for limitations.
We audited the backlink profiles of seven Miami-area agencies marketing AI-search services, using Ahrefs and Semrush data collected on 14 July 2026, and classified their referring domains against a five-category taxonomy distinguishing free auto-approve directories, vetted business directories, vertical-specific verified directories, PR-sourcing tools, and cheap-TLD link farms. The sampled providers exhibited a fractured directory strategy in which the categories bore no consistent relationship to one another: five of seven still used auto-approve directory farms that the SEO industry has publicly considered obsolete since 2012, only two of seven maintained presence on vetted business directories with editorial or membership thresholds, and none of the seven had built any local Miami editorial coverage. We propose — as a testable hypothesis, not a demonstrated finding — that AI retrieval systems may inherit entity trust signals from Google's Knowledge Graph, which is documented to ingest from vetted directories but not from auto-approve directories. If that dependency chain holds, the type of directory most consistently absent from the sampled providers is precisely the type most plausibly relevant to AI citation. This paper does not test that hypothesis. It identifies the evidence gap, proposes a study design that would close it, and documents observations sufficient for further researchers to replicate the classification exercise.
This paper investigates:
The paper is analytical rather than empirical: it does not measure AI citation outcomes directly, and it does not test causation. It documents observed patterns and identifies an evidence gap in the industry's public knowledge.
In April 2012, Google released the Penguin update. The SEO industry read Penguin — along with the disavow tool that arrived later that year, the parallel Panda updates, and Matt Cutts's public commentary — as a coordinated statement that the era of directory-based link-building was over. Submitting a business to hundreds of low-quality directories to inflate PageRank was declared dead. Practitioners largely accepted this. Guidance shifted toward earned editorial coverage and content marketing. "Citations" as a stand-alone tactic fell out of favour.
Fourteen years later, a new form of search retrieval has emerged in the form of generative AI systems — ChatGPT, Claude, Gemini, Perplexity, Google's AI Overviews and AI Mode. These systems do not rank pages the way traditional search did. They select and cite sources when composing answers, drawing on training corpora, retrieval-augmented generation pipelines, and, in several documented cases, external knowledge graphs.
This raises a question the industry has largely proceeded to answer by assumption rather than measurement: was the 2012 declaration correct for the new context? Or, more precisely — because the answer is unlikely to be a simple yes or no — which of the several distinct activities lumped together as "directory building" may or may not matter for AI citation? This report does not resolve that question. It attempts to sharpen it, using observations from a July 2026 audit of seven Miami-area providers marketing AI-search services.
Backlink analysis cannot determine intent or responsibility. A website's backlink profile may include links created by previous SEO providers, automated indexing services, unsolicited third parties, or historic campaigns that are no longer active. This analysis cannot determine whether a domain owner has submitted a Google Disavow file or taken other remediation steps. Accordingly, this report documents observable patterns rather than attributing motive or intent to any individual provider. Any pattern documented below may be attributable to one or more of the following:
Additionally, this report presents a hypothesis about the relationship between vetted directory presence and AI citation via Knowledge Graph inheritance. It does not present empirical measurement of causation. No controlled study of directory-type X versus AI-citation-outcome Y has been performed here, or, as far as we can determine, published elsewhere at time of writing. Section 6 outlines the study design that would be required to test the hypothesis rigorously. In the meantime, the report describes what is observable and identifies where the evidence gap lies.
A further methodological point: this study measures observable backlink presence rather than the underlying retrieval mechanisms of AI systems. The internal ranking or retrieval systems used by ChatGPT, Claude, Gemini, and Perplexity are proprietary and were not observable. Inferences about how those systems weight external signals are therefore reasoned from documented architectural descriptions and empirical citation patterns published elsewhere, not from direct examination of the systems' internals.
Providers are anonymised throughout this report as Agency A–G. This is a sector observation, not a call-out of any individual operator. The evidence is recoverable by any reader with an Ahrefs or Semrush account.
Much of the industry's discussion of "directory backlinks" or "business citations" collapses several categorically different phenomena into one label. The observations in Section 4 depend on distinguishing between them.
Sites such as freelistingusa.com, bizidex.com, agreatertown.com, find-us-here.com, citysquares.com, whatsyourhours.com, and their family variants (freelistinguk.com, getlistedcanada.com, and similar). These accept submissions with no editorial threshold; anyone can create a listing in minutes. Some carry a Domain Rating in the 60–75 range because the platform itself has volume — not because the individual listing has been vetted.
Sites such as chamberofcommerce.com, bbb.org (Better Business Bureau), dnb.com (Dun & Bradstreet), and crunchbase.com. Presence on these requires membership, business registration, editorial review, or a verified profile. A listing on Chamber of Commerce or BBB provides some evidence that the business exists as a genuine entity with an identifiable address and contact record.
Category-relevant sites: healthgrades.com, zocdoc.com, and vitals.com for health; avvo.com, justia.com, and martindale.com for legal; zillow.com and loopnet.com for real estate; clutch.co, goodfirms.co, designrush.com, and agencies.semrush.com for B2B agencies. These carry categorical signal — appearing on Avvo indicates the business is a law firm — which auto-approve directories do not carry.
Sites such as connectively.us (a successor to HARO / Help a Reporter Out), qwoted.com, featured.com, and muckrack.com. These are databases connecting individual subject-matter experts with journalists. A profile on Connectively is a personal PR-sourcing entry, not a business listing. The high Domain Rating of these platforms causes them to be counted as "authority backlinks" by amateur audits, but the semantic meaning of a personal expert profile is different from a vetted business directory listing.
A fifth category deserves brief mention: article syndication sites (articlescad.com, ezinearticles.com, and similar). These accept unedited content submissions and are widely used by PBN operators. They are frequently counted as "citations" but have no directory function.
The single label "directories" therefore obscures at least four categorically different activities: farming free auto-approve listings, maintaining vetted business directory presence, appearing in vertical-specific verified directories, and holding PR-sourcing profiles. Lumping them together produces meaningless conclusions in either direction — for or against — because the four categories almost certainly carry different weights in any downstream retrieval system.
We audited the backlink profiles of seven Miami-area agencies marketing AI-search services. For each provider we filtered to referring domains with Domain Rating ≥ 20 that were not flagged as spam by Ahrefs — the "real" tier of each profile — and classified each referrer manually into the categories set out in Section 3. The distribution:
| Agency | Total real ref. domains | Editorial | Educational | Vetted business dir | Vertical dir | Free auto-dir | PR-sourcing tool |
|---|---|---|---|---|---|---|---|
| A | 7 | 0 | 0 | 0 | 0 | 2 | 1 |
| B | 26 | 0 | 0 | 0 | 0 | 1 | 2 |
| C | 14 | 0 | 1 | 0 | 0 | 0 | 0 |
| D | 19 | 0 | 0 | 0 | 1 | 3 | 3 |
| E | 116 | 2 | 0 | 1 | 4 | 10 | 0 |
| F | 1,080 | 2 | 13 | 2 | 9 | 13 | 0 |
| G | 11 | 0 | 0 | 0 | 0 | 3 | 0 |
Several patterns are worth stating explicitly. These are observations about the sampled providers' backlink profiles, not claims about causal effect on AI citation.
Five of the seven sampled providers still use free auto-approve directories — the category of listing the industry has publicly considered obsolete for over a decade. Only two of the seven have any presence on vetted business directories. Three of the seven use PR-sourcing tools that appear to be counted as "directory citations" in amateur audits but are semantically different. And one provider (F) has substantial editorial and educational presence that the others do not. These are consistent with at least three separate underlying practices operating in parallel across the sample, rather than a single coherent strategy. This report does not investigate why.
Where vertical B2B agency directories appear (D, E, F), they are drawn from the same short list: designrush.com, goodfirms.co, themanifest.com, techbehemoths.com, selectedfirms.co, expertise.com. These are the directories that in fact appear in the top ten Google results for commercial queries such as "best AI search agency Miami" — agencies.semrush.com in particular ranks three separate pages for that query. Presence on these directories may function as both a buyer-touchpoint and an indirect ranking signal for the underlying commercial query, though this report does not test either mechanism.
Not one of the seven sampled providers has a single referring domain from a Miami-area publication in the DR≥20 non-spam tier: miamiherald.com, sun-sentinel.com, southfloridareporter.com, miamilivingmagazine.com, patch.com, wplg.com, nbcmiami.com, and their peers appear zero times across all seven profiles. Every one of the seven providers markets services to Miami-area businesses. None has built its own local editorial evidence.
Provider F carries Forbes and Entrepreneur editorial coverage, thirteen .edu referrers (Cornell, Yale, Rice, and others), and a BBB profile. This is atypical of the sample. It is also the only provider whose backlink profile approaches what would conventionally be recognised as a genuine reputation footprint. Whether this presence is associated with above-average AI citation frequency for provider F is not tested here.
The observed results are consistent with a hypothesis that AI systems preferentially surface entities with stronger external confirmation signals via vetted directories and editorial coverage. However, the present study does not directly test this mechanism and cannot establish causality. Section 5 develops the hypothesis; Section 6 outlines the study design that would be required to test it.
The observations in Section 4 describe a pattern. They do not, in themselves, explain why any particular directory type might or might not matter for AI citation. The following mechanism is offered as a testable hypothesis, not as a proven claim.
Google's Knowledge Graph is a documented, extensively-referenced entity graph that Google uses to disambiguate businesses, people, organisations, and other entities across its products. It is populated from a specific and limited set of sources — Wikipedia and Wikidata; government registries; vetted business directories (Chamber of Commerce, BBB, Dun & Bradstreet); vertical-specific authorities (Healthgrades for medical professionals, Avvo for lawyers, published academic sources for researchers); and, in recent years, structured data on the businesses' own websites where entity consistency can be verified.
Several generative AI systems have documented or plausible dependencies on Google's Knowledge Graph. Google Gemini and Google's AI Overviews use it directly, as documented in Google's own technical descriptions. Perplexity's retrieval layer depends heavily on Google-indexed properties. ChatGPT via its search integration with Bing inherits partial Knowledge Graph signals through Bing's own entity graph, which shares structural sources with Google's. Anthropic's Claude uses undisclosed retrieval pipelines but has shown empirical citation behaviour consistent with weighted entity signals from these graphs.
If this dependency chain holds — Knowledge Graph reads vetted directories, and AI retrieval reads Knowledge Graph — then the practical implication is that the type of directory presence most likely to influence AI citation is precisely the type most systematically absent from the sampled Miami providers. Auto-approve directories, PR-sourcing profiles, and site-analysis-tool spillover are not ingested by Knowledge Graph. Chamber of Commerce, BBB, Crunchbase, Wikidata, and vertical-specific verified directories are.
Restated as a testable hypothesis: vetted business directory presence and vertical-specific verified directory presence are more strongly associated with AI citation frequency than either auto-approve directory presence or the absence of any directory presence. This report does not test that hypothesis. It observes that the industry has largely abandoned the category of directory presence the hypothesis identifies as most likely to matter.
The hypothesis in Section 5 is not novel; variants of it are widely discussed in industry commentary. What is missing from the public record is a controlled test. The study design that would settle the question:
To our knowledge, no study of this design has been published as of the date of this report. The industry is proceeding on assumption in every direction — including the assumption that directory submissions are dead, and the assumption that directory submissions boost AI citation. Neither assumption is currently supported by the kind of controlled evidence that would justify confidence in either direction. This report identifies an evidence gap the industry — including AEOGeoAI — has an obligation to close through published research rather than through marketing copy.
The recommendations below reflect the balance of low-risk actions with plausible mechanisms — not proven-effective interventions.
freelistingusa.com family, along with bizidex.com, agreatertown.com, apsense.com, citysquares.com, and their peers, have no editorial threshold and are documented as low-trust by Ahrefs' spam classifier. They pollute the backlink profile without any documented benefit..shop-domain link-farm packages. Documented in the AEOGeoAI sector observation report as the dominant hidden practice in the Miami AI-search agency market. Spam-flag pollution of the backlink profile is well-documented; any AI-citation benefit is not.A note on epistemic modesty. The recommendations above are directional, not certain. They reflect the best current understanding of how AI retrieval systems build entity signals, and they favour actions with low downside risk and plausible mechanisms. They do not guarantee AI-citation lift, and providers who promise otherwise — including AEOGeoAI itself — should be asked to show controlled evidence.
Providers were identified from Miami-area agencies actively marketing "AI search," "AI SEO," "AEO" (Answer Engine Optimization), "GEO" (Generative Engine Optimization), or "AI visibility" services, as of July 2026. Inclusion was based solely on the provider's public marketing self-description and geographic scope. Neither backlink profile characteristics nor AI-citation performance played any role in inclusion or exclusion. Seven providers meeting these criteria were audited.
Backlink data for each provider was collected on 14 July 2026 from Ahrefs Site Explorer (referring domains, anchor text, Domain Rating, Ahrefs proprietary spam flag) and cross-referenced against Semrush Backlink Analytics collected the same day. Referring domains were filtered to those with Domain Rating ≥ 20 and not flagged as spam by Ahrefs; this "real tier" was retained for classification.
Retained referring domains were classified using a five-category taxonomy (free auto-approve directory, vetted business directory, vertical-specific verified directory, PR-sourcing tool, cheap-TLD link farm) together with residual categories (major editorial, local editorial, educational, wire service, platform/UGC, article syndication, site-analysis tool, and unclassified). Classification proceeded by domain-name pattern-matching against published lists of known category members, followed by manual review of the residual unclassified set for each provider. For the six smaller provider profiles, approximately 76% of referrers were classified with high confidence. Provider F's larger profile (1,080 domains) contains a substantial long tail of individually-classified referrers not exhaustively categorised; only the top categories are reported for F. Full classification lists are available in the archived dataset (see Section 10).
Another researcher following the published sample criteria, data collection dates, filtering thresholds (DR ≥ 20, non-spam), and classification taxonomy set out above should obtain qualitatively similar results, subject to the following sources of expected variance: Ahrefs and Semrush data are continuously updated, so a repeat audit on a later date may show altered referring-domain counts as new links are indexed or existing links are lost; the Ahrefs spam flag is a proprietary classifier and may re-classify individual domains over time; and any manual classification exercise carries a residual boundary-case error rate that a second reviewer may resolve differently in ambiguous cases (estimated at less than 5% for the top-DR tier). The direction and rough magnitude of the observations are expected to be stable across reasonable reviewer choices; specific counts may vary.
Provider identities are anonymised throughout this report. The data is recoverable by any reader with Ahrefs or Semrush subscriptions and a list of Miami-area agencies marketing the services listed in "Sample selection" above.
The observations and hypothesis in this paper suggest several directions for further investigation. AEOGeoAI intends to pursue several of the following, and openly invites other researchers to pursue any of them independently.
The classification taxonomy and the anonymised aggregated results (category counts per anonymised provider) presented in Section 4 are archived on Zenodo as supplementary material to this report, under a Creative Commons Attribution 4.0 International licence.
The raw referring-domain lists from Ahrefs and Semrush are commercial data subject to those platforms' terms of service and are not redistributed. Readers with Ahrefs or Semrush subscriptions can reproduce the underlying data by applying the sample-selection criteria in Section 8 to the same platforms at the same collection date. Provider identities are withheld consistent with the report's anonymisation policy; a reader who wishes to reconstruct the specific sample used here may do so by searching Google for Miami-area agencies marketing the services described in Section 8, applying reasonable judgement to include or exclude edge cases.
Correspondence regarding the classification taxonomy, or requests for further methodological detail, should be directed to [email protected].
@techreport{wilde2026directories,
author = {Wilde, Kevin H.},
title = {Directories Were Declared Dead in 2012. AI Didn't Get the Memo.},
year = {2026},
month = {July},
institution = {AEOGeoAI Research},
url = {https://aeogeoai.net/directories-ai-citation-report},
note = {ORCID: 0009-0002-3753-5482}
}
A permanent DOI-registered version of this report is available on Zenodo (DOI to be added on publication). This site version is the canonical reference; the Zenodo record is a mirror.
This report is part of the AEOGeoAI research programme on AI search visibility. See also the Miami AI Search Visibility Study 2026 (515 businesses tested across ChatGPT, Claude and Gemini) and all published research.
Browse all research →Open access · DOI-registered · CC BY 4.0