# AI publisher policy for feedthejoe.com — composed from the organization's engine tree (one module per provider) on 2026-09-28, node e812ddd76b This is a non-standard publisher policy and not a vendor ranking or access-control standard. The authoritative cooperative-crawler rules are in https://feedthejoe.com/robots.txt, generated from the same decisions. Server authentication and deny rules — not this file — protect non-public content. Decisions are recorded on the organizational record (properties.enginePolicy) and approved by the owner. ## What this site is feedthejoe.com is the central unifying node of THATDEVELOPERGUY LLC and its founder: the public record of the organization and the person as it can be proved — ensuring every website of the organization carries the full profile — with the SEO / AIO 26-Discipline Library, the catalogue of every fact with its evidence, and the MEGAMIND research project. ## Sections - The record (https://feedthejoe.com/): The organization (THATDEVELOPERGUY LLC) and its founder as the organizational record supports them: identifiers, registrations, the charter, the record of a working life in approved wording. Section index: https://feedthejoe.com/joseph/llms.txt. - The catalogue (https://feedthejoe.com/verify/): Six registers — identifiers, registrations, profiles, claims, works, credentials — each entry with its approved wording, the registry that holds it and how it is backed; plus how a fact gets onto the site. Section index: https://feedthejoe.com/verify/llms.txt. - The library (https://feedthejoe.com/library/): The SEO / AIO 26-Discipline Library with the Graph Systems expansion — 178 volumes on eight shelves, each with its review stamp; the organization's own methodology, published as written. Section index: https://feedthejoe.com/library/llms.txt. - MEGAMIND (https://feedthejoe.com/megamind/): The experimental federated AGI research project — the Chronicles, concepts, theory, model, papers and resources — preserved as written, with the entities referenced, never re-declared. Section index: https://feedthejoe.com/megamind/llms.txt. ## Google — Google Search, AI Overviews and AI Mode - Googlebot (search): allow; source: node properties.enginePolicy.bots. Google Search crawling; index and snippet rules are decided separately (meta robots / X-Robots-Tag). - Googlebot-Image (search): allow; source: node properties.enginePolicy.bots. Image crawling for Google Images and image features. - Googlebot-News (search): allow; source: node properties.enginePolicy.bots. Google News crawling. - Google-Extended (training): allow; source: node properties.enginePolicy.bots. A product token, not a separate HTTP user agent: controls Gemini training and specified grounding uses; not a Search inclusion switch. - Verification of real requests: reverse DNS with forward confirmation, or the published IP ranges — crawl-*.googlebot.com / *.google.com; JSON IP-range lists are published per crawler class. - What Google asks: No special file or format is required for AI features in Search; the same crawlable, indexable, useful pages qualify. Structured data must match visible content; Organization/Person declared once and referenced. One canonical per URL; sitemaps with accurate lastmod; no soft-404s. Google-Extended is an independent choice from Search access. - Where it reports: Search Console — Performance, and the generative AI performance report (AI Overviews and AI Mode impressions by page, country, device, date; no separate clicks or prompts). - Sources: S02 https://developers.google.com/search/docs/fundamentals/ai-optimization-guide · S03 https://developers.google.com/search/docs/appearance/ai-features · S12 https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers · S13 https://developers.google.com/crawling/docs/crawlers-fetchers/verify-google-requests · S17 https://developers.google.com/crawling/docs/robots-txt/robots-txt-spec · S05 https://support.google.com/webmasters/answer/16984139 ## Microsoft — Bing Search, Copilot and Bing AI answers - Bingbot (search): allow; source: node properties.enginePolicy.bots. Bing Search crawling; the Bing index also grounds Copilot and Bing AI answers. - BingPreview (preview): allow; source: wildcard (no explicit decision on the node yet). Page previews / snapshots for Bing surfaces. - Verification of real requests: reverse DNS with forward confirmation — *.search.msn.com - What Microsoft asks: Verify the site in Bing Webmaster Tools (BingSiteAuth.xml) and submit the sitemap as a feed. Notify changed URLs through IndexNow — a receipt, not indexing or citation. Use the AI Performance report (citation share is relative to all cited sites for a grounding query; not position or traffic). - Where it reports: Bing Webmaster Tools — Search performance, AI Performance (intents, topics, citation share, period comparison). - Sources: S06 https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview · S07 https://blogs.bing.com/search/June-2026/New-AI-Visibility-Insights-in-Bing-Webmaster-Tools-Intents-Topics-Citation-Share-Compare · S30 https://www.indexnow.org/documentation ## OpenAI — ChatGPT search and answers - OAI-SearchBot (search): allow; source: node properties.enginePolicy.bots. Search discovery for ChatGPT search results; separate from training. - GPTBot (training): allow; source: node properties.enginePolicy.bots. Crawling for model training. - ChatGPT-User (user-fetch): allow; source: node properties.enginePolicy.bots. User-triggered retrieval; robots rules may not apply — protection of private content must not depend on robots. Robots rules are only partly honoured for this token by the provider's own account. - OAI-AdsBot (ads): allow; source: node properties.enginePolicy.bots. Ad landing-page workflow; not a prerequisite for organic search visibility. - Verification of real requests: published IP ranges — OpenAI publishes IP address lists per crawler; reverse DNS is not the documented method. - What OpenAI asks: Decide search, training and user-fetch separately; each token is its own choice. Keep public pages crawlable and fast; the search crawler cites the crawled page. Do not rely on robots for anything private. - Sources: S08 https://developers.openai.com/api/docs/bots ## Anthropic — Claude search and answers - Claude-SearchBot (search): allow; source: node properties.enginePolicy.bots. Search discovery for Claude. - ClaudeBot (training): allow; source: node properties.enginePolicy.bots. Crawling for model training. - Claude-User (user-fetch): allow; source: node properties.enginePolicy.bots. User-directed retrieval; follows the stated robots behaviour — confirm real requests in logs. Robots rules are only partly honoured for this token by the provider's own account. - Verification of real requests: published IP ranges — Anthropic documents its crawlers and access information; recheck the official reference rather than a cached list. - What Anthropic asks: Decide search, training and user-fetch separately. Recheck the official crawler reference when policy depends on it. - Sources: S10 https://support.claude.com/en/articles/8896518-does-anthropic-crawl-data-from-the-web-and-how-can-site-owners-block-the-crawler ## Perplexity — Perplexity answers - PerplexityBot (search): allow; source: node properties.enginePolicy.bots. Search discovery, not foundation-model training. - Perplexity-User (user-fetch): allow; source: node properties.enginePolicy.bots. User-directed retrieval; generally ignores robots according to the provider. Robots rules are not honoured for this token by the provider's own account. - Verification of real requests: published IP ranges — Perplexity documents its crawlers and address ranges. - What Perplexity asks: Allow the search crawler for discovery; user-directed fetches happen regardless. Cited pages are the crawled canonical pages. - Sources: S11 https://docs.perplexity.ai/docs/resources/perplexity-crawlers ## Apple — Siri, Spotlight and Safari suggestions; Apple Intelligence - Applebot (search): allow; source: node properties.enginePolicy.bots. Search crawling for Siri, Spotlight and Safari. - Applebot-Extended (training): allow; source: wildcard (no explicit decision on the node yet). A separate token that controls use of crawled content for Apple foundation models; no separate crawler. - Verification of real requests: reverse DNS with forward confirmation — *.applebot.apple.com - What Apple asks: Decide search and training separately: Applebot-Extended is the training choice. - Sources: D-apple https://support.apple.com/en-us/119829 ## DuckDuckGo — DuckDuckGo search and DuckAssist - DuckDuckBot (search): allow; source: node properties.enginePolicy.bots. Search crawling. - DuckAssistBot (assistant): allow; source: wildcard (no explicit decision on the node yet). Fetches pages for cited AI-assisted answers; the provider says this data is not used to train AI models. - Verification of real requests: reverse DNS with forward confirmation, and a published IP list — *.duckduckgo.com - What DuckDuckGo asks: Allow the crawler; DuckAssist cites crawled pages. - Sources: D-ddg https://duckduckgo.com/duckduckgo-help-pages/results/duckduckbot/ · D-duckassist https://duckduckgo.com/duckduckgo-help-pages/results/duckassistbot ## Yandex — Yandex Search - YandexBot (search): allow; source: wildcard (no explicit decision on the node yet). Search crawling; receives IndexNow notifications. - Verification of real requests: reverse DNS with forward confirmation — *.yandex.ru / *.yandex.net / *.yandex.com - What Yandex asks: IndexNow-notified URLs are crawled promptly. - Where it reports: Yandex Webmaster (not connected). - Sources: S30 https://www.indexnow.org/documentation ## Meta — Meta AI and link previews - Meta-WebIndexer (search): allow; source: wildcard (no explicit decision on the node yet). Indexes content for Meta AI search quality and source citations. - Meta-ExternalFetcher (user-fetch): allow; source: wildcard (no explicit decision on the node yet). Retrieves links requested by users and supports agentic capability evaluation; user-requested fetches may bypass robots.txt. Robots rules are only partly honoured for this token by the provider's own account. - Meta-ExternalAgent (training): allow; source: wildcard (no explicit decision on the node yet). Crawls for foundation-model training and direct content indexing for products. - facebookexternalhit (preview): allow; source: wildcard (no explicit decision on the node yet). Link previews when a URL is shared. Robots rules are only partly honoured for this token by the provider's own account. - Verification of real requests: provider verification guidance — Use Meta’s official crawler identification guidance; a claimed user agent alone is not verified identity. - What Meta asks: Treat search indexing, user requests, training and link previews as distinct purposes. - Sources: D-meta https://developers.facebook.com/documentation/sharing/webmasters/web-crawlers ## Amazon — Alexa and Amazon services - Amzn-SearchBot (search): allow; source: wildcard (no explicit decision on the node yet). Crawls for Amazon search experiences without generative-model training. When its own rule is absent, it may follow other search-bot rules. - Amzn-User (user-fetch): allow; source: wildcard (no explicit decision on the node yet). Fetches information for user actions, without generative-model training; may not follow every robots directive. Robots rules are only partly honoured for this token by the provider's own account. - Amazonbot (training): allow; source: wildcard (no explicit decision on the node yet). Improves Amazon products and services; collected content may train Amazon AI models. - Verification of real requests: published IP ranges — Amazon publishes separate address lists for these three user agents. - What Amazon asks: Choose access separately for each user agent. - Sources: D-amazon https://developer.amazon.com/amazonbot ## Common Crawl — The open web corpus many models train on - CCBot (training): allow; source: wildcard (no explicit decision on the node yet). Builds the Common Crawl corpus used by many model trainers. - Verification of real requests: reverse DNS — Common Crawl publishes its crawler details. - What Common Crawl asks: A training decision that reaches many downstream models at once. - Sources: D-ccbot https://commoncrawl.org/ccbot ## Mistral — Mistral search, user-requested answers and model training - MistralAI-Index (search): allow; source: wildcard (no explicit decision on the node yet). Indexes content for Mistral search; the provider excludes generative model training from this crawler. - MistralAI-User (user-fetch): allow; source: wildcard (no explicit decision on the node yet). Retrieves pages for user requests; not automated crawling or generative model training. Its robots token controls eligible sites. - MistralAI-Training (training): allow; source: wildcard (no explicit decision on the node yet). Collects content for model-training datasets; separate from search indexing and live user queries. - Verification of real requests: provider documentation and published IP lists where supplied — The official reference links address lists for MistralAI-User and MistralAI-Index; do not infer request identity from a user agent alone. - What Mistral asks: Decide search, user fetch and training independently using the documented tokens. - Sources: D-mistral https://docs.mistral.ai/robots ## Brave — Brave Search - Verification of real requests: no differentiated user agent documented — The official crawler page does not supply a separate token that could identify Brave requests. - What Brave asks: Brave says it will not crawl pages that Googlebot cannot crawl; use the existing Googlebot policy, not an invented token. Robots controls crawling; index removal requires the documented noindex workflow. - Sources: D-brave https://search.brave.com/help/brave-search-crawler ## Consoles - Google Search Console: property sc-domain:feedthejoe.com; property recorded on the node; this manifest does not verify current ownership or access. Provides: Sitemap registered on the node: https://feedthejoe.com/sitemap-index.xml; Sitemap submission recorded: 2026-09-13. - Bing Webmaster Tools + IndexNow: property https://feedthejoe.com/; Verification status recorded: yes; verification file recorded: /BingSiteAuth.xml. Provides: Verification file registered on the node: /BingSiteAuth.xml; Sitemap feed registered on the node: https://feedthejoe.com/sitemap-index.xml; Sitemap feed submission recorded: 2026-09-13; IndexNow key files recorded: 3; IndexNow notification run recorded: 2026-09-13: 241 urls, HTTP 200. ## Non-public and non-content paths - /api/ - /admin/ - /private/ - /*?*utm_ - /wp-admin/ ## Canonical discovery - Sitemap: https://feedthejoe.com/sitemap-index.xml - Summary: https://feedthejoe.com/llms.txt (composed from the section indexes) - The engine tree as data: https://feedthejoe.com/engines.json - https://feedthejoe.com/entity.json - https://feedthejoe.com/entity-graph.json - https://feedthejoe.com/aeo.json - https://feedthejoe.com/site-dataset.json - https://feedthejoe.com/feed.xml - https://feedthejoe.com/library/feed.xml - https://feedthejoe.com/sitemap-index.xml