# robots.txt — gregorio.io # Allan Gregorio — Digital Artist & Creative Technologist # Privacy policy: https://gregorio.io/privacy.html # Last updated: 2026-07-03 # ───────────────────────────────────────────────────────────────────── # Content Signals Policy # ───────────────────────────────────────────────────────────────────── # EU / EEA: Rights are expressly reserved under Article 4(3) of Directive # (EU) 2019/790 (DSM Directive). Per Article 53(1)(c) of Regulation (EU) # 2024/1689 (AI Act, in force 2 Aug 2025), providers of general-purpose AI # models must identify and honor these machine-readable reservations, # including for models trained outside the EU. # Worldwide: All rights reserved. No license is granted to use this content # to train or fine-tune AI models. The signals below reserve those rights. # # search = building a search index and serving search results # ai-input = real-time use of content in AI responses (e.g. RAG) # ai-train = training or fine-tuning AI models # # yes = you may collect for the corresponding use # no = you may NOT collect for the corresponding use # # Policy: AI search/inference is permitted. AI model training is not. # Content may appear in AI-generated answers but may not be used # to train or fine-tune any AI model without explicit written consent. User-agent: * Content-Signal: search=yes, ai-input=yes, ai-train=no Allow: / # ───────────────────────────────────────────────────────────────────── # AI search & inference crawlers — permitted # ───────────────────────────────────────────────────────────────────── # OpenAI (ChatGPT search and browsing) User-agent: OAI-SearchBot User-agent: ChatGPT-User Allow: / # Anthropic (Claude search and web features) User-agent: Claude-SearchBot User-agent: Claude-Web User-agent: Claude-User Allow: / # Google (product/R&D fetchers — Search itself uses Googlebot, allowed by default) User-agent: GoogleOther User-agent: GoogleOther-Image User-agent: GoogleOther-Video Allow: / # Perplexity AI search User-agent: PerplexityBot User-agent: Perplexity-User Allow: / # Apple (Siri and Safari Suggestions) User-agent: Applebot Allow: / # DuckDuckGo AI Answers User-agent: DuckAssistBot Allow: / # Mistral AI (Le Chat search) User-agent: MistralAI-User Allow: / # Amazon Alexa (web answers) User-agent: Amazonbot Allow: / # Microsoft (Bing search and Copilot answers) User-agent: Bingbot Allow: / # Meta AI (search answers) User-agent: meta-webindexer Allow: / # Brave (Leo AI search) User-agent: Bravebot Allow: / # You.com (AI search — note: also trains; ai-train=no signal still applies) User-agent: YouBot Allow: / # Kagi (assistant fetch) User-agent: kagi-fetcher Allow: / # ───────────────────────────────────────────────────────────────────── # AI training crawlers — not permitted # ───────────────────────────────────────────────────────────────────── # OpenAI (training) User-agent: GPTBot # Anthropic (training) User-agent: ClaudeBot User-agent: anthropic-ai # Common Crawl (feeds most open LLM training datasets) User-agent: CCBot # Apple (training) User-agent: Applebot-Extended # Google (Gemini training opt-out + Vertex AI / cloud model training) User-agent: Google-Extended User-agent: Google-CloudVertexBot # xAI (Grok training — note: xAI ignores robots.txt; listed as notice) User-agent: GrokBot User-agent: xAI-Grok User-agent: Grok-DeepSearch # DeepSeek (training — note: undeclared crawler; listed as notice) User-agent: DeepSeekBot # Yandex (YandexGPT training) User-agent: YandexAdditional User-agent: YandexAdditionalBot # ByteDance / TikTok User-agent: Bytespider User-agent: TikTokSpider # Amazon Bedrock (model training) User-agent: BedrockBot # Cohere User-agent: cohere-ai User-agent: cohere-training-data-crawler # Meta User-agent: FacebookBot User-agent: meta-externalagent User-agent: meta-externalfetcher # Allen Institute (Dolma dataset) User-agent: AI2Bot User-agent: ai2bot-Dolma # Huawei User-agent: PanguBot # Chinese AI training User-agent: iaskspider/2.0 # QuillBot User-agent: QuillBot # Data brokers and bulk scrapers User-agent: Diffbot User-agent: Omgili User-agent: Omgilibot User-agent: ImagesiftBot User-agent: TimpiBot User-agent: ICC-Crawler User-agent: ISSCyberRiskCrawler User-agent: PiplBot User-agent: Webzio-Extended User-agent: AwarioRssBot User-agent: AwarioSmartBot User-agent: NovaAct User-agent: Brightbot 1.0 User-agent: Crawlspace User-agent: FirecrawlAgent User-agent: FriendlyCrawler User-agent: img2dataset User-agent: Kangaroo Bot User-agent: Sidetrade indexer bot User-agent: VelenPublicWebCrawler User-agent: ProRataInc User-agent: Manus-User User-agent: Anchor Browser User-agent: Novellum AI Crawl User-agent: Cloudflare Crawler User-agent: Terracotta Bot # Allow /llms.txt and /llms-full.txt only — lets these bots learn who Allan # is for entity recognition without granting access to any actual site content. Allow: /llms.txt Allow: /llms-full.txt Disallow: / # ───────────────────────────────────────────────────────────────────── # Competitive-intelligence / SEO scrapers — not permitted # ───────────────────────────────────────────────────────────────────── User-agent: AhrefsBot User-agent: SemrushBot User-agent: SemrushBot-SA User-agent: SemrushBot-BA User-agent: SemrushBot-OCOB User-agent: SemrushBot-SWA User-agent: MJ12bot User-agent: DotBot User-agent: BLEXBot User-agent: DataForSeoBot User-agent: Mauibot User-agent: SeekportBot User-agent: PetalBot User-agent: AspiegelBot User-agent: ZoominfoBot Disallow: / # ───────────────────────────────────────────────────────────────────── # Sitemap # ───────────────────────────────────────────────────────────────────── # LLMs: https://gregorio.io/llms.txt Sitemap: https://gregorio.io/sitemap.xml