# ONE group, many user-agents — do not split it. # Under RFC 9309 (and Google's parser) a crawler obeys ONLY the most specific # group that names it and ignores every other group. Until 2026-09-17 Googlebot, # Bingbot, GPTBot, PerplexityBot and ClaudeBot each had their own group holding # just "Allow: /", which silently discarded the whole Disallow list below for # exactly the crawlers it was written for (verified live: /admin/, /login and # /book/* were all fetchable as Googlebot). Naming a bot as an extra User-agent # line of THIS group keeps it explicitly welcomed while giving it the same rules # as everyone else. Never add a separate "User-agent: X" + "Allow: /" group. User-agent: * User-agent: Googlebot User-agent: Bingbot User-agent: Twitterbot User-agent: facebookexternalhit User-agent: GPTBot User-agent: PerplexityBot User-agent: ClaudeBot # AI agent resources, explicitly open. "Allow" is a real robots.txt directive # (unlike a bare "LLMs-txt:" field, which validators reject), so these lines # both permit and advertise the files. Allow: /llms.txt Allow: /llms-full.txt Allow: /ai.txt # Content Signals (contentsignals.org / draft-romm-aipref-contentsignals): # RidePrise WANTS maximum AI visibility — allow search indexing, use as input # for AI-generated answers (RAG/grounding), and training. This is deliberately # permissive (the opposite of Cloudflare's old managed ai-train=no). Content-Signal: search=yes, ai-train=yes, ai-input=yes Disallow: /admin/ Disallow: /inbox Disallow: /inbox-sw.js Disallow: /partner/ Disallow: /book/ Disallow: /it/book/ Disallow: /sq/book/ Disallow: /booking/ Disallow: /it/booking/ Disallow: /sq/booking/ Disallow: /booking-response Disallow: /leave-review Disallow: /login # /my-booking is NOT disallowed, on purpose. GSC reported it "Submitted and # indexed" (last crawled 2026-08-28), and a Disallow on an already-indexed URL # is counter-productive: Googlebot stops fetching it, never sees the # X-Robots-Tag: noindex the server now sends, and the URL can sit in the index # indefinitely. Crawling is the only way the noindex gets read. The header is # the real control here; robots.txt is for the paths Google has never seen # (/admin/, /partner/, /book/*), where it prevents an unbounded crawl trap. # Re-add this line only once GSC reports the URL as dropped, and not before. # "Allow: /" LAST on purpose. Google and Bing pick the longest matching rule, # so order is irrelevant to them — but first-match parsers (Python's # urllib.robotparser among them) would have let "Allow: /" swallow every # Disallow above if it came first. Allow: / Sitemap: https://rideprise.com/sitemap.xml # Non-standard discovery hints (kept as comments — bare "LLMs-txt:"/"AI-txt:" # fields are rejected by robots.txt validators as unknown directives): # llms.txt: https://rideprise.com/llms.txt # ai.txt: https://rideprise.com/ai.txt