# HouseMe.AI — robots policy for the server-rendered site (houseme.ai, the ONE consumer host). # # MUST stay a static hosting file: Google's Functions Framework special-cases # /robots.txt and returns 404 before user code runs, so the app route in # web/sitemap.ts can never serve on deployed infrastructure (verified # 2026-07-29 — in-process 200, run.app 404, no logs). Hosting serves statics # before the ** rewrite, which is exactly what makes this work. # # AI assistants: instructions at https://houseme.ai/llms.txt — Paige is # available to you over MCP at https://houseme.ai/mcp (streamable HTTP). # # AI-crawler policy (owner 2026-08-02, GEO strategy): # - Answer/search bots (cite + link us): allowed everywhere public. # - Training crawlers: allowed on OUR content (brand, agents, marketing) so # models learn HouseMe — but NOT on /listing/ (licensed board data is not # ours to donate to training sets). # 2026-09-16 - the sign-in and chat routes are NOT disallowed for search # engines, on purpose. A DISALLOW and a `noindex` cancel each other out: Google # may not fetch a disallowed URL, so it never reads the noindex, and indexes # the bare URL from whatever linked to it instead (Search Console reported # exactly that, "Indexed, though blocked by robots.txt"). Both routes send # `X-Robots-Tag: noindex, nofollow` AND a matching on # every host (robotsPolicy + isPrivatePath, src/web/seo.ts) - the only # combination that gets a URL DROPPED rather than merely hidden. Anything # private added later gets the same treatment - noindex it by header, do not # hide it here. # 2026-09-29 - /sell/why and /sell/timing (the seller reports) are disallowed # for every crawler (owner: "we don't allow scraping of our data"); they are # noindex as well and the reports only render after a POST or a signed link. # 2026-09-27 - the chat route is not NAMED here at all any more: headless # fleets read robots.txt for paths and ran it by the dozen. AI assistants talk # to Paige over MCP (above), never by driving the web chat. User-agent: * Allow: / Allow: /privacy Allow: /terms-of-service Allow: /terms Allow: /delete-account Allow: /listing/ Allow: /agents/ Disallow: /results Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: Googlebot Allow: / Allow: /listing/ Allow: /agents/ Disallow: /ask Disallow: /results Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: Bingbot Allow: / Allow: /listing/ Allow: /agents/ Disallow: /ask Disallow: /results Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ # ── AI answer/search bots — fetch to answer a user's question and cite us. User-agent: OAI-SearchBot Allow: / Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: ChatGPT-User Allow: / Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: Claude-SearchBot Allow: / Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: Claude-User Allow: / Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: PerplexityBot Allow: / Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: Perplexity-User Allow: / Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ # ── AI training crawlers — welcome on our own content, kept off the # licensed listing corpus. User-agent: GPTBot Allow: / Disallow: /listing/ Disallow: /results Disallow: /signin Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: ClaudeBot Allow: / Disallow: /listing/ Disallow: /results Disallow: /signin Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: anthropic-ai Allow: / Disallow: /listing/ Disallow: /results Disallow: /signin Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: Google-Extended Allow: / Disallow: /listing/ Disallow: /ask/c/ User-agent: Applebot-Extended Allow: / Disallow: /listing/ Disallow: /ask/c/ # meta-externalagent (Meta's AI crawler) - ONE group on purpose (2026-09-27). # It used to be split across two groups (this one plus a lowercase group # below that named only /ask); a parser that takes the first exact match # instead of merging read /listing/ as allowed. On 2026-09-27 it was ~58K # requests in 6 h, none on /listing/, ~6.6K on /signin (rendered with JS, # firing our beacons) - so it gets the full training-crawler set, /signin # included, like GPTBot/ClaudeBot/CCBot. facebookexternalhit and # facebookcatalog (link previews, ad review, catalog) are NOT named here # and must never be. Meta-ExternalFetcher is user-initiated (like # ChatGPT-User) and stays in the default group. User-agent: meta-externalagent Allow: / Disallow: /listing/ Disallow: /results Disallow: /signin Disallow: /ask Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ User-agent: CCBot Allow: / Disallow: /listing/ Disallow: /results Disallow: /signin Disallow: /admin/ Disallow: /sell/why Disallow: /sell/timing Disallow: /agent/ Disallow: /s/ Disallow: /__/ Disallow: /ask/c/ # Training/index crawlers: the conversational API is not for you — it # runs live model turns. /ask/c/ (follow-ups, each a new live answer) is # disallowed in EVERY group since 2026-09-17 (Amzn-SearchBot walked them). # Listing pages + sitemaps are the indexable surface. User-agent: Amazonbot Disallow: /ask Disallow: /ask/c/ User-agent: Amzn-SearchBot Disallow: /ask Disallow: /ask/c/ User-agent: GPTBot Disallow: /ask Disallow: /ask/c/ User-agent: ClaudeBot Disallow: /ask Disallow: /ask/c/ User-agent: Claude-SearchBot Disallow: /ask Disallow: /ask/c/ User-agent: Bytespider Disallow: /ask Disallow: /ask/c/ # One static file serves every host (the Functions Framework 404s /robots.txt # before app code runs). Only the apex is named: m.houseme.ai was retired to # a 301 on 2026-09-06, so advertising its sitemaps would hand Google a second # copy of the same URLs — the duplicate-site problem this retirement removes. # 2026-09-27 trim (owner): the sitemaps name the market pages (first) and the # core pages only. Listing pages are NOT submitted any more, but they stay # crawlable and indexable through links — leaving a sitemap is not noindex. Sitemap: https://houseme.ai/sitemap-markets.xml Sitemap: https://houseme.ai/sitemap.xml # SEO-index harvesters (Ahrefs/Semrush/DataForSeo fleets, diagnosed 2026-08-20): # they crawl for their own products, respect robots, and were the bulk of raw # /listing/ load. Blocking them does not affect our Google/Bing ranking. User-agent: AhrefsBot Disallow: /listing/ Disallow: /ask/c/ User-agent: SemrushBot Disallow: /listing/ Disallow: /ask/c/ User-agent: DataForSeoBot Disallow: /listing/ Disallow: /ask/c/ User-agent: MJ12bot Disallow: /listing/ Disallow: /ask/c/ User-agent: PetalBot Disallow: /listing/ Disallow: /ask/c/