Who respects your robots.txt? AI crawler compliance, measured in the open.
This site runs honeypot paths explicitly blocked in /robots.txt.
Every crawler that visits them is logged as a violation.
No guessing — just evidence.
Bot Compliance Scoreboard
| Bot | Operator | Verdict | Signed | Violations | Visits |
|---|---|---|---|---|---|
| Loading… | |||||
/robots.txt. A
crawler that follows a link and visits the target has unambiguously
ignored the rule — a "via link" violation.
The other three honeypot paths (
/private/, /honeypot/,
/robots-test/) have no links anywhere on the site.
A hit there means the bot used the robots.txt
Disallow list as a crawl seed ("treasure map") or guessed the
path — a "via guess" violation.
Live Violation Feed
Discovery Reads — who fetches llms.txt, agents.md, grounding pages & the API/AI catalogs
These files are invitations, not traps:
/llms.txt is the LLM-oriented site summary, and
/agents.md (served at /AGENTS.md,
/agents.md and /.well-known/agents.md) is a
probe surface for agents that look for usage instructions. The
/.well-known/api-catalog goes one step further: it offers a
machine-readable description of the JSON stats API, so an agent can use the
interface instead of scraping this page. The
/.well-known/ai-catalog.json Agentic Resource Discovery
manifest goes further still — it lets agent-facing registries index this
site so agents can find it in the first place. Reading any of them is
not a violation — it is the opposite signal: an agent
deliberately doing discovery. So far —
discovery reads have been logged.
| Bot | Operator | llms.txt | agents.md | grounding | api catalog | ai catalog | Total | Last read |
|---|---|---|---|---|---|---|---|---|
| Loading… | ||||||||