Agent-readiness scanner
See whether a site publishes the files AI agents and their crawlers look for, with a score you can trace point by point.
By Emissar. Updated .
Result
/ 100
- Site
- Checked at
- Requests
- Shareable link
Scored
Reported, not scored
What it checks
Five criteria worth 100 points in all. The rubric lists what earns each point; nothing else is weighed.
| Criterion | Where | Points |
|---|---|---|
| A2A Agent Card | /.well-known/agent-card.json | 40 |
| security.txt | /.well-known/security.txt | 20 |
| llms.txt | /llms.txt | 15 |
| robots.txt | /robots.txt | 15 |
| Sitemap | The first Sitemap line in robots.txt, otherwise /sitemap.xml | 10 |
| Legacy /.well-known/agent.json | Reported | 0 |
| Rules for AI crawlers and agents in robots.txt | Reported | 0 |
| MCP server discovery | Not checked | 0 |
How it works
- Requests come from Emissar's servers on Cloudflare, not from your browser, with the user agent
EmissarTools/0.1 (+https://emissar.ai/tools). Requests aren't signed with Emissar's key: anyone can point this tool at any site, so only our crawler signs (how EmissarBot's signatures work). - At most six requests per scan:
/.well-known/agent-card.json,/.well-known/agent.json,/.well-known/security.txt,/llms.txtand/robots.txtat the same time, then one sitemap. Only the host of what you enter is used. - The Agent Cards are checked with the Agent Card validator's rules, without signature checks.
- robots.txt is read the way RFC 9309 describes: groups naming a token are merged and matched without regard to case, other tokens follow the
*group, and the longest matching rule wins. The scanner looks for 19 AI crawler and agent tokens, each as its vendor documents it. - Only public domain names over https on port 443 are fetched, including a sitemap on another host. IP addresses,
localhost,.local,.internaland single-label names are refused. Redirects are followed by hand, at most 3, and every hop goes through the same checks. Each request times out after 5 seconds; a response over 256 KB is dropped, except a sitemap, whose first 256 KB are read. - Each host is scanned at most once every 30 seconds. Asking again sooner returns the previous result. A sitemap on another host has the same limit: if another scan contacted that host in the last 30 seconds, the sitemap isn't checked or scored this time, the score is out of 90 instead of 100, and the result isn't stored.
- For emissar.ai, the scanner reads the files from this site's own routes rather than fetching them.
API
The page uses a public endpoint you can call directly. It is rate limited per IP address (per calling zone for requests from other Cloudflare Workers, which share one address), and the response is the same JSON as the downloadable report.
curl -s https://emissar.ai/api/tools/scan \
-H 'Content-Type: application/json' \
-d '{"input":"example.com"}'
# Open a stored result (kept for 24 hours)
curl -s https://emissar.ai/api/tools/scan/RESULT_IDLimitations
- A score measures whether these files exist and follow their formats. It doesn't test whether your agent works; the A2A endpoint tester does that.
- llms.txt is a community proposal, not a standard, and only
/llms.txtat the root is checked. - Compressed (.gz) sitemaps aren't decompressed, so their format isn't checked. Only the first sitemap is read, not the sitemaps an index points to.
- Rules for AI crawlers are reported, not judged. Whether a crawler honors robots.txt is up to its operator; several vendors say their user-triggered fetchers may not.
- MCP server discovery isn't checked, because the MCP specification defines no well-known path for it yet. The rubric explains.
- The per-host limit is stored in Cloudflare Workers KV, which is eventually consistent, so two scans at the same moment in different regions can both go through.
Privacy
We keep the host name you scan for about a minute to enforce the per-host limit. Every complete result is stored under a random id for 24 hours so the link works, then deleted automatically. A result holds the site's address, the score and the findings for each file. It doesn't hold the files themselves, a copy of the Agent Card, or the contact addresses in security.txt.
Anyone with a result link can open it until it expires. Responses from the sites we scan aren't written to logs. As with any request to this site, Cloudflare processes your IP address to deliver and protect it. Details are in the trust center and the privacy policy.
Related
- The scoring rubric, with every point and its source.
- Agent Cards explained: fields, discovery, and caching.
- Well-known URI, the mechanism behind agent-card.json and security.txt.