Glossary · Identity and trust

Bot detection

Techniques that decide whether a web request comes from a person or automated software, by scoring traffic signals or checking bots that identify themselves.

Bot detection is the set of techniques a website or API uses to decide whether a request comes from a person or from automated software, and if automated, which software.

Two approaches.

Approach How it works Limits
Inference Score each request from signals such as headers, browser characteristics, behaviour and JavaScript checks Probabilistic; capable bots imitate real browsers
Declared identity The bot proves who it is: DNS checks on its IP address, published IP ranges, or cryptographic request signatures Works only for bots that choose to take part

Cloudflare’s bot score is an example of inference: a value from 1 to 99 produced by heuristics, machine learning and optional JavaScript detections, where 1 means almost certainly automated. Google’s crawler verification is an example of declared identity: the site runs a reverse DNS lookup on the requesting IP address, checks that the name belongs to one of the Google domains the documentation lists, then runs a forward lookup to confirm the name maps back to the same address.

Signed bots. The IETF Web Bot Auth draft points out that User-Agent strings can be spoofed and that IP lists are hard to maintain and attribute. It has automated clients sign requests with HTTP Message Signatures (RFC 9421), using keys the operator publishes in a JWKS-based directory, located through a Signature-Agent header or a well-known URI. The working group’s charter covers AI agents that fetch content for end users. It excludes agent-to-agent interfaces, authentication of the end user, and techniques for telling non-participating bots from people.

robots.txt is a request. The Robots Exclusion Protocol (RFC 9309) says its rules are not a form of access authorization. Crawlers that follow the protocol respect them; nothing forces others to.

For agent-facing services. An A2A endpoint expects automated clients, so the useful question shifts from “is this a bot?” to “which agent is this, who runs it, and should I trust this request?” That is the territory of agent identity and Know Your Agent checks.

Neighbouring terms. Reputation scoring rates identified agents over time. Rate limiting is often the first action a detection result triggers.

Sources

  1. Cloudflare documentation: Bot scores (accessed )
  2. Google Crawling Infrastructure: Verifying requests from Google's crawlers and fetchers (accessed )
  3. IETF draft: HTTP Message Signatures for automated traffic (draft-ietf-webbotauth-httpsig-protocol) (accessed )
  4. IETF Web Bot Auth (webbotauth) working group charter (accessed )
  5. RFC 9421: HTTP Message Signatures (accessed )
  6. RFC 9309: Robots Exclusion Protocol (accessed )