Learn

Detecting AI agents on your website

How to recognize AI crawlers and agents on your site, from user agents and IP ranges to robots.txt and signed requests (RFC 9421, Web Bot Auth).

You detect AI agents on a website by combining signals of increasing strength. A user-agent string is a claim that anyone can copy. Published IP ranges and DNS checks confirm that claim for operators that publish them. robots.txt records what you want and what well-behaved operators say they honor, but it enforces nothing. A request signed with HTTP Message Signatures, using the Web Bot Auth profile, proves which operator’s key sent it.

None of these signals tells you what the agent is trying to do or for whom. If agents come to your site to complete tasks, you can also give them a direct route: an agent endpoint that accepts structured requests. This guide covers each signal from weakest to strongest, with a tested script for IP checks, and ends with practical steps.

Three kinds of automated traffic

Operators now document several distinct bots, and they behave differently:

Kind What it does Examples from vendor docs robots.txt
Training crawler Collects pages that may be used to train models GPTBot (OpenAI), ClaudeBot (Anthropic) Vendors document robots.txt controls
Search crawler Indexes pages for an AI search or answer product OAI-SearchBot, Claude-SearchBot, PerplexityBot, Googlebot Vendors document robots.txt controls
User-triggered fetcher or agent Fetches a page, or acts on a site, because a person asked ChatGPT-User, Claude-User, Perplexity-User, Google-Agent Varies; several vendors say these may not follow robots.txt

Google adds a fourth case. Google-Extended is a robots.txt product token with no user-agent string of its own. It controls whether content Google crawls may be used for Gemini model training and grounding, and Google says it does not affect inclusion or ranking in Google Search.

The third row matters most to businesses, because those requests often come from a customer trying to get something done. Treat them differently from bulk crawling.

Signal 1: user-agent strings

Most operators publish the tokens their bots send. Google’s user-triggered fetchers page, for example, lists Google-Agent, used by agents hosted on Google infrastructure to navigate sites and act on a user’s request. Its user agent is a normal Chrome string with a compatible; Google-Agent token added. OpenAI lists GPTBot, OAI-SearchBot, ChatGPT-User and OAI-AdsBot. Anthropic lists ClaudeBot, Claude-User and Claude-SearchBot. Perplexity lists PerplexityBot and Perplexity-User.

Use these tokens to sort traffic that identifies itself honestly: logs, analytics, rate-limit buckets. Never use them alone to grant access. The string is set by the client, so an abusive scraper can send GPTBot as easily as OpenAI can. And a client that sends an ordinary browser user agent looks like any other visitor.

Signal 2: published IP ranges and DNS

Operators that want to be verifiable publish where their traffic comes from.

Operator What it publishes
OpenAI One JSON file per bot: gptbot.json, searchbot.json, chatgpt-user.json, adsbot.json under openai.com
Anthropic One file, claude.com/crawling/bots.json, covering its bots
Google JSON files per category (common crawlers, special-case crawlers, user-triggered fetchers, user-triggered agents), plus reverse DNS
Perplexity perplexitybot.json and perplexity-user.json under perplexity.com

All of these files share one shape: a creationTime and a list of prefixes, each with an ipv4Prefix or ipv6Prefix.

Google also documents a DNS check. Run a reverse DNS lookup on the source IP, confirm the name ends in googlebot.com, google.com or googleusercontent.com, then run a forward lookup on that name and confirm it returns the same IP.

A tested check

This script checks an IP against a published range file, or runs Google’s DNS check. It uses only the Python standard library.

import ipaddress, json, socket, sys, urllib.request

# IP range files published by each operator (all use the same JSON shape).
RANGES = {
    "openai-chatgpt-user": "https://openai.com/chatgpt-user.json",
    "openai-gptbot": "https://openai.com/gptbot.json",
    "anthropic": "https://claude.com/crawling/bots.json",
    "google-agent": "https://developers.google.com/static/crawling/ipranges/user-triggered-agents.json",
    "perplexity-user": "https://www.perplexity.com/perplexity-user.json",
}

def in_published_range(ip: str, url: str) -> bool:
    req = urllib.request.Request(url, headers={"User-Agent": "ip-range-check/1.0"})
    with urllib.request.urlopen(req, timeout=10) as resp:
        prefixes = json.load(resp)["prefixes"]
    addr = ipaddress.ip_address(ip)
    return any(addr in ipaddress.ip_network(p.get("ipv4Prefix") or p.get("ipv6Prefix")) for p in prefixes)

def google_dns_check(ip: str) -> bool:
    """Reverse DNS, check the domain, then forward DNS back to the same IP."""
    try:
        host = socket.gethostbyaddr(ip)[0].rstrip(".")
    except OSError:
        return False
    if not host.endswith((".googlebot.com", ".google.com", ".googleusercontent.com")):
        return False
    return ip in socket.gethostbyname_ex(host)[2]

if __name__ == "__main__":
    operator, ip = sys.argv[1], sys.argv[2]
    if operator == "google-dns":
        print(ip, "passes Google's DNS check:", google_dns_check(ip))
    else:
        print(ip, "in", operator, "ranges:", in_published_range(ip, RANGES[operator]))

Output from a run on September 26, 2026. The first address sits inside a prefix in OpenAI’s file that day; 203.0.113.7 is a documentation address that belongs to no one; 35.247.243.240 is the example IP from Google’s verification page.

$ python check_bot_ip.py openai-chatgpt-user 104.208.184.193
104.208.184.193 in openai-chatgpt-user ranges: True
$ python check_bot_ip.py openai-chatgpt-user 203.0.113.7
203.0.113.7 in openai-chatgpt-user ranges: False
$ python check_bot_ip.py anthropic 216.73.216.10
216.73.216.10 in anthropic ranges: True
$ python check_bot_ip.py google-agent 2001:4860:c::5
2001:4860:c::5 in google-agent ranges: True
$ python check_bot_ip.py google-dns 35.247.243.240
35.247.243.240 passes Google's DNS check: True
$ python check_bot_ip.py google-dns 203.0.113.7
203.0.113.7 passes Google's DNS check: False

One practical detail from testing: the first run against Anthropic’s file failed with HTTP 403 until the script sent its own User-Agent header. Range files can sit behind bot protection too. In production, fetch the files on a schedule, cache them, and match IPs locally instead of fetching on every request.

Limits of IP checks

  • IP ranges identify an operator’s infrastructure. They don’t say which product, which user, or which task sent the request.
  • The lists change. The files carry a creationTime for that reason.
  • AWS points out, in its Security Blog post on Web Bot Auth, that IP filtering and reverse DNS break down on multi-tenant platforms, where many unrelated workloads share the same address space.
  • Anthropic notes that blocking its IP addresses may not work as a reliable opt-out, because it stops its crawler from reading your robots.txt. Use robots.txt for preferences.

An individual Internet-Draft in the IETF webbotauth group, draft-illyes-webbotauth-jafar, proposes a common JSON format for publishing IP ranges of automated clients. It is not adopted as a standard.

Signal 3: robots.txt conventions

robots.txt is standardized as RFC 9309 (September 2022, Standards Track). It lets a site state which paths each crawler may fetch, by user-agent token. Section 1 is explicit about what it is not: “These rules are not a form of access authorization.” Compliance is voluntary.

What robots.txt gives you is a published expectation. Each vendor documents what its bots do with it:

  • Google’s common crawlers always obey robots.txt when crawling automatically. Google’s user-triggered fetchers generally ignore it, because a user requested the fetch.
  • OpenAI says robots.txt rules may not apply to ChatGPT-User, since its actions are initiated by a user. GPTBot and OAI-SearchBot are separate controls.
  • Perplexity says Perplexity-User generally ignores robots.txt for the same reason.
  • Anthropic says its bots honor robots.txt directives, and describes disallowing Claude-User as the way to stop retrieval for user queries.

Illustrative robots.txt that allows search crawlers, opts out of training, and keeps a private area closed to everyone:

User-agent: GPTBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: *
Disallow: /account/

A request that ignores these rules while claiming a token whose vendor says it obeys them is a useful signal in itself. Verify the claim with signal 2 before you conclude which is happening: a misbehaving bot, or someone impersonating it.

Signal 4: signed requests with HTTP Message Signatures

The strongest signal is cryptographic. RFC 9421 (February 2024, Standards Track) defines how a client signs parts of an HTTP message. The Signature-Input field lists the covered components, such as @authority or @path, plus parameters including created, expires, keyid and tag. The Signature field carries the signature value.

Web Bot Auth is a profile of RFC 9421 for automated clients. The IETF chartered the webbotauth working group for it. As of September 26, 2026, its main document is the working group draft draft-ietf-webbotauth-httpsig-protocol-00 (September 1, 2026), which grew out of the earlier individual drafts draft-meunier-web-bot-auth-architecture and draft-meunier-http-message-signatures-directory and now defines both the signing profile and the key directory. It is an Internet-Draft, so details can change.

How it works

  1. The operator generates a signing key and publishes the public key as a JWK set at /.well-known/http-message-signatures-directory on its domain, over HTTPS.
  2. Each request carries three headers: Signature-Input, Signature, and Signature-Agent, which tells the site where to find the key directory.
  3. Signature-Input must cover @authority or @target-uri, include created and expires, set keyid to the key’s JWK thumbprint, and set tag="web-bot-auth".
  4. The site checks the tag, finds the key by keyid (fetching and caching the directory if the key is new), verifies the signature, and checks the time window.

The example in the working group draft, with line wrapping as the draft shows it:

GET /path/to/resource HTTP/1.1
Host: origin.example.com
Signature: sig=abc==
Signature-Input: sig=("@authority" "signature-agent";key="sig");\
                 created=1700000000;\
                 expires=1700011111;\
                 keyid="ba3e64==";\
                 tag="web-bot-auth"
Signature-Agent: sig="https://signer.example.com"

Two cautions from the draft itself. First, a signature that covers only @authority binds to your host, not to one request, so anyone who observes it can replay it against your site until it expires. The draft recommends expiry of no more than 24 hours; shorter is safer. Second, formats have moved between revisions. The working group draft writes Signature-Agent as a dictionary member (sig="..."), while Cloudflare’s documentation, which cites the earlier drafts, uses a plain quoted string. Check which revision your verifier supports.

Who uses it

  • OpenAI documents that ChatGPT Work’s Cloud browser signs outbound requests with Web Bot Auth, sends Signature-Agent: "https://chatgpt.com", and publishes keys at https://chatgpt.com/.well-known/http-message-signatures-directory.
  • Cloudflare lets bot operators register for verification with request signatures, and treats Web Bot Auth as one of the ways a bot can become a verified bot.
  • AWS WAF Bot Control (rule group version 4.0 and later) validates signatures at the edge and adds labels such as web_bot_auth:verified, web_bot_auth:invalid, web_bot_auth:expired and web_bot_auth:unknown_bot.

What a signature proves

A valid signature proves the request came from whoever holds a key published under a given domain. That is a strong statement about the operator. It still says nothing about which end user the agent serves, whether that user authorized the action, or whether the operator is one you want to deal with. You decide which operators to trust.

What detection can’t tell you

Every signal above answers “who sent this request?” None answers “what does this agent want, and may it have it?” An agent that reads your checkout pages might be comparing prices for a customer, or it might be scraping inventory. From the request alone, you often can’t tell.

If agents visit your site to get things done, such as order status, returns, quotes or bookings, give them a direct route. Publish an A2A Agent Card at /.well-known/agent-card.json that lists the tasks you accept and the authentication you require. An agent that uses it:

  • sends typed requests you can validate, instead of driving pages built for people;
  • authenticates with a scheme you chose and declared in the card;
  • gets clear answers and errors, and a task you can track;
  • stays out of your page-rendering path, which you can then protect more strictly.

Tutorial: publish your first Agent Card shows how. Making your customer service agent-ready covers which tasks to expose and how to protect them.

Practical steps for site owners

  1. Log the evidence. Keep the user agent, source IP, and any Signature-Agent, Signature-Input and Signature headers for automated traffic.
  2. Verify claims. When a request claims a known bot token, check the IP against that operator’s published file or DNS method. Treat failures as impersonation.
  3. State preferences in robots.txt, per token, knowing that user-triggered fetchers may not follow it and that it enforces nothing.
  4. Enforce at the edge. Put the rules that matter, such as blocking training crawlers or rate-limiting unverified bots, in your CDN or WAF.
  5. Prefer signatures where available. If your CDN verifies Web Bot Auth, turn it on and write rules on the verified result. If you verify yourself, cache key directories, reject expired signatures, and keep the accepted time window short.
  6. Don’t block user-triggered agents by reflex. A request from ChatGPT-User or Google-Agent may be your customer trying to reach you.
  7. Offer an agent endpoint for the tasks agents come to do, and move that traffic off your pages.

Questions

Can I trust a request whose user agent says GPTBot or ClaudeBot?
Not on the string alone, because any client can send it. Check the source IP against the operator's published range file, or use reverse and forward DNS where the operator documents it. If the request is signed with Web Bot Auth, verify the signature instead.
If I block GPTBot, does my site disappear from ChatGPT search?
OpenAI documents GPTBot and OAI-SearchBot as separate controls. GPTBot relates to model training and OAI-SearchBot to search results in ChatGPT, so disallowing one does not disallow the other. Check each vendor's documentation, because the split differs between companies.
Is Web Bot Auth a finished standard?
No. HTTP Message Signatures is RFC 9421, a Standards Track RFC from February 2024. Web Bot Auth builds on it and is being developed in the IETF webbotauth working group. As of September 26, 2026 its protocol document is a working group Internet-Draft, draft-ietf-webbotauth-httpsig-protocol-00.

Sources

  1. Overview of OpenAI Crawlers (OpenAI developer documentation) (accessed )
  2. ChatGPT Work's Cloud browser allowlisting (OpenAI Help Center) (accessed )
  3. Does Anthropic crawl data from the web, and how can site owners block the crawler? (Claude Help Center) (accessed )
  4. Google's common crawlers (Google Crawling Infrastructure) (accessed )
  5. Google User-Triggered Fetchers (Google Crawling Infrastructure) (accessed )
  6. Verify Requests from Google Crawlers and Fetchers (accessed )
  7. Perplexity Crawlers (Perplexity documentation) (accessed )
  8. RFC 9309: Robots Exclusion Protocol (accessed )
  9. RFC 9421: HTTP Message Signatures (accessed )
  10. Web Bot Auth (webbotauth) working group documents, IETF Datatracker (accessed )
  11. draft-ietf-webbotauth-httpsig-protocol-00: HTTP Message Signatures for automated traffic (accessed )
  12. Web Bot Auth (Cloudflare bot solutions docs) (accessed )
  13. Verified bots (Cloudflare bot solutions docs) (accessed )
  14. Authenticate legitimate AI agent traffic with AWS WAF Bot Control (AWS Security Blog) (accessed )
  15. A2A Protocol Specification, section 8: Agent Discovery (accessed )