{"slug": "someone-is-running-mass-vulnerability-scans-spoofing-ai-bots-like-claudebot", "title": "Someone is running mass vulnerability scans, spoofing AI bots like ClaudeBot", "summary": "Cloudflare reported that 2% of bot traffic across 5,000+ websites using Agent Analytics and AI Chat Referral Tracking is AI-related, up 11% from the previous 90 days, while overall bot visits fell 2%. The company identified a surge in mass vulnerability scans spoofing AI bots like ClaudeBot, with ClaudeBot accounting for 3.2% of agent activity.", "body_md": "Key ecosystem metrics across **5,000+** websites using [Agent Analytics](/products/agent-analytics) and [AI Chat Referral Tracking](/products/ai-chat-referral-tracking).\n\n↓ 2%\n\nCompared to the previous 90 days\n\nThe amount of visits from bots vs. humans\n\n↑ 11%\n\nCompared to the previous 90 days\n\nThe percentage of bot traffic that's AI-related\n\n↓ 9%\n\nCompared to the previous 90 days\n\nAI Agent\n\nAI Agent\n\nUses an actual web browser to autonomously complete complex tasks on behalf of a human user\n\nAI Assistant\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\nAI Coding Agent\n\nAI Coding Agent\n\nFetches documentation and other resources to help build software\n\nAI Data Provider\n\nAI Data Provider\n\nCrawls websites to supply structured content to AI systems as a third-party service\n\nAI Data Scraper\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\nAI Search Crawler\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\nArchiver\n\nArchiver\n\nCaptures and stores historical website snapshots for long-term digital preservation\n\nAutomated Agent\n\nAutomated Agent\n\nAutomates browser interactions programmatically without direct human supervision\n\nDeveloper Helper\n\nDeveloper Helper\n\nAssists with testing, debugging, and ensuring website functionality\n\nFetcher\n\nFetcher\n\nRetrieves web page metadata to power app features like link previews or feeds\n\nIntelligence Gatherer\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\nScraper\n\nScraper\n\nExtracts large amounts of web data, often without explicit website permission\n\nSearch Engine Crawler\n\nSearch Engine Crawler\n\nSystematically scans and indexes web pages to include in search results\n\nSecurity Scanner\n\nSecurity Scanner\n\nScans websites for security vulnerabilities, threats, and configuration weaknesses\n\nSEO Crawler\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\nUncategorized\n\nUncategorized\n\nNot yet assigned a type\n\nUndocumented AI Agent\n\nUndocumented AI Agent\n\nCrawls websites without disclosing its purpose, collecting data for an unknown AI use case\n\nHover over each agent type for more information about what they do\n\nAgent types with the most activity\n\n[bingbot](/agents/bingbot)\nSRCH\n\nSearch Engine Crawler\n\nSystematically scans and indexes web pages to include in search results\n\n8.2%\n\n[Googlebot](/agents/googlebot)\nSRCH\n\nSearch Engine Crawler\n\nSystematically scans and indexes web pages to include in search results\n\n7.9%\n\n[AhrefsBot](/agents/ahrefsbot)\nSEO\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\n6.2%\n\n[Known Agent](/agents/known-agent)\nDEV\n\nDeveloper Helper\n\nAssists with testing, debugging, and ensuring website functionality\n\n5.5%\n\n[ChatGPT-User](/agents/chatgpt-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n3.3%\n\n[ClaudeBot](/agents/claudebot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n3.2%\n\n[PetalBot](/agents/petalbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n3.2%\n\n[SemrushBot](/agents/semrushbot)\nSEO\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\n3.0%\n\n[facebookexternalhit](/agents/facebookexternalhit)\nFTCH\n\nFetcher\n\nRetrieves web page metadata to power app features like link previews or feeds\n\n2.6%\n\n[meta-externalagent](/agents/meta-externalagent)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n2.3%\n\n[Amazonbot](/agents/amazonbot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n2.3%\n\n[MJ12bot](/agents/mj12bot)\nSEO\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\n2.2%\n\n[Amzn-SearchBot](/agents/amzn-searchbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n2.2%\n\n[DotBot](/agents/dotbot)\nSEO\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\n2.1%\n\n[Applebot](/agents/applebot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n2.1%\n\nAgents with the most activity\n\nOperators with the most activity\n\nThese bots scrape website content to train AI models. Some belong to AI companies, while others belong to third-party services that resell the data. [Automatic Robots.txt](/products/automatic-robots-txt) can block unwanted scraping. Included agent types include [AI Data Providers](/agents?agent_type_url_slug=ai-data-provider) and [AI Data Scrapers](/agents?agent_type_url_slug=ai-data-scraper).\n\nAI Data Provider\n\nAI Data Provider\n\nCrawls websites to supply structured content to AI systems as a third-party service\n\nAI Data Scraper\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\nAI scraping activity by agent type over time\n\nComputers and Electronics\n\n5.3%\n\nBusiness and Industrial\n\n5.0%\n\nInternet and Telecom\n\n4.5%\n\nWebsite categories with most activity\n\n[ClaudeBot](/agents/claudebot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n27.0%\n\n[meta-externalagent](/agents/meta-externalagent)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n19.5%\n\n[Amazonbot](/agents/amazonbot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n19.2%\n\n[GPTBot](/agents/gptbot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n9.4%\n\n[Bytespider](/agents/bytespider)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n7.3%\n\n[ShapBot](/agents/shapbot)\nPVDR\n\nAI Data Provider\n\nCrawls websites to supply structured content to AI systems as a third-party service\n\n5.6%\n\n[GoogleOther](/agents/googleother)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n3.9%\n\n[CCBot](/agents/ccbot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n2.3%\n\n[Timpibot](/agents/timpibot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n1.8%\n\n[YouBot](/agents/youbot)\nPVDR\n\nAI Data Provider\n\nCrawls websites to supply structured content to AI systems as a third-party service\n\n1.4%\n\n[VelenPublicWebCrawler](/agents/velenpublicwebcrawler)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n0.8%\n\n[Diffbot](/agents/diffbot)\nPVDR\n\nAI Data Provider\n\nCrawls websites to supply structured content to AI systems as a third-party service\n\n0.5%\n\n[DeepSeekBot](/agents/deepseekbot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n0.4%\n\n[TerraCotta](/agents/terracotta)\nPVDR\n\nAI Data Provider\n\nCrawls websites to supply structured content to AI systems as a third-party service\n\n0.3%\n\n[AIWebIndex](/agents/aiwebindex)\nPVDR\n\nAI Data Provider\n\nCrawls websites to supply structured content to AI systems as a third-party service\n\n0.2%\n\nAgents doing the most AI scraping\n\nOperators doing the most AI scraping\n\nThese bots fetch website content in real time to power AI assistants, coding agents, and other retrieval-augmented generation (RAG) tasks. Pages inform responses on the spot, such as when an assistant summarizes an article or a coding agent references documentation. Included agent types include [AI Assistants](/agents?agent_type_url_slug=ai-assistant) and [AI Coding Agents](/agents?agent_type_url_slug=ai-coding-agent).\n\nAI Assistant\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\nAI Coding Agent\n\nAI Coding Agent\n\nFetches documentation and other resources to help build software\n\nAI fetching activity by agent type over time\n\nComputers and Electronics\n\n2.3%\n\nTravel and Transportation\n\n1.9%\n\nBusiness and Industrial\n\n1.8%\n\nInternet and Telecom\n\n1.7%\n\nWebsite categories with most activity\n\n[ChatGPT-User](/agents/chatgpt-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n86.4%\n\n[DuckAssistBot](/agents/duckassistbot)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n3.3%\n\n[Perplexity-User](/agents/perplexity-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n2.7%\n\n[Claude-User](/agents/claude-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n2.5%\n\n[MistralAI-User](/agents/mistralai-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n2.1%\n\n[Claude-Code](/agents/claude-code)\nCODE\n\nAI Coding Agent\n\nFetches documentation and other resources to help build software\n\n1.1%\n\n[Shap-User](/agents/shap-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n1.0%\n\n[Google-NotebookLM](/agents/google-notebooklm)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n0.4%\n\n[Cursor](/agents/cursor)\nCODE\n\nAI Coding Agent\n\nFetches documentation and other resources to help build software\n\n0.4%\n\n[Gemini-Deep-Research](/agents/gemini-deep-research)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n0.0%\n\n[meta-externalfetcher](/agents/meta-externalfetcher)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n0.0%\n\n[GoogleAgent-URLContext](/agents/googleagent-urlcontext)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n0.0%\n\n[Code](/agents/code)\nCODE\n\nAI Coding Agent\n\nFetches documentation and other resources to help build software\n\n0.0%\n\n[kagi-fetcher](/agents/kagi-fetcher)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n0.0%\n\n[QualifiedBot](/agents/qualifiedbot)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n0.0%\n\nAgents doing the most AI fetching\n\nOperators doing the most AI fetching\n\nThese bots crawl website content so it can be surfaced in AI search engines and AI-generated answers. Those answers often include citations or links back to the source pages. Included agent types include [AI Search Crawlers](/agents?agent_type_url_slug=ai-search-crawler).\n\nAI Search Crawler\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\nAI search indexing activity by agent type over time\n\nTravel and Transportation\n\n6.8%\n\nBooks and Literature\n\n4.9%\n\nComputers and Electronics\n\n4.6%\n\nArts and Entertainment\n\n4.5%\n\nWebsite categories with most activity\n\n[PetalBot](/agents/petalbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n25.4%\n\n[Amzn-SearchBot](/agents/amzn-searchbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n17.3%\n\n[Applebot](/agents/applebot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n16.7%\n\n[OAI-SearchBot](/agents/oai-searchbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n11.9%\n\n[meta-webindexer](/agents/meta-webindexer)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n10.8%\n\n[Claude-SearchBot](/agents/claude-searchbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n10.4%\n\n[PerplexityBot](/agents/perplexitybot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n3.4%\n\n[LinkupBot](/agents/linkupbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n3.4%\n\n[Google-CloudVertexBot](/agents/google-cloudvertexbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n0.3%\n\n[xAI-SearchBot](/agents/xai-searchbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n0.2%\n\n[AzureAI-SearchBot](/agents/azureai-searchbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n0.2%\n\n[AddSearchBot](/agents/addsearchbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n0.0%\n\n[Anomura](/agents/anomura)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n0.0%\n\n[MistralAI-Index](/agents/mistralai-index)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n0.0%\n\n[amazon-kendra](/agents/amazon-kendra)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n0.0%\n\nAgents doing the most AI search indexing\n\nOperators doing the most AI search indexing\n\nThese bots use browsers to autonomously navigate websites, click through pages, and make decisions to complete tasks for people. [Agentic UX best practices](https://web.dev/articles/ai-agent-site-ux) and [Google PageSpeed Insights](https://pagespeed.web.dev/) help evaluate how well websites support them. Included agent types include [AI Agents](/agents?agent_type_url_slug=ai-agent).\n\n↓ 5%\n\nCompared to the previous 90 days\n\nThe average duration of a session\n\n↓ 1%\n\nCompared to the previous 90 days\n\nThe average number of pages visited per session\n\nAI Agent\n\nAI Agent\n\nUses an actual web browser to autonomously complete complex tasks on behalf of a human user\n\nAI browsing activity by agent type over time\n\nComputers and Electronics\n\n0.0%\n\nInternet and Telecom\n\n0.0%\n\nTravel and Transportation\n\n0.0%\n\nBusiness and Industrial\n\n0.0%\n\nBooks and Literature\n\n0.0%\n\nArts and Entertainment\n\n0.0%\n\nWebsite categories with most activity\n\n[Google-Agent](/agents/google-agent)\nAGNT\n\nAI Agent\n\nUses an actual web browser to autonomously complete complex tasks on behalf of a human user\n\n43.0%\n\n[Manus-User](/agents/manus-user)\nAGNT\n\nAI Agent\n\nUses an actual web browser to autonomously complete complex tasks on behalf of a human user\n\n40.1%\n\n[ChatGPT Agent](/agents/chatgpt-agent)\nAGNT\n\nAI Agent\n\nUses an actual web browser to autonomously complete complex tasks on behalf of a human user\n\n16.9%\n\n[NovaAct](/agents/novaact)\nAGNT\n\nAI Agent\n\nUses an actual web browser to autonomously complete complex tasks on behalf of a human user\n\n0.0%\n\n[GoogleAgent-Mariner](/agents/googleagent-mariner)\nAGNT\n\nAI Agent\n\nUses an actual web browser to autonomously complete complex tasks on behalf of a human user\n\n0.0%\n\n[AmazonBuyForMe](/agents/amazonbuyforme)\nAGNT\n\nAI Agent\n\nUses an actual web browser to autonomously complete complex tasks on behalf of a human user\n\n0.0%\n\n[TwinAgent](/agents/twinagent)\nAGNT\n\nAI Agent\n\nUses an actual web browser to autonomously complete complex tasks on behalf of a human user\n\n0.0%\n\nAgents doing the most AI browsing\n\nOperators doing the most AI browsing\n\nSee which robots.txt rules are set across the web and how well agents follow them. An agent's [Robots.txt Effectiveness](#measuring-robots-txt-effectiveness) measures the effectiveness of a disallow rule for it by estimating how much the agent reduces its traffic after it's blocked.\n\nHow often\n\n[all agents](/agents) across all agent types follow robots.txt rules\n\n[LinkupBot](/agents/linkupbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n100.0%\n\n[serpstatbot](/agents/serpstatbot)\nSEO\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\n100.0%\n\n[proximic](/agents/proximic)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n100.0%\n\n[ClarityBot](/agents/claritybot)\nSEO\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\n100.0%\n\n[IAS crawler](/agents/ias-crawler)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n100.0%\n\n[Barkrowler](/agents/barkrowler)\nSEO\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\n100.0%\n\n[um-IC](/agents/um-ic)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n100.0%\n\n[SEOkicks](/agents/seokicks)\nSEO\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\n100.0%\n\n[Dragonfly](/agents/dragonfly)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n100.0%\n\n[AffsignalCrawler](/agents/affsignalcrawler)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n100.0%\n\n[AdsBot-Google-Mobile](/agents/adsbot-google-mobile)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n100.0%\n\n[meta-externalads](/agents/meta-externalads)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n100.0%\n\n[UptimeRobot](/agents/uptimerobot)\nDEV\n\nDeveloper Helper\n\nAssists with testing, debugging, and ensuring website functionality\n\n100.0%\n\n[Leikibot](/agents/leikibot)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n100.0%\n\n[TTD-Content](/agents/ttd-content)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n100.0%\n\nAgents with the best [Robots.txt Effectiveness](#measuring-robots-txt-effectiveness) percentages\n\n[Baiduspider](/agents/baiduspider)\nSRCH\n\nSearch Engine Crawler\n\nSystematically scans and indexes web pages to include in search results\n\n82.6%\n\n[SirdataBot](/agents/sirdatabot)\nUNC\n\nUncategorized\n\nNot yet assigned a type\n\n84.5%\n\n[ShapBot](/agents/shapbot)\nPVDR\n\nAI Data Provider\n\nCrawls websites to supply structured content to AI systems as a third-party service\n\n90.4%\n\n[YandexBot](/agents/yandexbot)\nSRCH\n\nSearch Engine Crawler\n\nSystematically scans and indexes web pages to include in search results\n\n91.2%\n\n[Sogou web spider](/agents/sogou-web-spider)\nSRCH\n\nSearch Engine Crawler\n\nSystematically scans and indexes web pages to include in search results\n\n92.7%\n\n[Mediapartners-Google](/agents/mediapartners-google)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n93.2%\n\n[Scrapy](/agents/scrapy)\nSCRP\n\nScraper\n\nExtracts large amounts of web data, often without explicit website permission\n\n93.8%\n\n[Amazonbot](/agents/amazonbot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n94.7%\n\n[HeadlessChrome](/agents/headlesschrome)\nAUTO\n\nAutomated Agent\n\nAutomates browser interactions programmatically without direct human supervision\n\n95.4%\n\n[linkfluence](/agents/linkfluence)\nINT\n\nIntelligence Gatherer\n\nAnalyzes web content for brand safety, competitive insights, and ad targeting\n\n95.9%\n\n[DotBot](/agents/dotbot)\nSEO\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\n96.1%\n\n[OAI-SearchBot](/agents/oai-searchbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n96.3%\n\n[ChatGPT-User](/agents/chatgpt-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n96.8%\n\n[archive.org_bot](/agents/archive-org-bot)\nARCH\n\nArchiver\n\nCaptures and stores historical website snapshots for long-term digital preservation\n\n96.9%\n\nAgents with the worst [Robots.txt Effectiveness](#measuring-robots-txt-effectiveness) percentages\n\n●\n\n[GPTBot](/agents/gptbot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n24.6%\n\n●\n\n[CCBot](/agents/ccbot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n22.6%\n\n●\n\n[ClaudeBot](/agents/claudebot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n21.7%\n\n●\n\n[Bytespider](/agents/bytespider)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n18.9%\n\n●\n\n[PerplexityBot](/agents/perplexitybot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n17.9%\n\n●\n\n[ChatGPT-User](/agents/chatgpt-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n16.1%\n\n●\n\n[anthropic-ai](/agents/anthropic-ai)\nUND\n\nUndocumented AI Agent\n\nCrawls websites without disclosing its purpose, collecting data for an unknown AI use case\n\n15.8%\n\n●\n\n[meta-externalagent](/agents/meta-externalagent)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n15.6%\n\n●\n\n[Amazonbot](/agents/amazonbot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n15.6%\n\n●\n\n[Diffbot](/agents/diffbot)\nPVDR\n\nAI Data Provider\n\nCrawls websites to supply structured content to AI systems as a third-party service\n\n14.0%\n\n●\n\n[Claude-Web](/agents/claude-web)\nUND\n\nUndocumented AI Agent\n\nCrawls websites without disclosing its purpose, collecting data for an unknown AI use case\n\n14.0%\n\n●\n\n[cohere-ai](/agents/cohere-ai)\nUND\n\nUndocumented AI Agent\n\nCrawls websites without disclosing its purpose, collecting data for an unknown AI use case\n\n14.0%\n\n●\n\n[omgili](/agents/omgili)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n13.5%\n\nAgents blocked by the most top websites\n\nSee which agents are most frequently impersonated, and how spoofing activity changes over time. A visit is considered spoofed when it claims a recognized agent identity but fails that agent's supported authentication method, such as verified IP or Web Bot Auth.\n\nActive Threat: AI Bot Spoofing Campaign\n\nWe are observing a widespread campaign impersonating AI bots to scan websites for vulnerabilities. The attacker appears to be targeting credential and configuration paths used by AI coding tools.\n\n[Contact us](#) for more information.\n\nThe percentage of impersonated website traffic for each agent identity over time\n\n●\n\n[Googlebot](/agents/googlebot)\nSRCH\n\nSearch Engine Crawler\n\nSystematically scans and indexes web pages to include in search results\n\n0.5%\n\n●\n\n[ChatGPT-User](/agents/chatgpt-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n0.1%\n\n●\n\n[GPTBot](/agents/gptbot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n0.1%\n\n●\n\n[OAI-SearchBot](/agents/oai-searchbot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n0.1%\n\n●\n\n[PerplexityBot](/agents/perplexitybot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n0.1%\n\n●\n\n[ClaudeBot](/agents/claudebot)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n0.0%\n\n●\n\n[Applebot](/agents/applebot)\nSRCH\n\nAI Search Crawler\n\nIndexes website content to possibly include as citations in AI-powered search results\n\n0.0%\n\n●\n\n[bingbot](/agents/bingbot)\nSRCH\n\nSearch Engine Crawler\n\nSystematically scans and indexes web pages to include in search results\n\n0.0%\n\n●\n\n[Perplexity-User](/agents/perplexity-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n0.0%\n\n●\n\n[MistralAI-User](/agents/mistralai-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n0.0%\n\n●\n\n[GoogleOther](/agents/googleother)\nSCRP\n\nAI Data Scraper\n\nDownloads website content to include in datasets used for training AI models such as LLMs\n\n0.0%\n\n●\n\n[AhrefsBot](/agents/ahrefsbot)\nSEO\n\nSEO Crawler\n\nAnalyzes website structure and content to identify SEO improvement opportunities\n\n0.0%\n\n●\n\n[Claude-User](/agents/claude-user)\nASST\n\nAI Assistant\n\nFetches website content in response to a user prompt, to include in an AI-generated answer\n\n0.0%\n\nThe most impersonated agent identities\n\n/.config/anthropic/credentials/default.json\n\n/.claude/settings.json\n\n/.claude.json\n\n/.hermes/.env\n\n/.openclaw/.env\n\n/.codex/config.toml\n\n/.continue/config.json\n\n/.aider.conf.yml\n\n/service-account.json\n\n/firebase-adminsdk.json\n\n/firebase-config.json\n\n/.aws/credentials\n\n/.aws/config\n\n/proc/self/environ\n\n/.git-credentials\n\n/.npmrc\n\n/.env.local\n\n/.env.production\n\n/.env.backup\n\n/backend/.env\n\n/api/.env\n\n/admin/.env\n\n/config/.env\n\n/laravel/.env\n\n/credentials.json\n\n/secrets.json\n\n/secrets.yml\n\n/key.json\n\n/rclone.conf\n\nA selection of top recently targeted paths\n\nSee which AI platforms like ChatGPT, Perplexity, and Gemini cite websites and send them human referral traffic. Citations are estimated. [Google's guide](https://developers.google.com/search/docs/fundamentals/ai-optimization-guide) explains how websites can optimize their content to be more visible in AI chat responses (GEO).\n\nComputers and Electronics\n\n2.3%\n\nTravel and Transportation\n\n1.9%\n\nBusiness and Industrial\n\n1.8%\n\nInternet and Telecom\n\n1.6%\n\nWebsite categories most frequently cited in AI chat responses\n\nTravel and Transportation\n\n0.3%\n\nComputers and Electronics\n\n0.1%\n\nBusiness and Industrial\n\n0.1%\n\nInternet and Telecom\n\n0.0%\n\nWebsite categories receiving the most referrals from AI chat\n\nThe Index is updated daily with completed days of traffic, security, and referral data from more than 5,000 websites using [Agent Analytics](/products/agent-analytics) and [AI Chat Referral Tracking](/products/ai-chat-referral-tracking). The current partial day is excluded. Percentage-change tags compare the current period with the preceding period of the same duration. Agent names, operators, and classifications come from the [Agent Directory](/agents), which is updated as new agents are discovered or existing agents change. Website categories follow the taxonomy used by [Google AdSense](https://adsense.google.com/start). Participating websites are not a random sample of the entire web, and the qualifying set can change as websites connect, disconnect, or cross activity thresholds. Results characterize the observed network and broader directional trends; they should not be interpreted as a precise census of global web traffic.\n\nOnly websites meeting minimum activity and data-quality requirements are included. Internal, test, incomplete, or anomalous data is excluded. Bot traffic percentages use total server traffic as their denominator. AI chat referral percentages use estimated human traffic, calculated by excluding identified bot visits from total server traffic. Rates are calculated for each qualifying website first, then averaged across websites and completed days. This gives each website equal weight regardless of traffic volume and prevents a small number of high-traffic websites from dominating the results. Daily charts are not smoothed, allowing normal seasonality to remain visible.\n\nAn agent's Robots.txt Effectiveness estimates the reduction in its request rate associated with a full disallow rule. For each completed day, Known Agents establishes an agent-specific baseline from qualifying websites where that agent is allowed, adjusts the baseline for the overall traffic of each website where the agent is disallowed, and compares the expected activity with the activity actually observed. Only website-day observations with sufficient site traffic, agent activity, cross-site coverage, and expected volume qualify. Scores also require repeated observations across multiple websites and days. When an agent publishes a supported authentication method, only verified traffic is attributed to it.\n\nEach qualifying website-day contributes equally. Scores range from 0%, meaning no measurable reduction, to 100%, meaning no qualifying requests were observed where the agent was disallowed. The headline Robots.txt Effectiveness metric gives each qualifying agent equal weight. Because this is an observational estimate rather than a controlled experiment, it measures an association with robots.txt rules but does not claim that robots.txt caused every observed difference. Top Blocked Bots is calculated separately using daily robots.txt scans of [Similarweb's top 1,000 websites](https://www.similarweb.com/top-websites).\n\nSpoofing statistics measure traffic from visits that claim the identity of a known agent but fail a supported authentication method, such as published IP verification or HTTP message signatures. Each agent's daily rate is calculated against total server traffic for every qualifying website, then averaged across websites. A failed check indicates that the visit was likely impersonating the named agent; it does not identify the software or operator that actually made the request. Agents without a supported authentication method are not included in these measurements.\n\nAI chat referral statistics count directly observed human visits carrying a recognized AI platform in the referring URL or campaign source. Visits without usable referral information cannot be attributed to an AI platform. Citation statistics are estimates based on requests from agents known to retrieve content for AI platforms. Those requests indicate that content may have informed a response, but they do not confirm that a source appeared as a citation to a user. Because AI platforms do not provide a complete public record of their sources, citation results should be interpreted as directional patterns rather than exact citation counts.\n\n##\n### Can journalists and media organizations use this data?\n\nAbsolutely. You may cite The Agentic Web Index with attribution and a link to this page. For interviews, fact-checking, background context, or a more specific breakdown for a story, [contact us](#) and include your deadline.\n\n##\n### Do you work with researchers?\n\nAbsolutely. We welcome thoughtful research into how agents and bots are changing the web. [Tell us about your research question](#), timeframe, and intended use. Depending on the scope and data constraints, we may be able to provide additional context, compare approaches, or explore a joint analysis.\n\n##\n### Can I request a specific analysis?\n\nYes. If you need a breakdown by agent, operator, activity type, website category, or time period that is not shown here, [contact us](#). When the underlying data supports it, we can examine the question and provide a focused analysis.\n\n##\n### How do I see these trends on my own website?", "url": "https://wpnews.pro/news/someone-is-running-mass-vulnerability-scans-spoofing-ai-bots-like-claudebot", "canonical_source": "https://knownagents.com/insights", "published_at": "2026-08-12 14:02:46+00:00", "updated_at": "2026-08-12 14:13:44.100179+00:00", "lang": "en", "topics": ["ai-agents", "ai-policy", "ai-infrastructure"], "entities": ["Cloudflare", "Agent Analytics", "AI Chat Referral Tracking", "ClaudeBot", "bingbot", "Googlebot", "AhrefsBot", "ChatGPT-User"], "alternates": {"html": "https://wpnews.pro/news/someone-is-running-mass-vulnerability-scans-spoofing-ai-bots-like-claudebot", "markdown": "https://wpnews.pro/news/someone-is-running-mass-vulnerability-scans-spoofing-ai-bots-like-claudebot.md", "text": "https://wpnews.pro/news/someone-is-running-mass-vulnerability-scans-spoofing-ai-bots-like-claudebot.txt", "jsonld": "https://wpnews.pro/news/someone-is-running-mass-vulnerability-scans-spoofing-ai-bots-like-claudebot.jsonld"}}