AI Bot Directory
Every bot that shows up in your server logs, from AI crawlers to search engines to uptime monitors: who operates each one, what it does, and how to control it.
167 bots, from the catalog Temso uses to detect bot traffic · Reviewed July 2026
167 of 167
ChatGPT-User
OpenAI
Fetches pages in real time when a ChatGPT user asks about them or clicks a link in a conversation.
Claude-User
Anthropic
Fetches pages on behalf of Claude users when a conversation needs live web content.
Perplexity-User
Perplexity
Visits pages a Perplexity user asks about, separate from the PerplexityBot index crawler.
DuckAssistBot
DuckDuckGo
Retrieves page content used to generate DuckAssist AI answers on DuckDuckGo.
Meta-ExternalFetcher
Meta
Fetches individual links on behalf of Meta AI when a user's request requires them.
MistralAI-User
Mistral AI
Fetches pages cited in Le Chat answers when users ask about live web content.
OAI-SearchBot
OpenAI
Builds the search index behind ChatGPT search. Not used for model training.
Claude-SearchBot
Anthropic
Indexes the web to improve search results referenced inside Claude.
PerplexityBot
Perplexity
Builds Perplexity's search index: the pool of pages its answers can cite.
YouBot
You.com
Crawls the web for You.com's AI search and chat products.
GPTBot
OpenAI
OpenAI's training crawler. It collects public web content that may be used to train future GPT models.
ClaudeBot
Anthropic
Anthropic's training crawler. It gathers public web data that may improve future Claude models.
CCBot
Common Crawl
Builds the open Common Crawl dataset, a widely used source of LLM training data.
Bytespider
ByteDance
ByteDance's crawler, associated with data collection for its AI products including Doubao.
Amazonbot
Amazon
Amazon's web crawler, used to improve services including Alexa's question answering.
Meta-ExternalAgent
Meta
Meta's crawler for AI training data and content indexing across its products.
Google-Extended
Not a crawler: a robots.txt token that controls whether Google uses your content for Gemini training and grounding.
Applebot-Extended
Apple
Not a crawler: a robots.txt token that opts your content out of Apple foundation model training.
Googlebot
Google's main indexer: the crawl behind Search, AI Overviews, and AI Mode.
GoogleOther
Google's generic crawler for product teams: research and development fetches outside Search.
Bingbot
Microsoft
Microsoft's indexer: the crawl behind Bing and the grounding index for Copilot.
Applebot
Apple
Apple's crawler for Siri, Spotlight, and Safari suggestions, and the fetch layer for Apple Intelligence.
DuckDuckBot
DuckDuckGo
DuckDuckGo's crawler for its search index and instant answers.
Claude-Web
Anthropic
Real-time browsing fetch used by Claude to compose answers.
anthropic-ai
Anthropic
Anthropic's general-purpose data collection agent.
Google-Agent
Google agent that fetches pages on behalf of generative experiences.
FacebookBot
Meta
Meta's general-purpose crawler used across Facebook surfaces.
Diffbot
Diffbot
Extracts structured data from pages for Diffbot's knowledge graph.
Google-NotebookLM
Fetches sources a user adds to NotebookLM so it can summarize and cite them.
MistralAI-Index
Mistral AI
Indexes web content for Mistral's Le Chat search feature.
Bravebot
Brave
Crawls pages for the Brave Search index that grounds Brave's Leo assistant.
LinerBot
LINER
Gathers web content for the Liner AI research and answer assistant.
Channel3Bot
Channel3
Indexes product detail pages for AI-powered product discovery.
ShapBot
Parallel
Crawls content for Parallel's search and extraction APIs.
Meta-WebIndexer
Meta
Crawls web content to provide search results for Meta AI users.
Amazon Bedrock
Amazon
Fetches pages added as data sources to Amazon Bedrock knowledge bases.
Google-CloudVertexBot
Crawls site-owner-requested pages to build Vertex AI Search agents.
Amazon Kendra
Amazon
Indexes content for Amazon Kendra enterprise search.
Amazon Q
Amazon
Crawls content for the Amazon Q Business generative assistant.
Coveobot
Coveo
Indexes content for Coveo's enterprise search and generative experiences.
Atlassian Rovo
Atlassian
Crawls and indexes content for Atlassian Rovo's AI search and agents.
AI2Bot
Ai2
Crawls the web for content to train Allen Institute open-source AI models.
Cotoyogi
ROIS
Collects Japanese-language data resources for AI training.
Omgilibot
Webz.io
Crawls public web content for Webz.io data feeds licensed for AI training.
VelenPublicWebCrawler
Webz.io
Collects public web content for Webz.io's licensed data feeds.
SBIntuitionsBot
SB Intuitions
Collects web data for SB Intuitions' AI development and analysis.
SemanticScholarBot
Ai2
Crawls academic PDFs for Semantic Scholar research discovery.
ImagesiftBot
Hive
Scrapes public images to support Hive's web intelligence products.
YandexAdditional
Yandex
Collects web content for Yandex's YandexGPT and generative AI products.
Googlebot-Image
Indexes images for Google Image Search.
Googlebot-Video
Indexes video content for Google Video Search.
Googlebot-News
Indexes articles for Google News.
YandexBot
Yandex
Yandex's primary search-index crawler.
Baiduspider
Baidu
Baidu's web crawler for Chinese-language search.
Slurp
Yahoo
Yahoo's legacy search-index crawler.
yacybot
YaCy
Crawler for the decentralized YaCy peer-to-peer search engine.
Pinterestbot
Indexes pages so Pinterest can show rich pin previews.
Sogou
Sogou
Sogou's web/news/inst spider used by its Chinese search products.
Mojeekbot
Mojeek
Independent search-index crawler operated by Mojeek.
SeznamBot
Seznam
Crawler for the Czech Seznam search engine.
Storebot-Google
Crawls shopping content for all Google Shopping surfaces.
Yeti
Naver
Naver's web crawler (Yeti) for South Korea's largest search engine.
SeekportBot
Seekport
Crawler for the German Seekport search engine, operated by SISTRIX.
Qwantbot
Qwant
Crawls and indexes content for the Qwant search engine.
PetalBot
Huawei
Web crawler operated by Huawei's Petal Search engine.
Algolia Crawler
Algolia
Extracts site content and makes it searchable via Algolia.
GeedoProductSearch
Geedo
Indexes product information from e-commerce sites for Geedo.
AmazonProductDiscovery
Amazon
Collects public product details to improve product info on Amazon.
AhrefsSiteAudit
Ahrefs
Powers Ahrefs' Site Audit tool for technical and on-page SEO issues.
AhrefsBot
Ahrefs
Builds the Ahrefs marketing-intelligence backlink and content database.
SiteAuditBot
Semrush
Semrush's Site Audit crawler for on-page and technical SEO issues.
SemrushBot
Semrush
Semrush's crawler for SEO, content, and competitive research.
DataForSeoBot
DataForSEO
Builds and maintains DataForSEO's backlink database.
SERankingBacklinksBot
SE Ranking
Discovers and analyzes backlink profiles for SE Ranking.
Barkrowler
Babbar
Babbar's crawler that fuels its graph representation of the web for SEO tools.
DotBot
Moz
Moz's crawler for its Link Explorer tool and Links API.
MJ12bot
Majestic
Majestic-12's crawler for backlink analysis and web-structure mapping.
ClarityBot
seoClarity
seoClarity's crawler for technical SEO audits and content analysis.
Lumar
Lumar
Lumar (formerly Deepcrawl) website-intelligence technical health crawler.
Marfeel Audits
Marfeel
Re-crawls traffic-receiving URLs to detect structured data and HTML issues.
Screaming Frog
Screaming Frog
The Screaming Frog SEO Spider used for site audits and technical SEO.
Sitebulb
Sitebulb
Website crawler for technical SEO audits.
Seobility
Seobility
Online SEO software crawler that audits pages to improve rankings.
ContentKing
Conductor
Continuously audits sites for performance and visibility (Conductor Monitoring).
Brightbot
Bright Data
Bright Data's crawler that monitors website health and data-collection ethics.
GTmetrix
GTmetrix
Measures site loading speed and performance.
Chrome Lighthouse
Lighthouse page-experience and performance auditing.
LogRocket
LogRocket
Captures and caches web assets for LogRocket session-replay playback.
AwarioBot
Awario
Collects web data for Awario's social-media and brand-mention monitoring.
AwarioSmartBot
Awario
One of Awario's primary crawlers discovering new and updated web data.
Trendiction
Talkwalker
Collects public web data for social-media and media-intelligence monitoring.
AdsBot-Google-Mobile
Checks mobile ad landing-page quality for Google Ads.
AdsBot-Google
Checks ad landing-page quality for Google Ads.
Google-Display-Ads-Bot
Verifies site eligibility during the AdSense approval process.
Mediapartners-Google
Crawls participating sites to serve relevant AdSense/AdMob ads.
AdIdxBot
Microsoft
Bing Ads crawler for quality control of ads and destination sites.
Amazon AdBot
Amazon
Determines site content for Amazon advertising relevance.
CriteoBot
Criteo
Analyzes web content to serve relevant contextual ads for Criteo.
IAS Crawler
Integral Ad Science
IAS crawler analyzing content for brand safety and suitability.
Proximic
Comscore
Comscore's crawler for contextual content analysis to match ad campaigns.
Quantcastbot
Quantcast
Ad quality assurance and interest-based-audience content analysis.
TTD-Content
The Trade Desk
Verifies content and quality of ad placements for The Trade Desk's DSP.
meta-externalads
Meta
Crawls the web to improve Meta's advertising and business products.
OAI-AdsBot
OpenAI
Validates the safety of web pages submitted as ads on ChatGPT.
Yahoo Ad Monitoring
Yahoo
Crawls Yahoo advertising landing pages to analyze content quality.
Amazon Route 53 Health Check
Amazon
AWS Route 53 endpoint health-check service.
Better Stack
Better Stack
Better Stack uptime monitoring and alerting.
Checkly
Checkly
Checkly synthetic monitoring and alerting.
Datadog Synthetics
Datadog
Datadog synthetic tests verifying availability and performance.
DigitalOceanUptimeBot
DigitalOcean
Checks the health of a URL or IP address.
Google Stackdriver
Google Cloud uptime checks and availability monitoring.
Google-InspectionTool
Search Console URL Inspection and Rich Results Test.
Google-Safety
Abuse-specific crawling such as malware discovery.
HetrixTools
HetrixTools
Uptime and performance checks.
Hydrozen
Hydrozen
Monitors availability of sites, cronjobs, APIs, domains, and SSL.
OhDearBot
Oh Dear
Uptime checks, broken-link detection, and mixed-content scanning.
Pingdom
SolarWinds
Pingdom uptime and performance checks.
Redirect.pizza
Redirect.pizza
Ensures redirect destination URLs are reachable.
Sansec
Sansec
Monitors online stores for malicious code and digital skimming.
Sentry Uptime
Sentry
Health checks on configured URLs for availability.
SISTRIX Uptime
SISTRIX
Continuous website availability monitoring.
Site24x7
ManageEngine
Uptime and performance checks.
StatusCake
StatusCake
Monitors uptime, page speed, and SSL certificates.
Updown.io
Updown.io
Uptime and performance checks.
Uptime Robot
UptimeRobot
Uptime monitoring and alerting.
CensysInspectBot
Censys
Internet-wide scanning of publicly accessible services.
Detectify
Detectify
Web security scanner for automated tests and attack-surface monitoring.
InternetMeasurementBot
Driftnet
Discovers and measures publicly exposed services.
Cookiebot
Usercentrics
Automates cookie-law compliance and consent scanning.
TermlyBot
Termly
Detects and categorizes first- and third-party cookies.
UsercentricsBot
Usercentrics
Scans sites for data-processing services and third-party technologies.
Stripebot
Stripe
Crawls Stripe merchant sites for service-delivery and compliance.
Google Site Verifier
Fetches Search Console verification tokens.
facebookexternalhit
Meta
Fetches shared links on Meta platforms to generate rich previews.
Slackbot-LinkExpanding
Slack
Fetches metadata from shared links to create rich previews in Slack.
Slack-ImgProxy
Slack
Fetches and caches images posted in Slack channels.
Slackbot
Slack
Slack's general-purpose bot for API requests and integrations.
Buffer Link Preview
Buffer
Generates rich previews when Buffer users share links.
Iframely
Iframely
Fetches page metadata to generate rich link previews.
SnapURLPreviewBot
Snap
Generates previews of URLs shared on Snap platforms.
MicrosoftPreview
Microsoft
Generates page snapshots for Microsoft products.
Google Image Proxy
Image caching proxy used by Gmail and other Google services.
Marfeel Preview
Marfeel
Renders preview experiences for mobile and desktop.
Marfeel Social
Marfeel
Used for social experiences across Facebook, X, Telegram, Reddit, LinkedIn.
Marfeel Flowcards
Marfeel
Fetches content for Flowcards that load from specific URLs.
Vemetric Favicon
Vemetric
Fetches favicons in the highest quality available.
GitHub Camo
GitHub
GitHub's image proxy service.
Chrome Prefetch Proxy
Fetches traffic-advice to enable privacy-preserving prefetch hints.
Stripe Webhooks
Stripe
Real-time event notifications for payment processing.
PayPal IPN
PayPal
Instant Payment Notification callbacks for PayPal events.
Shopify Webhooks
Shopify
Keeps apps in sync with Shopify data after events.
GitHub Hookshot
GitHub
GitHub webhooks for events like push and pull request.
Customer.io Webhooks
Customer.io
Webhook service for event-driven marketing automation.
Sanity Webhooks
Sanity
Real-time notifications for Sanity content changes.
Twilio Proxy
Twilio
Handles communications between end-users and apps via Twilio.
Atlassian HttpClient
Atlassian
Atlassian's HTTP client library — Jira/Confluence integrations, app calls, and webhook deliveries.
Razorpay Webhook
Razorpay
Real-time callbacks for Razorpay payment events.
Google Feedfetcher
Crawls RSS/Atom feeds for Google News and PubSubHubbub.
Google Publisher Center
Fetches feeds publishers supply for Google News landing pages.
Apple Podcasts
Apple
Accesses URLs for registered Apple Podcasts content.
FlipboardProxy
Fetches and prepares website content for the Flipboard app.
Awario RSS
Awario
One of Awario's crawlers specialized in collecting RSS feed data.
Google Read Aloud
Fetches and reads out web pages using text-to-speech on request.
APIs-Google
Delivers push-notification messages for Google APIs.
RyeBot
Rye
Powers automated checkout on behalf of shoppers with consent.
Amazon Seller Listing
Amazon
Lets sellers provide a URL to create Amazon product pages.
TangibleeBot
Tangiblee
Collects product data for visualization and virtual try-on.
Every kind of bot, and what blocking it costs you
A robots.txt policy that treats every AI bot the same is usually a mistake. Each category collects content for a different purpose, with different consequences for your AI search visibility.
AI Assistant Fetchers
9Agents that fetch a page in real time because a user asked an AI assistant about it. They don't build an index; each visit maps to a live conversation.
Blocking these removes your pages from live AI answers at the exact moment a user is asking about you or your category.
AI Search Crawlers
16Crawlers that build the retrieval indexes behind AI search products like ChatGPT search and Perplexity. They decide which pages can be cited as sources.
Blocking these keeps your content out of the source index: AI search products can't cite pages they can't crawl.
AI Training Crawlers
18Crawlers that collect content to train foundation models. Data gathered today shapes what future model versions know about your brand.
Blocking these doesn't affect live citations, but it limits what future models learn about you from your own site, leaving third-party sources to fill the gap.
Search Engine Crawlers
24Classic search indexers that now also feed AI features: Google's AI Overviews and AI Mode build on Googlebot's index, and Microsoft Copilot builds on Bingbot's.
Blocking these removes you from both traditional rankings and the AI answers built on top of those indexes, usually the most costly block of all.
SEO & Analytics Crawlers
23Crawlers behind SEO and marketing-intelligence platforms: backlink databases, site audits, rank tracking, and brand monitoring.
Blocking these has no effect on search or AI visibility. It mainly limits what third-party tools (including your competitors' tools) can see about your site.
Advertising Bots
14Bots run by ad platforms to check landing pages, verify ad placements, and analyze content for brand safety and targeting.
If you run ads on a platform, its ad bots usually need access; blocking them can break ad approval and quality checks for your own campaigns.
Operational Bots
63Monitoring, security scanning, link previews, webhooks, and feed fetchers: the utility traffic that keeps integrations and shared links working.
Most of this traffic is service-critical or user-triggered. Webhooks and previews don't follow robots.txt, and blocking them tends to break things you rely on.
See which bots visit your site
Temso helps marketers track every AI crawler and bot that visits your site, and runs regular audits to make sure your website is always performing at its full potential.
Explore Website Audits