AI Bot Directory

Every bot that shows up in your server logs, from AI crawlers to search engines to uptime monitors: who operates each one, what it does, and how to control it.

167 bots, from the catalog Temso uses to detect bot traffic · Reviewed July 2026

167 of 167

ChatGPT-User

OpenAI

Fetches pages in real time when a ChatGPT user asks about them or clicks a link in a conversation.

AI AssistantRespects robots.txt

Claude-User

Anthropic

Fetches pages on behalf of Claude users when a conversation needs live web content.

AI AssistantRespects robots.txt

Perplexity-User

Perplexity

Visits pages a Perplexity user asks about, separate from the PerplexityBot index crawler.

AI AssistantUser-triggered; robots.txt may not apply

DuckAssistBot

DuckDuckGo

Retrieves page content used to generate DuckAssist AI answers on DuckDuckGo.

AI AssistantRespects robots.txt

Meta-ExternalFetcher

Meta

Fetches individual links on behalf of Meta AI when a user's request requires them.

AI AssistantUser-triggered; robots.txt may not apply

MistralAI-User

Mistral AI

Fetches pages cited in Le Chat answers when users ask about live web content.

AI AssistantRespects robots.txt

OAI-SearchBot

OpenAI

Builds the search index behind ChatGPT search. Not used for model training.

AI SearchRespects robots.txt

Claude-SearchBot

Anthropic

Indexes the web to improve search results referenced inside Claude.

AI SearchRespects robots.txt

PerplexityBot

Perplexity

Builds Perplexity's search index: the pool of pages its answers can cite.

AI SearchRespects robots.txt

YouBot

You.com

Crawls the web for You.com's AI search and chat products.

AI SearchRespects robots.txt

GPTBot

OpenAI

OpenAI's training crawler. It collects public web content that may be used to train future GPT models.

AI TrainingRespects robots.txt

ClaudeBot

Anthropic

Anthropic's training crawler. It gathers public web data that may improve future Claude models.

AI TrainingRespects robots.txt

CCBot

Common Crawl

Builds the open Common Crawl dataset, a widely used source of LLM training data.

AI TrainingRespects robots.txt

Bytespider

ByteDance

ByteDance's crawler, associated with data collection for its AI products including Doubao.

AI TrainingReported to ignore robots.txt

Amazonbot

Amazon

Amazon's web crawler, used to improve services including Alexa's question answering.

AI TrainingRespects robots.txt

Meta-ExternalAgent

Meta

Meta's crawler for AI training data and content indexing across its products.

AI TrainingRespects robots.txt

Google-Extended

Google

Not a crawler: a robots.txt token that controls whether Google uses your content for Gemini training and grounding.

AI Trainingrobots.txt control token

Applebot-Extended

Apple

Not a crawler: a robots.txt token that opts your content out of Apple foundation model training.

AI Trainingrobots.txt control token

Googlebot

Google

Google's main indexer: the crawl behind Search, AI Overviews, and AI Mode.

Search EngineRespects robots.txt

GoogleOther

Google

Google's generic crawler for product teams: research and development fetches outside Search.

Search EngineRespects robots.txt

Bingbot

Microsoft

Microsoft's indexer: the crawl behind Bing and the grounding index for Copilot.

Search EngineRespects robots.txt

Applebot

Apple

Apple's crawler for Siri, Spotlight, and Safari suggestions, and the fetch layer for Apple Intelligence.

Search EngineRespects robots.txt

DuckDuckBot

DuckDuckGo

DuckDuckGo's crawler for its search index and instant answers.

Search EngineRespects robots.txt

Claude-Web

Anthropic

Real-time browsing fetch used by Claude to compose answers.

AI Assistant

anthropic-ai

Anthropic

Anthropic's general-purpose data collection agent.

AI Training

Google-Agent

Google

Google agent that fetches pages on behalf of generative experiences.

AI Assistant

FacebookBot

Meta

Meta's general-purpose crawler used across Facebook surfaces.

Operational

Diffbot

Diffbot

Extracts structured data from pages for Diffbot's knowledge graph.

AI Training

Google-NotebookLM

Google

Fetches sources a user adds to NotebookLM so it can summarize and cite them.

AI Assistant

MistralAI-Index

Mistral AI

Indexes web content for Mistral's Le Chat search feature.

AI Search

Bravebot

Brave

Crawls pages for the Brave Search index that grounds Brave's Leo assistant.

AI Search

LinerBot

LINER

Gathers web content for the Liner AI research and answer assistant.

AI Search

Channel3Bot

Channel3

Indexes product detail pages for AI-powered product discovery.

AI Search

ShapBot

Parallel

Crawls content for Parallel's search and extraction APIs.

AI Search

Meta-WebIndexer

Meta

Crawls web content to provide search results for Meta AI users.

AI Search

Amazon Bedrock

Amazon

Fetches pages added as data sources to Amazon Bedrock knowledge bases.

AI Search

Google-CloudVertexBot

Google

Crawls site-owner-requested pages to build Vertex AI Search agents.

AI Search

Amazon Kendra

Amazon

Indexes content for Amazon Kendra enterprise search.

AI Search

Amazon Q

Amazon

Crawls content for the Amazon Q Business generative assistant.

AI Search

Coveobot

Coveo

Indexes content for Coveo's enterprise search and generative experiences.

AI Search

Atlassian Rovo

Atlassian

Crawls and indexes content for Atlassian Rovo's AI search and agents.

AI Search

AI2Bot

Ai2

Crawls the web for content to train Allen Institute open-source AI models.

AI Training

Cotoyogi

ROIS

Collects Japanese-language data resources for AI training.

AI Training

Omgilibot

Webz.io

Crawls public web content for Webz.io data feeds licensed for AI training.

AI Training

VelenPublicWebCrawler

Webz.io

Collects public web content for Webz.io's licensed data feeds.

AI Training

SBIntuitionsBot

SB Intuitions

Collects web data for SB Intuitions' AI development and analysis.

AI Training

SemanticScholarBot

Ai2

Crawls academic PDFs for Semantic Scholar research discovery.

AI Training

ImagesiftBot

Hive

Scrapes public images to support Hive's web intelligence products.

AI Training

YandexAdditional

Yandex

Collects web content for Yandex's YandexGPT and generative AI products.

AI Training

Googlebot-Image

Google

Indexes images for Google Image Search.

Search Engine

Googlebot-Video

Google

Indexes video content for Google Video Search.

Search Engine

Googlebot-News

Google

Indexes articles for Google News.

Search Engine

YandexBot

Yandex

Yandex's primary search-index crawler.

Search Engine

Baiduspider

Baidu

Baidu's web crawler for Chinese-language search.

Search Engine

Slurp

Yahoo

Yahoo's legacy search-index crawler.

Search Engine

yacybot

YaCy

Crawler for the decentralized YaCy peer-to-peer search engine.

Search Engine

Pinterestbot

Pinterest

Indexes pages so Pinterest can show rich pin previews.

Search Engine

Sogou

Sogou

Sogou's web/news/inst spider used by its Chinese search products.

Search Engine

Mojeekbot

Mojeek

Independent search-index crawler operated by Mojeek.

Search Engine

SeznamBot

Seznam

Crawler for the Czech Seznam search engine.

Search Engine

Storebot-Google

Google

Crawls shopping content for all Google Shopping surfaces.

Search Engine

Yeti

Naver

Naver's web crawler (Yeti) for South Korea's largest search engine.

Search Engine

SeekportBot

Seekport

Crawler for the German Seekport search engine, operated by SISTRIX.

Search Engine

Qwantbot

Qwant

Crawls and indexes content for the Qwant search engine.

Search Engine

PetalBot

Huawei

Web crawler operated by Huawei's Petal Search engine.

Search Engine

Algolia Crawler

Algolia

Extracts site content and makes it searchable via Algolia.

Search Engine

GeedoProductSearch

Geedo

Indexes product information from e-commerce sites for Geedo.

Search Engine

AmazonProductDiscovery

Amazon

Collects public product details to improve product info on Amazon.

Search Engine

AhrefsSiteAudit

Ahrefs

Powers Ahrefs' Site Audit tool for technical and on-page SEO issues.

SEO Tools

AhrefsBot

Ahrefs

Builds the Ahrefs marketing-intelligence backlink and content database.

SEO Tools

SiteAuditBot

Semrush

Semrush's Site Audit crawler for on-page and technical SEO issues.

SEO Tools

SemrushBot

Semrush

Semrush's crawler for SEO, content, and competitive research.

SEO Tools

DataForSeoBot

DataForSEO

Builds and maintains DataForSEO's backlink database.

SEO Tools

SERankingBacklinksBot

SE Ranking

Discovers and analyzes backlink profiles for SE Ranking.

SEO Tools

Barkrowler

Babbar

Babbar's crawler that fuels its graph representation of the web for SEO tools.

SEO Tools

DotBot

Moz

Moz's crawler for its Link Explorer tool and Links API.

SEO Tools

MJ12bot

Majestic

Majestic-12's crawler for backlink analysis and web-structure mapping.

SEO Tools

ClarityBot

seoClarity

seoClarity's crawler for technical SEO audits and content analysis.

SEO Tools

Lumar

Lumar

Lumar (formerly Deepcrawl) website-intelligence technical health crawler.

SEO Tools

Marfeel Audits

Marfeel

Re-crawls traffic-receiving URLs to detect structured data and HTML issues.

SEO Tools

Screaming Frog

Screaming Frog

The Screaming Frog SEO Spider used for site audits and technical SEO.

SEO Tools

Sitebulb

Sitebulb

Website crawler for technical SEO audits.

SEO Tools

Seobility

Seobility

Online SEO software crawler that audits pages to improve rankings.

SEO Tools

ContentKing

Conductor

Continuously audits sites for performance and visibility (Conductor Monitoring).

SEO Tools

Brightbot

Bright Data

Bright Data's crawler that monitors website health and data-collection ethics.

SEO Tools

GTmetrix

GTmetrix

Measures site loading speed and performance.

SEO Tools

Chrome Lighthouse

Google

Lighthouse page-experience and performance auditing.

SEO Tools

LogRocket

LogRocket

Captures and caches web assets for LogRocket session-replay playback.

SEO Tools

AwarioBot

Awario

Collects web data for Awario's social-media and brand-mention monitoring.

SEO Tools

AwarioSmartBot

Awario

One of Awario's primary crawlers discovering new and updated web data.

SEO Tools

Trendiction

Talkwalker

Collects public web data for social-media and media-intelligence monitoring.

SEO Tools

AdsBot-Google-Mobile

Google

Checks mobile ad landing-page quality for Google Ads.

Advertising

AdsBot-Google

Google

Checks ad landing-page quality for Google Ads.

Advertising

Google-Display-Ads-Bot

Google

Verifies site eligibility during the AdSense approval process.

Advertising

Mediapartners-Google

Google

Crawls participating sites to serve relevant AdSense/AdMob ads.

Advertising

AdIdxBot

Microsoft

Bing Ads crawler for quality control of ads and destination sites.

Advertising

Amazon AdBot

Amazon

Determines site content for Amazon advertising relevance.

Advertising

CriteoBot

Criteo

Analyzes web content to serve relevant contextual ads for Criteo.

Advertising

IAS Crawler

Integral Ad Science

IAS crawler analyzing content for brand safety and suitability.

Advertising

Proximic

Comscore

Comscore's crawler for contextual content analysis to match ad campaigns.

Advertising

Quantcastbot

Quantcast

Ad quality assurance and interest-based-audience content analysis.

Advertising

TTD-Content

The Trade Desk

Verifies content and quality of ad placements for The Trade Desk's DSP.

Advertising

meta-externalads

Meta

Crawls the web to improve Meta's advertising and business products.

Advertising

OAI-AdsBot

OpenAI

Validates the safety of web pages submitted as ads on ChatGPT.

Advertising

Yahoo Ad Monitoring

Yahoo

Crawls Yahoo advertising landing pages to analyze content quality.

Advertising

Amazon Route 53 Health Check

Amazon

AWS Route 53 endpoint health-check service.

Operational

Better Stack

Better Stack

Better Stack uptime monitoring and alerting.

Operational

Checkly

Checkly

Checkly synthetic monitoring and alerting.

Operational

Datadog Synthetics

Datadog

Datadog synthetic tests verifying availability and performance.

Operational

DigitalOceanUptimeBot

DigitalOcean

Checks the health of a URL or IP address.

Operational

Google Stackdriver

Google

Google Cloud uptime checks and availability monitoring.

Operational

Google-InspectionTool

Google

Search Console URL Inspection and Rich Results Test.

Operational

Google-Safety

Google

Abuse-specific crawling such as malware discovery.

Operational

HetrixTools

HetrixTools

Uptime and performance checks.

Operational

Hydrozen

Hydrozen

Monitors availability of sites, cronjobs, APIs, domains, and SSL.

Operational

OhDearBot

Oh Dear

Uptime checks, broken-link detection, and mixed-content scanning.

Operational

Pingdom

SolarWinds

Pingdom uptime and performance checks.

Operational

Redirect.pizza

Redirect.pizza

Ensures redirect destination URLs are reachable.

Operational

Sansec

Sansec

Monitors online stores for malicious code and digital skimming.

Operational

Sentry Uptime

Sentry

Health checks on configured URLs for availability.

Operational

SISTRIX Uptime

SISTRIX

Continuous website availability monitoring.

Operational

Site24x7

ManageEngine

Uptime and performance checks.

Operational

StatusCake

StatusCake

Monitors uptime, page speed, and SSL certificates.

Operational

Updown.io

Updown.io

Uptime and performance checks.

Operational

Uptime Robot

UptimeRobot

Uptime monitoring and alerting.

Operational

CensysInspectBot

Censys

Internet-wide scanning of publicly accessible services.

Operational

Detectify

Detectify

Web security scanner for automated tests and attack-surface monitoring.

Operational

InternetMeasurementBot

Driftnet

Discovers and measures publicly exposed services.

Operational

Cookiebot

Usercentrics

Automates cookie-law compliance and consent scanning.

Operational

TermlyBot

Termly

Detects and categorizes first- and third-party cookies.

Operational

UsercentricsBot

Usercentrics

Scans sites for data-processing services and third-party technologies.

Operational

Stripebot

Stripe

Crawls Stripe merchant sites for service-delivery and compliance.

Operational

Google Site Verifier

Google

Fetches Search Console verification tokens.

Operational

facebookexternalhit

Meta

Fetches shared links on Meta platforms to generate rich previews.

Operational

Slackbot-LinkExpanding

Slack

Fetches metadata from shared links to create rich previews in Slack.

Operational

Slack-ImgProxy

Slack

Fetches and caches images posted in Slack channels.

Operational

Slackbot

Slack

Slack's general-purpose bot for API requests and integrations.

Operational

Buffer Link Preview

Buffer

Generates rich previews when Buffer users share links.

Operational

Iframely

Iframely

Fetches page metadata to generate rich link previews.

Operational

SnapURLPreviewBot

Snap

Generates previews of URLs shared on Snap platforms.

Operational

MicrosoftPreview

Microsoft

Generates page snapshots for Microsoft products.

Operational

Google Image Proxy

Google

Image caching proxy used by Gmail and other Google services.

Operational

Marfeel Preview

Marfeel

Renders preview experiences for mobile and desktop.

Operational

Marfeel Social

Marfeel

Used for social experiences across Facebook, X, Telegram, Reddit, LinkedIn.

Operational

Marfeel Flowcards

Marfeel

Fetches content for Flowcards that load from specific URLs.

Operational

Vemetric Favicon

Vemetric

Fetches favicons in the highest quality available.

Operational

GitHub Camo

GitHub

GitHub's image proxy service.

Operational

Chrome Prefetch Proxy

Google

Fetches traffic-advice to enable privacy-preserving prefetch hints.

Operational

Stripe Webhooks

Stripe

Real-time event notifications for payment processing.

Operational

PayPal IPN

PayPal

Instant Payment Notification callbacks for PayPal events.

Operational

Shopify Webhooks

Shopify

Keeps apps in sync with Shopify data after events.

Operational

GitHub Hookshot

GitHub

GitHub webhooks for events like push and pull request.

Operational

Customer.io Webhooks

Customer.io

Webhook service for event-driven marketing automation.

Operational

Sanity Webhooks

Sanity

Real-time notifications for Sanity content changes.

Operational

Twilio Proxy

Twilio

Handles communications between end-users and apps via Twilio.

Operational

Atlassian HttpClient

Atlassian

Atlassian's HTTP client library — Jira/Confluence integrations, app calls, and webhook deliveries.

Operational

Razorpay Webhook

Razorpay

Real-time callbacks for Razorpay payment events.

Operational

Google Feedfetcher

Google

Crawls RSS/Atom feeds for Google News and PubSubHubbub.

Operational

Google Publisher Center

Google

Fetches feeds publishers supply for Google News landing pages.

Operational

Apple Podcasts

Apple

Accesses URLs for registered Apple Podcasts content.

Operational

FlipboardProxy

Flipboard

Fetches and prepares website content for the Flipboard app.

Operational

Awario RSS

Awario

One of Awario's crawlers specialized in collecting RSS feed data.

Operational

Google Read Aloud

Google

Fetches and reads out web pages using text-to-speech on request.

Operational

APIs-Google

Google

Delivers push-notification messages for Google APIs.

Operational

RyeBot

Rye

Powers automated checkout on behalf of shoppers with consent.

Operational

Amazon Seller Listing

Amazon

Lets sellers provide a URL to create Amazon product pages.

Operational

TangibleeBot

Tangiblee

Collects product data for visualization and virtual try-on.

Operational

Every kind of bot, and what blocking it costs you

A robots.txt policy that treats every AI bot the same is usually a mistake. Each category collects content for a different purpose, with different consequences for your AI search visibility.

AI Assistant Fetchers

9

Agents that fetch a page in real time because a user asked an AI assistant about it. They don't build an index; each visit maps to a live conversation.

Blocking these removes your pages from live AI answers at the exact moment a user is asking about you or your category.

AI Search Crawlers

16

Crawlers that build the retrieval indexes behind AI search products like ChatGPT search and Perplexity. They decide which pages can be cited as sources.

Blocking these keeps your content out of the source index: AI search products can't cite pages they can't crawl.

AI Training Crawlers

18

Crawlers that collect content to train foundation models. Data gathered today shapes what future model versions know about your brand.

Blocking these doesn't affect live citations, but it limits what future models learn about you from your own site, leaving third-party sources to fill the gap.

Search Engine Crawlers

24

Classic search indexers that now also feed AI features: Google's AI Overviews and AI Mode build on Googlebot's index, and Microsoft Copilot builds on Bingbot's.

Blocking these removes you from both traditional rankings and the AI answers built on top of those indexes, usually the most costly block of all.

SEO & Analytics Crawlers

23

Crawlers behind SEO and marketing-intelligence platforms: backlink databases, site audits, rank tracking, and brand monitoring.

Blocking these has no effect on search or AI visibility. It mainly limits what third-party tools (including your competitors' tools) can see about your site.

Advertising Bots

14

Bots run by ad platforms to check landing pages, verify ad placements, and analyze content for brand safety and targeting.

If you run ads on a platform, its ad bots usually need access; blocking them can break ad approval and quality checks for your own campaigns.

Operational Bots

63

Monitoring, security scanning, link previews, webhooks, and feed fetchers: the utility traffic that keeps integrations and shared links working.

Most of this traffic is service-critical or user-triggered. Webhooks and previews don't follow robots.txt, and blocking them tends to break things you rely on.

See which bots visit your site

Temso helps marketers track every AI crawler and bot that visits your site, and runs regular audits to make sure your website is always performing at its full potential.

Explore Website Audits
Website audit dashboard showing bot visits, citation-bot visits, and a table of AI bots crawling the site

FAQ

Common questionsabout AI bots

An AI crawler is an automated agent that visits websites to collect content for AI systems. They fall into distinct groups: training crawlers (like GPTBot) gather data to train foundation models, AI search crawlers (like OAI-SearchBot and PerplexityBot) build the indexes AI search products cite from, and assistant fetchers (like ChatGPT-User) retrieve pages live when a user asks about them.

It depends on what you want future models to know. Blocking GPTBot keeps your content out of OpenAI's training data, but it does not remove you from ChatGPT search; that index is built by OAI-SearchBot, a separate crawler. If AI visibility matters to your business, most sites benefit from allowing both: training data shapes what models inherently say about you, and the search index determines whether you can be cited.

They are three separate OpenAI agents with different jobs. GPTBot collects content that may train future GPT models. OAI-SearchBot builds and maintains the ChatGPT search index and is not used for training. ChatGPT-User fetches individual pages in real time when a ChatGPT user asks about them. Each honors its own robots.txt token, so you can allow or block them independently.

Not directly. Live citations in ChatGPT search, Perplexity, and similar products come from search indexes and real-time fetches, which use different bots than training crawlers. But there is a long-term effect: models trained without your content only know what third parties say about you, which can degrade how accurately AI assistants describe your brand even when they cite other sources.

Add a robots.txt rule using the bot's user-agent token. For example, 'User-agent: GPTBot' followed by 'Disallow: /' blocks GPTBot from the whole site, while 'Disallow:' (empty) explicitly allows it. Each bot page in this directory shows its exact token and a ready-to-use snippet. Note that robots.txt is a voluntary standard: most major operators honor it, but compliance varies and is noted per bot.

Never trust the user-agent string alone; it is trivially spoofed. Major operators publish official IP ranges (OpenAI, Perplexity, Google, Microsoft, and DuckDuckGo all do) or support reverse-DNS verification (Googlebot, Applebot). Check requests claiming to be a given bot against the operator's published list, linked from each bot's page in this directory.

Want to know how visible your brand actually is in the AI systems these bots feed?