Anti-Bot System
An anti-bot system is the layer a website uses to separate automated traffic from human visitors.
Key terms in AI search, web data, and the infrastructure that powers LLMs and AI agents.
An anti-bot system is the layer a website uses to separate automated traffic from human visitors.
An antidetect browser is a Chromium or Firefox build plus a management layer that gives each profile its own internally consistent…
An ASN is the identifier assigned to an autonomous system, meaning a network operator that announces its IP ranges to the rest of the…
A backconnect proxy is a single stable gateway hostname and port that brokers each request out through a different peer in the…
Behavioral analysis is the detection layer that scores what a client does over time rather than what it claims to be.
A browser pool is a managed set of browser instances or browser contexts that a crawler checks out, uses for a page or a session, and…
A browser profile is the persisted bundle of state that makes a session look like a returning visitor rather than a first-time arrival.
Canvas fingerprinting draws text or geometry to an offscreen HTML canvas, reads the pixels back, and hashes the result.
A CAPTCHA solving service accepts a challenge from your crawler and returns something you can submit.
A concurrency limit is the number of requests a provider will let you have in flight at the same moment.
Crawl-delay is a directive publishers place in robots.txt asking crawlers to wait a minimum number of seconds between requests, written…
Data parsing is the step that turns a fetched response into typed records.
A dedicated proxy is an IP address assigned to one customer and nobody else.
Deduplication is the machinery that stops a crawl from fetching the same thing twice and from storing the same record twice.
Distributed crawling is running one logical crawl across many workers, whether those are processes on one box, containers in a cluster…
The DOM is the in-memory tree a browser builds from an HTML document and then mutates as scripts run.
Exponential backoff is the standard retry policy for transient failures: wait a short interval after the first failure, then multiply…
HTTP/2 fingerprinting identifies a client by how it sets up and drives the connection rather than by what it puts in the headers.
Incremental crawling means re-fetching only the pages likely to have changed since the last run, instead of re-crawling the full set on…
IP reputation is the composite score that third-party intelligence services attach to an address based on its history and characteristics.
An IPv6 proxy routes your request through an IPv6 exit address rather than the IPv4 addresses that almost all proxy supply uses.
An ISP proxy is a static IP address registered to a consumer internet service provider's ASN but hosted on server hardware in a data center.
JA3 is a fingerprint of a TLS client computed from the ClientHello, the first message a client sends when opening a connection.
A JavaScript challenge is an interstitial page served in place of the content you asked for.
Pagination handling is how a crawler walks a result set that a site serves in slices.
Per-GB pricing is bandwidth-metered billing, the dominant model for residential and mobile proxy supply.
Proxy anonymity levels are a three-tier classification that proxy vendors and free-proxy lists still use to describe how much a proxy…
Proxy authentication is how a provider decides that a connection belongs to a paying customer.
Proxy geo-targeting is choosing the geographic location of the exit IP that carries your request, so the target site sees traffic…
Proxyware is the consumer-facing software that supplies residential proxy pools.
Request success rate is the share of scraping requests that succeeded.
A residential proxy routes your request through an IP address that an internet service provider assigned to a real home connection.
A rotating proxy changes the outbound IP address across your requests, either on every request or on a set schedule, drawing from a…
Scraper middleware is the interception layer a crawling framework exposes around every request and response, so behavior can be changed…
An SLA is the contractual promise a proxy or scraping API vendor makes about the service it delivers, plus what you get when the promise…
SOCKS5 is a transport layer proxy protocol that forwards raw TCP, and optionally UDP, without parsing what it carries.
A sticky session is a proxy gateway binding that holds one exit IP for a fixed time to live, or until you stop sending the session…
Subnet diversity is how widely a proxy pool spreads across distinct network blocks, usually counted at the /24 level and, one layer up…
The URL frontier is the crawler's queue of URLs that have been discovered but not yet fetched.
A web application firewall is the rule-based filtering layer that sits in front of an origin server and decides which requests reach it.
Web scraping is the practice of retrieving web pages programmatically and extracting structured records from them.
A web unblocker is a product category, not a technique: one endpoint that bundles proxy selection and rotation, TLS and header…
A WebRTC leak is the exposure of a machine's real network addresses through the browser's ICE candidate gathering.