Cloudflare's Content Signals policy put ai-train=no into 3.8 million robots.txt files by default. What that advisory line means for teams collecting web data.
For three years residential proxy rates only fell. In 2026 the majors pulled their coupons and dropped pay-as-you-go entry plans. What that means at renewal.
Every major scraping vendor now ships an MCP server. What handing an AI agent a metered scraping tool changes for cost and reliability, and what it doesn't.
Web Bot Auth lets a bot prove identity with a cryptographic signature. Cloudflare and AWS already check it. What that closing lane means for collection teams.
A federal judge dismissed Google's DMCA anti-circumvention suit against SerpApi. What the ruling means for any scraper that gets past an anti-bot system.
The Ninth Circuit freed Perplexity's Comet agent because the user, not the company, accessed Amazon. Why that CFAA reasoning doesn't cover server-side scraping.
The RSL licensing standard turns robots.txt from a yes/no crawl rule into a machine-readable license. Why a robots-compliant crawler can still miss it.
The 2026 State of Web Scraping survey found most teams still don't run AI in the collection loop. Where AI web scraping landed, and where the money went.
Cloudflare pay per crawl meters only declared AI crawlers, not the proxied collection most teams run. The web is splitting into two lanes. How to sort targets.
Browser agents keep failing benchmarks that involve changing state, yet they are being deployed in production on portal and back-office workflows. The split is not a contradiction, and it tells you where to point them.
Fifteen vendors sell isolated, fingerprint-spoofed browser profiles. Hardware-bound sessions and server-side signals are steadily making the fingerprint the least interesting part of the product.
Scrapy, Crawlee and the newer LLM-oriented crawlers are all healthy. The reason teams still buy an API has almost nothing to do with the framework and everything to do with what sits between it and the origin.
California's deletion regime turns registered data brokers into an operational dependency with dates attached. What that means when person-level records are an input to your pipeline.
Running headless Chrome at crawl scale now means running something much closer to a full browser. A look at what that did to per-session cost and when a non-Chromium engine or a rented fleet wins.
Two dozen point-and-click scrapers now compete with prompt-based extraction on one side and agent runtimes on the other. The surviving argument for the category is maintenance, not ease of use.
Proxy vendors keep shipping finer geographic and network targeting. For most collection jobs the granularity you need is coarser than the granularity you are sold, and paying for it costs you pool size.
Acquisitions and outside capital have quietly reshaped the proxy market. What that means when you renew a contract with a vendor that no longer answers only to itself.
Headline dollar-per-gigabyte rates keep falling while the number that decides your bill moved somewhere else entirely. A look at how proxy billing is fragmenting and how to compare offers that no longer share a unit.
A live case in the Southern District of New York puts a residential proxy operator and a proxy network on the same caption as the AI company that used them. The contract questions that follow for anyone buying collection capacity.
Law enforcement actions against proxy operators and the SDK monetisation model behind consumer IP pools have made peer sourcing a diligence item rather than a footnote.
Nearly every large proxy network now sells a managed unblocking endpoint on top of its pool. That product, not the raw proxy, is what most teams are actually buying in 2026.