Scraper Suddenly Returning 403? Read the Response Headers First
If a scraper that worked last week now returns HTTP 403 from Cloudflare, you are hitting web scraping blocked by cloudflare at the transport layer, not at the page layer. Cloudflare scored the request before your code saw a byte of HTML, using IP reputation, TLS and HTTP/2 fingerprints, JavaScript execution and session history. You fix it by matching a real browser’s transport fingerprint, moving to clean residential IPs, and keeping one coherent session per identity.

📌 TL;DR Executive Summary
- Core Takeaway: A 403 carrying
cf-mitigated: challengemeans Cloudflare served a challenge your client never executed. The fix is transport-level: TLS impersonation via curl_cffi, or a driver-patched browser, plus residential IPs. - Key Risk/Challenge: Datacenter IPs, a Chrome user agent sitting on a Python TLS fingerprint, and clearance cookies that expire in about 30 minutes all break long crawls mid-run.
- Recommended Solution: One isolated browser profile per scraper identity, each with its own fingerprint, sticky proxy exit and cookie jar, so a burned identity never contaminates the rest.
What Actually Changed: Four Triggers Behind a Sudden Block
A Cloudflare block rarely appears because you edited your parser. It appears because the request’s identity crossed a threshold. Most cases of web scraping blocked by cloudflare trace back to one of four triggers: the site switched on Bot Fight Mode or moved behind Cloudflare; your proxy range accumulated abuse history from other scrapers; your concurrency climbed past a configured rate limit; or a long session outlived the clearance cookie it started with.
The response tells you which layer failed. Match it against this table before you change a line of code.
| What you received | Signal | What it means | First move |
|---|---|---|---|
| 403 with a “Just a moment…” page | cf-mitigated: challenge |
A challenge page was served; your client never ran the JavaScript | Use a TLS-impersonating client or a real patched browser |
| 403, Cloudflare error 1020 | Error 1020 page | A firewall rule blocked you by country, user-agent pattern or IP range | Change the blocked attribute — this is a rule, not a bot score |
| 403, Cloudflare error 1010 | Error 1010 page | The site owner manually banned your IP address | Move to a new IP identity; retry logic will not help |
| 429 or Cloudflare error 1015 | Error 1015 page | You passed the rate limit the owner configured | Lower concurrency, add randomised delays |
| Connection or handshake error | Error 1012 | The handshake looked like a compromised TLS connection | Fix protocol and cipher settings, then impersonate a browser |
| 200 OK, generic nonsense content | No Cloudflare markers | Likely an AI Labyrinth honeypot page, not real data | Validate against your selectors and stop the run |
Log the full header set on every request, not just the status code. A server: cloudflare value confirms the request reached the edge, and a cf-ray ending in a datacenter code such as LAX shows where it was evaluated. Without those two signals you cannot separate an IP ban from a rate limit from a fingerprint mismatch, and you keep applying the wrong fix.
What Cloudflare Checks Before Your Parser Sees Any HTML
Cloudflare fronts more than 24 million active sites — an estimated 22% of the web — and its bot score runs from 1 to 99, where 1 means definitely a bot and 99 definitely human. Site owners set their own block and challenge thresholds, which is why one target lets your scraper through and the next one kills it.
The order of those checks matters, because web scraping blocked by cloudflare at the IP stage needs a completely different fix from a block at the JavaScript stage.
1. IP reputation is checked first
Before any fingerprint is examined, the edge looks at where the request came from: ASN, whether the range belongs to a hosting provider, and the abuse history attached to it. Datacenter IPs from cloud providers fail that check almost on sight. Residential proxies run about $5 to $8 per GB pay-as-you-go, and a clean residential IP alone still will not pass every site — but a dirty datacenter IP guarantees failure no matter what else you fix.
2. TLS and HTTP/2 fingerprints expose your client library
The TLS ClientHello your library sends produces a JA3 or JA4 hash, and the HTTP/2 handshake produces its own fingerprint through SETTINGS frame order, window sizes and header order. Python’s requests and httpx emit hashes no real browser produces. Rewriting your user agent to Chrome while the TLS hash still says Python is worse than sending an honest header, because the mismatch is a single-rule catch. The walkthrough on how to prevent browser fingerprinting breaks down which layers get checked and in what order.
3. JavaScript execution and clearance tokens
Non-interactive challenges are silent JavaScript proof-of-work checks that usually finish in under five seconds and need no visitor action. A plain HTTP client has no JavaScript engine, so it never computes the proof and never receives a clearance cookie. That single mechanism explains most of the 403s returned to Python scrapers.
4. Behaviour, and now patterns across requests
Cloudflare watches request timing, header consistency, cookie reuse and, in browser sessions, interaction signals. Its JA4 Signals work tracks inter-request patterns over time, so a session can look clean request by request and still be flagged as a whole. Machine-precise intervals between requests are among the easiest patterns to spot.
Non-Interactive, Managed, Interactive: What Each Challenge Tells You
Cloudflare documents four challenge behaviours. Which one you hit decides whether a lightweight fix works or you need a full browser.
| Challenge type | Visitor action | Typical time | What defeats it |
|---|---|---|---|
| Non-Interactive Challenge | None — injected JavaScript runs automatically | Usually under five seconds | A client that executes JavaScript |
| Managed Challenge | Cloudflare picks based on browser signals; often none | Seconds | A browser with a coherent fingerprint |
| Interactive Challenge | A visible checkbox or interaction | Depends on the visitor | A real browser session, or a solver API |
| Turnstile Pre-clearance | None once the token is issued | Persistent until the cookie expires | A stored cf_clearance cookie tied to the same IP and session |
One limitation matters most for API-style scraping: challenge pages break when the browser expects a non-HTML response, such as a fetch or XHR call. According to Cloudflare’s own challenge documentation, Turnstile Pre-clearance is the recommended path for single-page apps and API integrations for exactly that reason. If your target loads data in background requests, parsing the rendered HTML is usually more reliable than replaying the endpoint. The write-up on Cloudflare Turnstile challenges covers the pre-clearance flow in more depth.
Fixes Ordered by Effort: Quick Wins First
Work through them in this order. Most cases of web scraping blocked by cloudflare clear within the first three fixes; the last two matter when the target runs full Bot Management.
Fix 1: Confirm what you actually received
Add header logging before anything else. Print the status code, server, cf-ray and cf-mitigated for every request and keep them in your crawl log. This costs ten minutes and prevents days of guessing, because error 1020 needs a completely different response than a managed challenge.
Fix 2: Replace datacenter IPs with residential
If your proxy list is AWS, Hetzner or DigitalOcean, stop debugging fingerprints — the IP reputation check already failed. Residential proxies are the baseline for Cloudflare-protected targets. Two details matter more than the provider: keep one sticky IP per session so the clearance cookie stays valid, and watch bandwidth, because challenge pages return no data while still consuming your traffic.
Fix 3: Stop claiming to be Chrome from a non-browser client
Either send a user agent that matches your actual client, or make your client genuinely look like a browser. Half-matched headers — a Chrome UA, a Python TLS hash, a missing sec-ch-ua set — are a stronger bot signal than an honest one.
Fix 4: Match the TLS fingerprint with curl_cffi
curl_cffi impersonates real browser TLS and HTTP/2 fingerprints from Python and clears the most common fingerprint-level block without launching a browser. It handles JA3 and the Akamai HTTP/2 digest, and its current documented impersonation targets include chrome136, safari184 and safari184_ios, with the plain chrome and safari aliases following the latest releases.
from curl_cffi import requests
# Impersonate a real Chrome build so JA3/JA4 and the HTTP/2 fingerprint match.
session = requests.Session(impersonate="chrome136")
resp = session.get(
"https://example.com/catalog",
headers={"Accept-Language": "en-US,en;q=0.9"},
timeout=30,
)
print(resp.status_code)
print(resp.headers.get("cf-ray"), resp.headers.get("cf-mitigated"))
Two caveats. Impersonation targets age fast — a fingerprint that works today needs a library release to keep working as browsers change. And HTTP/3 now carries roughly 35% of internet traffic, with many sites showing less or no detection over it, but HTTP/3 impersonation in curl_cffi is a commercial-only feature, so do not plan on it as a free fix.
Fix 5: Move to a driver-patched browser
When the challenge requires JavaScript execution, you need a real browser that does not leak automation markers. Nodriver is undetected-chromedriver’s official successor: no webdriver flag, no Selenium in the stack, asynchronous by default, a fresh profile per run, cookie save and load, and an iframe-aware lookup that can find and click a human-verification checkbox by its text.
import asyncio
import nodriver as uc
async def scrape():
# Fresh profile per run; no webdriver flag, no Selenium in the stack.
browser = await uc.start(headless=False)
page = await browser.get("https://example.com/catalog")
await page.sleep(3)
# Iframe-aware lookup: finds a human-verification checkbox by its text.
await page.find("Verify you are human", best_match=True)
html = await page.get_content()
print(len(html))
browser.stop()
asyncio.run(scrape())
SeleniumBase UC Mode is the drop-in alternative if your stack is already Selenium: it patches fingerprints, prevents CDP leaks and ships Turnstile helpers behind a uc=True flag. Camoufox takes another route — a modified Firefox build with genuine fingerprints, useful when a target profiles Chrome specifically. Whichever you pick, do not disable web security, force shadow roots open, or run a bare headless browser. Nodriver’s own guidance is that those settings increase detectability; if you need headless on a server, run a virtual display instead.
Why requests and httpx Cannot Pass a Managed Challenge
No header, retry or user-agent combination gets requests or httpx through a managed challenge, because the challenge is not a check your client can answer — it is a task your client cannot perform. The response is a 403 with a challenge page, the browser is expected to run injected JavaScript and compute a proof of work, and only then does Cloudflare issue a clearance cookie. A client with no JavaScript engine never reaches step one.
Automate Web Scraping Blocked By Cloudflare With Send.win
Send.win pairs isolated, fingerprint-managed browser profiles with a full Automation API, so your scripts run in profiles that look and behave like real, separate users:
- Selenium, Puppeteer & Playwright support – drive any profile programmatically (Team plan)
- Isolated profiles – each with its own fingerprint, cookies, and storage
- Built-in residential proxies – with automatic timezone, locale, and WebRTC matching
- Desktop app for Windows, macOS & Linux – plus cloud sessions when you don’t want a local install
Try the instant cloud browser demo — no install, straight from your browser. Then compare plans: a 30-day free trial with no credit card, and paid plans from $6.99/month billed annually.
What those libraries can still do is pass sites that apply only cheap fingerprint and reputation rules, which is a large share of free-plan sites running Bot Fight Mode. Once a target uses Bot Management with JA4 fields and Turnstile, you need a browser that executes JavaScript under a coherent identity.
Honeypots, Cookie Expiry and Failures That Look Like Success
AI Labyrinth: a 200 response is not proof of data
Since mid-2026, Cloudflare has routed suspicious crawlers into AI Labyrinth honeypot pages — AI-generated content served instead of a block, designed to waste crawler resources silently. You get a 200, plausible-looking text, and nothing usable. Watch for content that never matches your selectors, links that loop back on themselves, missing structured data and oddly generic copy. A September 2026 report also describes Cloudflare blocking AI agents by default on pages that carry ads; treat the exact scope as unconfirmed, but the direction is clear.
Clearance cookies expire in about 30 minutes
Both cf_clearance and __cf_bm typically expire around the 30-minute mark, so a crawl that starts clean gets re-challenged mid-run. Design for renewal instead of treating it as a failure: detect the challenge response, pause, re-solve in a browser, persist the new cookie, and continue in the same session.
DIY Stack vs Managed Unlocker: Cost and Control
The build-versus-buy decision comes down to page volume and how much fingerprint maintenance you can absorb. Managed endpoints return HTML for a submitted URL and handle Cloudflare, CAPTCHAs and TLS internally at a per-request price. DIY gives you control and no per-request fee, but you own every regression when a browser update changes a hash.
| Approach | Cost signal | What it solves | What it does not |
|---|---|---|---|
| curl_cffi impersonation | Free library | TLS JA3/JA4 and HTTP/2 fingerprint mismatches | No JavaScript execution, so no managed challenges |
| Residential proxies | About $5–$8 per GB pay-as-you-go | IP reputation and geo consistency | Does not fix fingerprints or challenge tokens |
| Nodriver, SeleniumBase UC, Camoufox | Free, your compute and bandwidth | Driver leaks and real JavaScript execution | Slower per page; you maintain it when browsers change |
| Solver APIs (for example CapSolver) | Pay per solve | Turnstile and challenge pages you cannot pass | You still supply the proxy and the session |
| Managed unlocker (for example Bright Data Web Unlocker) | $2.49–$5.40 per 1,000 requests by volume | Cloudflare, CAPTCHAs and TLS in one endpoint | Less control; cost scales linearly with pages |
Run the arithmetic on your own volume. At $2.49 to $5.40 per 1,000 requests, 10,000 pages lands between roughly $25 and $54, and 100,000 pages between $249 and $540 — predictable, but expensive if all you needed was a fingerprint fix. DIY inverts that: near-zero marginal cost plus engineering time whenever an impersonation target goes stale. Aggregator benchmarks rank providers by success rate and latency, but pricing and results move quickly, so verify current numbers before committing. For the wider tooling picture, the comparison of the best browser for web scraping covers what each engine is good at.
Preventing Recurrence: One Isolated Profile per Scraper Identity
Repeat cases of web scraping blocked by cloudflare are usually identity hygiene problems, not tooling problems. If ten scrapers share one cookie jar, one browser profile and one rotating proxy pool, a single flagged session poisons the others, and you cannot tell which request caused it. The structural fix is one isolated profile per identity: its own fingerprint, its own sticky proxy exit, its own cookies and storage, and its own clearance tokens.
Send.win’s desktop browser is built for this shape of work. Each profile gets a coherent fingerprint — canvas, WebGL, audio, fonts and hardware spoofed at the engine level rather than injected by scripts — so no two profiles share a fingerprint. Timezone, locale, WebRTC and geolocation follow the proxy’s exit IP automatically, which matters because a US residential IP paired with a European browser locale is its own detection signal. Built-in residential proxies ship on every plan (10 with 1 GB on the free trial, 20 with 5 GB on Pro, 20 with 20 GB on Team), you can bring your own HTTP or SOCKS5 pool, and extra bandwidth is $6 per GB.
If your scraper is Playwright-based, you attach to a profile instead of launching a fresh anonymous browser, which keeps the session warm across runs. The local automation API covers Selenium, Puppeteer and Playwright and is available on the Team plan:
from playwright.sync_api import sync_playwright
# Copy this from the profile's automation settings — the port differs per profile.
CDP_URL = "http://127.0.0.1:PORT"
with sync_playwright() as p:
browser = p.chromium.connect_over_cdp(CDP_URL)
page = browser.contexts[0].new_page()
page.goto("https://example.com/catalog", wait_until="domcontentloaded")
print(page.title())
browser.close()
Three habits keep blocks from returning. Keep one proxy exit per profile permanently — never rotate an IP underneath a live session. Give each target site its own profile rather than its own thread, so cookie jars stay separate. And run profiles concurrently instead of hammering one: Send.win applies no concurrency cap to local profiles, so scaling out beats pushing up a request rate. If you are still mapping the broader strategy, the guide to web scraping without getting blocked covers the layered approach end to end.
🏆 Send.win Verdict
When your problem is identity rather than parsing, isolated profiles fix it at the root. Each Send.win profile carries its own fingerprint, its own sticky proxy exit and its own cookie jar, so a clearance cookie stays valid, a flagged session does not spread, and you can run many profiles at once instead of squeezing one identity until Cloudflare throttles it. Built-in residential proxies are included on every plan with extra bandwidth at $6 per GB, and the Team plan’s local Automation API lets you attach Selenium, Puppeteer or Playwright to a profile that already looks like a real machine.
Try Send.win free today — start the 30-day desktop trial with 10 profiles and 1 GB of residential bandwidth, or open the cloud browser preview with nothing to install.
Frequently Asked Questions
Why does my scraper suddenly get 403 from Cloudflare?
Something in the request’s identity crossed a threshold: an IP range picked up abuse history, the site enabled Bot Fight Mode, your rate passed a configured limit, or your clearance cookie expired mid-run. Check headers first — cf-mitigated: challenge means a challenge was served, while error 1020 means a firewall rule blocked you. The two need different fixes.
What does cf-mitigated: challenge mean?
It means Cloudflare returned a challenge page instead of the content you requested. Your client received a 403, no data, and a set of instructions meant for a JavaScript engine. Because plain HTTP clients cannot execute that JavaScript, they never receive a clearance cookie, and retrying returns the same page.
Can requests or httpx bypass Cloudflare at all?
Not a managed or interactive challenge — those require JavaScript execution neither library provides. They can still work against sites applying only cheap fingerprint and reputation checks, and curl_cffi fixes the TLS and HTTP/2 mismatch behind many of those blocks. For anything involving Turnstile or Bot Management, you need a real browser session.
How long does the cf_clearance cookie last, and can I reuse it across IPs?
Typically around 30 minutes, after which the site re-challenges you. Treat it as bound to the session that earned it: same IP, same profile, same cookie jar. Reusing a clearance token from a different proxy exit is a strong signal the session is being replayed, and it usually gets the token invalidated.
Is FlareSolverr still working in 2026?
FlareSolverr and puppeteer-stealth are widely treated as legacy tools now, detected on sight by current Cloudflare setups. The maintained alternatives are driver-level projects such as Nodriver, SeleniumBase UC Mode and Camoufox, which patch automation leaks close to the browser instead of through injected scripts.
What is Cloudflare error 1020 and how do I fix it?
Error 1020 means the site’s firewall rules explicitly denied your request — by country, user-agent pattern or IP range. It is not an automated bot score. Identify which attribute you match, then change it: a different exit country, or a header set that fits the client you are actually running.
Do residential proxies alone stop Cloudflare blocks?
No. Residential IPs fix the reputation check, which is only the first gate. A residential IP paired with a Python TLS fingerprint and no JavaScript engine still fails. Proxies, fingerprint matching and JavaScript execution are three separate layers, and a block clears only when all three line up.
Is scraping a Cloudflare-protected site legal?
It depends on jurisdiction, the data involved and the site’s terms of service. Court decisions on scraping public data are not uniform, and circumventing an access control can breach terms of service even where the data is public. Keep to public, non-personal data, respect stated limits, and get legal advice for anything commercial and contested.