Rotating Proxies for Web Scraping: Python requests, Scrapy Middleware, and Rank Tracking
A practical guide to rotating proxies in Python and Scrapy - pool selection, sticky sessions versus per-request rotation, working middleware code, ban detection, and why rank trackers need location consistency more than raw IP volume.
If you scrape a few hundred pages from a single IP, you will eventually meet a rate limit, a CAPTCHA, or a soft block. Rotation is the standard fix - but bad rotation is worse than none: it burns proxy budget, makes debugging harder, and can look even more bot-like than a steady, polite crawler.
This guide covers how rotation works, how to implement it in Python with requests and in Scrapy with a downloader middleware, and how the same techniques apply to SERP rank tracking.
Before you rotate, check whether rotation is the problem
Rotation hides IP-level blocks. It does not fix:
- A pattern problem. No number of IPs will save a crawler that requests the same URL fifty times a second with a default user agent.
- A session problem. If your flow requires staying logged in, a new IP mid-session looks like account theft.
- A geo problem. If the content you need is region-locked, you need the right country, not more IPs.
Diagnose first: log status codes, response bodies, and response times per IP. If every IP gets a 200 and one specific page returns a block, your issue is request behaviour, not IP reputation.
Pick a proxy type that matches the job
| Type | Typical strength | Watch out for |
|---|---|---|
| Datacenter | Fast, high concurrency, widely available | Easily identified; blocked by strict sites |
| Residential | Blends with real user traffic | Slower; bandwidth limits; session handling matters |
| Mobile | High trust, hardest to block | Expensive; unstable; fewer concurrent sessions |
| ISP / static residential | Stable, consistent IP | Smaller pools; less rotation |
For structured data on permissive sites, datacenter is usually enough. For search results, marketplaces, and travel sites, residential or mobile is typically required. Evaluate providers on rotation granularity, session control, geo coverage, and how bans are handled - not on a headline count of IPs.
Rotation strategies
Per-request rotation
Every request leaves from an IP chosen by the pool. Many providers expose this as a single rotating gateway hostname plus a port, or as a session token in the username. Best for stateless page fetches.
Sticky sessions
The same IP is reused for a defined window or until you release it. Use this for multi-step flows: search, open listing, read detail page. In most commercial pools, stickiness is controlled by a session identifier in the proxy username, and a new identifier produces a new IP.
Geographic stickiness
For rank tracking and price checks you usually want the same city or country across a query set, because results vary by location. Rotate within a geo, not across geos.
Concurrency throttling
Rotation and rate limiting are complements, not substitutes. A crawler with five concurrent requests that respects Retry-After will outperform one with two hundred threads hammering every IP until it is banned.
Python: rotating proxies with requests
Install SOCKS support if you need it. The requests library uses PySocks underneath:
pip install "requests[socks]"
Then rotate across a pool:
import itertools
import requests
POOL = [
'http://user:[email protected]:8000',
'http://user:[email protected]:8000',
'socks5h://user:[email protected]:1080',
]
proxy_cycle = itertools.cycle(POOL)
def fetch(url, timeout=20):
proxy = next(proxy_cycle)
proxies = {'http': proxy, 'https': proxy}
response = requests.get(url, proxies=proxies, timeout=timeout)
response.raise_for_status()
return response.text
A few things worth adding in production:
- Retry with a different proxy on
429or403, not the same one. - Use
requests.Session()for connection reuse, but create one session per IP rather than one session for the whole crawl. - Log which proxy produced which result so you can drop bad IPs instead of redrawing them.
- Use
socks5h://(note theh) when you want hostnames resolved at the proxy instead of locally.
Scrapy: a rotating proxy downloader middleware
Scrapy already handles proxy authentication through its built-in HTTP proxy middleware, so you only need to set request.meta['proxy'] before the request is sent. A simple rotation middleware looks like this:
import random
class RotatingProxyMiddleware:
def __init__(self, proxies):
self.proxies = proxies
@classmethod
def from_crawler(cls, crawler):
proxies = crawler.settings.getlist('ROTATING_PROXIES')
if not proxies:
raise ValueError('ROTATING_PROXIES is empty')
return cls(proxies)
def process_request(self, request, spider):
request.meta['proxy'] = random.choice(self.proxies)
Register it in your settings module:
DOWNLOADER_MIDDLEWARES = {
'myproject.middlewares.RotatingProxyMiddleware': 350,
}
ROTATING_PROXIES = [
'http://user:[email protected]:8000',
'http://user:[email protected]:8000',
]
AUTOTHROTTLE_ENABLED = True
RETRY_HTTP_CODES = [403, 429, 500, 502, 503, 504]
Two notes:
- A naive
random.choicewill keep reusing banned IPs. Community projects such asscrapy-rotating-proxiesadd ban detection and per-proxy statistics on top of the same idea, which is worth the dependency once you scale past a handful of requests. - Scrapy's retry middleware normally retries with the same proxy. If bans are your problem, make the retry path re-run proxy selection.
Applying the same approach to rank tracking
Rank trackers have a proxy requirement that generic scrapers do not: location consistency. A keyword checked from one city and then from another will produce different SERPs and a meaningless movement graph.
Practical rules for rank tracking:
- Choose one location per tracked keyword and keep it stable across runs.
- Refresh the IP periodically so you are not flagged, but stay inside the same city or region.
- Use mobile proxies when tracking mobile results, because desktop and mobile SERPs differ enough to matter.
- Cache aggressively and check on a schedule rather than continuously.
Also worth stating plainly: scraping search engines usually violates their terms of service. If you need dependable ranking data, official APIs and licensed rank-tracking feeds are the lower-risk path. Proxies are for the cases where an API genuinely does not exist.
How to tell whether rotation is working
Track these over a whole crawl, not per request:
- Success rate per proxy, meaning 2xx responses divided by attempts.
- Block rate per proxy, with the block type recorded:
403, CAPTCHA, empty body, consent page. - Median response time per proxy. A proxy that works but takes thirty seconds a page is a net loss.
- Unique IPs actually used, versus the size of the pool you configured.
If success rate is flat across IPs, your problem is not IP reputation, and no amount of rotation will help.
Takeaway
Rotate only after you have fixed request rate, headers, and session logic. Use per-request rotation for stateless fetches and sticky sessions for anything multi-step. In Python, requests needs the SOCKS extra and a proxy per session, while in Scrapy a downloader middleware that assigns request.meta['proxy'] gets you most of the way, with ban-tracking packages handling the rest. And for rank tracking, treat geography as the requirement and rotation as the maintenance.