Scrapy Proxy Rotation: How to Rotate Proxies in Scrapy (and When a Dedicated WireGuard Gateway Is Better)
Implement proxy rotation in Scrapy with custom middleware, handle bans gracefully, and learn when a stable WireGuard gateway like NordLayer is the right choice for SEO monitoring and web scraping.
Overview
Web scraping at scale often hits rate limits, IP bans, or geo-restrictions. Scrapy, a popular Python framework, doesn't rotate proxies out of the box. This tutorial shows you how to build a rotating proxy middleware for Scrapy, configure retries, and decide when a dedicated WireGuard VPN gateway (like NordLayer) is a better fit than a pool of rotating proxies.
You'll learn two approaches:
- Rotating proxies: use many IPs per session to distribute requests and avoid blocks.
- Dedicated WireGuard gateway: use a single, stable IP for tasks that require consistency, such as rank tracking, ad verification, or accessing geo-specific content that penalises IP changes.
NordLayer provides business-grade WireGuard tunnels with dedicated gateways and flat monthly billing. It is not a rotating proxy service, but it can serve as the stable egress IP for your Scrapy spiders when consistency matters.
Prerequisites
- Python 3.7 or later
- Basic familiarity with Scrapy
- A list of HTTP/HTTPS or SOCKS5 proxies (for the rotation approach)
- Optional: a NordLayer account with a dedicated WireGuard gateway (for the stable IP approach)
- A Linux server or local machine (commands shown for Debian/Ubuntu; adapt for other systems)
Understanding Proxy Rotation in Scrapy
Scrapy uses downloader middlewares to process requests and responses. To rotate proxies, you intercept each request, assign a proxy from a pool, and optionally retry with a new proxy when a request fails.
Key concepts:
- Proxy pool: a list of proxy URLs (with credentials if needed).
- Rotation strategy: random selection, round-robin, or weighted by proxy health.
- Retry logic: re-queue failed requests with a different proxy.
When to consider a dedicated WireGuard gateway instead:
- You need a consistent IP for SEO rank tracking across multiple locations.
- Your target site blocks frequent IP changes more aggressively than static IPs.
- You want encrypted, team-managed access without maintaining a proxy list.
Steps
1. Install Scrapy and create a project
If you haven't already, install Scrapy and scaffold a new project.
pip install scrapy
scrapy startproject proxy_scraper
cd proxy_scraper
scrapy genspider example example.com
2. Create a rotating proxy middleware
Open proxy_scraper/middlewares.py and add the following middleware. It selects a random proxy for each request and retries on common block status codes.
# proxy_scraper/middlewares.py
import random
from scrapy import signals
from scrapy.exceptions import NotConfigured
class RotatingProxyMiddleware:
def __init__(self, proxy_list):
self.proxy_list = proxy_list
self.current_proxy = None
@classmethod
def from_crawler(cls, crawler):
proxy_list = crawler.settings.getlist('ROTATING_PROXY_LIST')
if not proxy_list:
raise NotConfigured('ROTATING_PROXY_LIST is empty')
return cls(proxy_list)
def process_request(self, request, spider):
self.current_proxy = random.choice(self.proxy_list)
request.meta['proxy'] = self.current_proxy
spider.logger.debug(f'Using proxy: {self.current_proxy}')
def process_response(self, request, response, spider):
if response.status in [403, 429]:
spider.logger.warning(f'Blocked with proxy {self.current_proxy}. Retrying...')
new_request = request.copy()
new_request.dont_filter = True
return new_request
return response
3. Configure settings to enable the middleware and define proxies
Edit proxy_scraper/settings.py to activate the middleware and supply your proxy list. Replace the example proxies with your own.
# proxy_scraper/settings.py
DOWNLOADER_MIDDLEWARES = {
'proxy_scraper.middlewares.RotatingProxyMiddleware': 543,
}
ROTATING_PROXY_LIST = [
'http://user:[email protected]:8080',
'http://user:[email protected]:8080',
'socks5://user:[email protected]:1080',
]
# Optional: retry settings
RETRY_TIMES = 5
RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408, 429, 403]
4. Test rotation with a simple spider
Run your spider and watch the DEBUG logs to confirm different proxies are used.
scrapy crawl example -L DEBUG
If you see lines like Using proxy: http://..., the middleware is working. If all requests use the same proxy, check the middleware order and that ROTATING_PROXY_LIST is not empty.
5. Improve reliability with health checks and backoff
For production scraping, add basic proxy health tracking. This example marks a proxy as failed after repeated blocks and temporarily removes it.
# proxy_scraper/middlewares.py (enhanced)
import random
from collections import defaultdict
class RotatingProxyMiddleware:
def __init__(self, proxy_list):
self.proxy_list = proxy_list
self.failure_count = defaultdict(int)
self.max_failures = 3
@classmethod
def from_crawler(cls, crawler):
proxy_list = crawler.settings.getlist('ROTATING_PROXY_LIST')
return cls(proxy_list)
def process_request(self, request, spider):
available = [p for p in self.proxy_list if self.failure_count[p] < self.max_failures]
if not available:
spider.logger.error('All proxies exhausted')
return
proxy = random.choice(available)
request.meta['proxy'] = proxy
def process_response(self, request, response, spider):
proxy = request.meta.get('proxy')
if response.status in [403, 429]:
self.failure_count[proxy] += 1
new_request = request.copy()
new_request.dont_filter = True
return new_request
else:
self.failure_count[proxy] = 0
return response
6. When to switch to a dedicated WireGuard gateway
Rotating proxies are excellent for large-scale, stateless scraping. But some tasks need a stable, dedicated IP:
- SEO rank tracking where you need consistent location data
- Accessing APIs that tie sessions to an IP
- Team-based scraping where multiple users share the same egress IP
NordLayer WireGuard tunnels give you a dedicated gateway with flat monthly billing. You can route all Scrapy traffic through the tunnel, effectively using a single static IP.
7. Set up WireGuard on Linux and route Scrapy through it
Install WireGuard and configure a tunnel using the credentials from your NordLayer gateway.
sudo apt update
sudo apt install wireguard
Create a configuration file at /etc/wireguard/wg0.conf (replace placeholders with your NordLayer values).
[Interface]
PrivateKey = <your-private-key>
Address = 10.0.0.2/32
DNS = 1.1.1.1
[Peer]
PublicKey = <nordlayer-public-key>
Endpoint = <gateway-ip>:51820
AllowedIPs = 0.0.0.0/0
PersistentKeepalive = 25
Bring the tunnel up:
sudo wg-quick up wg0
Now all Scrapy traffic (and any other traffic) exits through the NordLayer gateway. Verify your public IP has changed:
curl ifconfig.me
Run your Scrapy spider as usual. It will use the stable gateway IP.
8. Combine approaches for hybrid scraping
You can use both: route most requests through rotating proxies, but send requests that need a consistent IP through the WireGuard tunnel. In Scrapy, you can conditionally set the proxy to None (which uses the system's default route, i.e., the WireGuard tunnel) for specific requests.
# Example: in a spider, for rank-tracking requests, don't set a proxy
# The default route (WireGuard) will be used.
def start_requests(self):
# Rotating proxies for general pages
yield scrapy.Request('https://example.com/general', meta={'proxy': None})
# For rank tracking, bypass the rotating middleware
yield scrapy.Request('https://example.com/serp', meta={'proxy': 'direct'})
Then modify the middleware to skip rotation when proxy == 'direct':
def process_request(self, request, spider):
if request.meta.get('proxy') == 'direct':
return None # use default route
# ... rest of rotation logic
Troubleshooting
- Middleware not applied: Ensure
DOWNLOADER_MIDDLEWARESincludes the correct path and that the value (543) doesn't conflict with other middlewares. Check Scrapy's default middleware order. - All requests still use the same IP: Verify that
ROTATING_PROXY_LISTis populated and that the middleware'sprocess_requestis running. Add aprintor a debug log. - Frequent 403/429 errors: Your proxies may be blocked or too slow. Reduce
CONCURRENT_REQUESTSandDOWNLOAD_DELAY, or switch to higher-quality proxies. Consider a dedicated WireGuard gateway if the target site blocks rotating IPs aggressively. - WireGuard tunnel fails to connect: Double-check the private key, public key, endpoint, and firewall rules (UDP port 51820 by default). Run
sudo wg showto see handshake status. - Scrapy ignores the WireGuard tunnel: If you set
AllowedIPs = 0.0.0.0/0, all traffic should route through the tunnel. Confirm withip routeand test withcurl ifconfig.me.
Summary
You now know how to implement proxy rotation in Scrapy using custom middleware, add retry logic, and monitor proxy health. For projects that require a stable, dedicated IP—such as SEO rank tracking or team-based access—a WireGuard gateway like NordLayer can complement or replace rotating proxies. Choose the approach that matches your scraping goals: rotation for scale and anonymity, dedicated gateways for consistency and control.