Skip to content
Advanced / 7 min read

Scrapy Proxy Rotation: How to Rotate Proxies in Scrapy (and When a Dedicated WireGuard Gateway Is Better)

Implement proxy rotation in Scrapy with custom middleware, handle bans gracefully, and learn when a stable WireGuard gateway like NordLayer is the right choice for SEO monitoring and web scraping.

Linux WireGuard Web Scraping SEO Monitoring

Overview

Web scraping at scale often hits rate limits, IP bans, or geo-restrictions. Scrapy, a popular Python framework, doesn't rotate proxies out of the box. This tutorial shows you how to build a rotating proxy middleware for Scrapy, configure retries, and decide when a dedicated WireGuard VPN gateway (like NordLayer) is a better fit than a pool of rotating proxies.

You'll learn two approaches:

  • Rotating proxies: use many IPs per session to distribute requests and avoid blocks.
  • Dedicated WireGuard gateway: use a single, stable IP for tasks that require consistency, such as rank tracking, ad verification, or accessing geo-specific content that penalises IP changes.

NordLayer provides business-grade WireGuard tunnels with dedicated gateways and flat monthly billing. It is not a rotating proxy service, but it can serve as the stable egress IP for your Scrapy spiders when consistency matters.

Prerequisites

  • Python 3.7 or later
  • Basic familiarity with Scrapy
  • A list of HTTP/HTTPS or SOCKS5 proxies (for the rotation approach)
  • Optional: a NordLayer account with a dedicated WireGuard gateway (for the stable IP approach)
  • A Linux server or local machine (commands shown for Debian/Ubuntu; adapt for other systems)

Understanding Proxy Rotation in Scrapy

Scrapy uses downloader middlewares to process requests and responses. To rotate proxies, you intercept each request, assign a proxy from a pool, and optionally retry with a new proxy when a request fails.

Key concepts:

  • Proxy pool: a list of proxy URLs (with credentials if needed).
  • Rotation strategy: random selection, round-robin, or weighted by proxy health.
  • Retry logic: re-queue failed requests with a different proxy.

When to consider a dedicated WireGuard gateway instead:

  • You need a consistent IP for SEO rank tracking across multiple locations.
  • Your target site blocks frequent IP changes more aggressively than static IPs.
  • You want encrypted, team-managed access without maintaining a proxy list.

Steps

1. Install Scrapy and create a project

If you haven't already, install Scrapy and scaffold a new project.

pip install scrapy
scrapy startproject proxy_scraper
cd proxy_scraper
scrapy genspider example example.com

2. Create a rotating proxy middleware

Open proxy_scraper/middlewares.py and add the following middleware. It selects a random proxy for each request and retries on common block status codes.

# proxy_scraper/middlewares.py
import random
from scrapy import signals
from scrapy.exceptions import NotConfigured

class RotatingProxyMiddleware:
    def __init__(self, proxy_list):
        self.proxy_list = proxy_list
        self.current_proxy = None

    @classmethod
    def from_crawler(cls, crawler):
        proxy_list = crawler.settings.getlist('ROTATING_PROXY_LIST')
        if not proxy_list:
            raise NotConfigured('ROTATING_PROXY_LIST is empty')
        return cls(proxy_list)

    def process_request(self, request, spider):
        self.current_proxy = random.choice(self.proxy_list)
        request.meta['proxy'] = self.current_proxy
        spider.logger.debug(f'Using proxy: {self.current_proxy}')

    def process_response(self, request, response, spider):
        if response.status in [403, 429]:
            spider.logger.warning(f'Blocked with proxy {self.current_proxy}. Retrying...')
            new_request = request.copy()
            new_request.dont_filter = True
            return new_request
        return response

3. Configure settings to enable the middleware and define proxies

Edit proxy_scraper/settings.py to activate the middleware and supply your proxy list. Replace the example proxies with your own.

# proxy_scraper/settings.py

DOWNLOADER_MIDDLEWARES = {
    'proxy_scraper.middlewares.RotatingProxyMiddleware': 543,
}

ROTATING_PROXY_LIST = [
    'http://user:[email protected]:8080',
    'http://user:[email protected]:8080',
    'socks5://user:[email protected]:1080',
]

# Optional: retry settings
RETRY_TIMES = 5
RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408, 429, 403]

4. Test rotation with a simple spider

Run your spider and watch the DEBUG logs to confirm different proxies are used.

scrapy crawl example -L DEBUG

If you see lines like Using proxy: http://..., the middleware is working. If all requests use the same proxy, check the middleware order and that ROTATING_PROXY_LIST is not empty.

5. Improve reliability with health checks and backoff

For production scraping, add basic proxy health tracking. This example marks a proxy as failed after repeated blocks and temporarily removes it.

# proxy_scraper/middlewares.py (enhanced)
import random
from collections import defaultdict

class RotatingProxyMiddleware:
    def __init__(self, proxy_list):
        self.proxy_list = proxy_list
        self.failure_count = defaultdict(int)
        self.max_failures = 3

    @classmethod
    def from_crawler(cls, crawler):
        proxy_list = crawler.settings.getlist('ROTATING_PROXY_LIST')
        return cls(proxy_list)

    def process_request(self, request, spider):
        available = [p for p in self.proxy_list if self.failure_count[p] < self.max_failures]
        if not available:
            spider.logger.error('All proxies exhausted')
            return
        proxy = random.choice(available)
        request.meta['proxy'] = proxy

    def process_response(self, request, response, spider):
        proxy = request.meta.get('proxy')
        if response.status in [403, 429]:
            self.failure_count[proxy] += 1
            new_request = request.copy()
            new_request.dont_filter = True
            return new_request
        else:
            self.failure_count[proxy] = 0
        return response

6. When to switch to a dedicated WireGuard gateway

Rotating proxies are excellent for large-scale, stateless scraping. But some tasks need a stable, dedicated IP:

  • SEO rank tracking where you need consistent location data
  • Accessing APIs that tie sessions to an IP
  • Team-based scraping where multiple users share the same egress IP

NordLayer WireGuard tunnels give you a dedicated gateway with flat monthly billing. You can route all Scrapy traffic through the tunnel, effectively using a single static IP.

7. Set up WireGuard on Linux and route Scrapy through it

Install WireGuard and configure a tunnel using the credentials from your NordLayer gateway.

sudo apt update
sudo apt install wireguard

Create a configuration file at /etc/wireguard/wg0.conf (replace placeholders with your NordLayer values).

[Interface]
PrivateKey = <your-private-key>
Address = 10.0.0.2/32
DNS = 1.1.1.1

[Peer]
PublicKey = <nordlayer-public-key>
Endpoint = <gateway-ip>:51820
AllowedIPs = 0.0.0.0/0
PersistentKeepalive = 25

Bring the tunnel up:

sudo wg-quick up wg0

Now all Scrapy traffic (and any other traffic) exits through the NordLayer gateway. Verify your public IP has changed:

curl ifconfig.me

Run your Scrapy spider as usual. It will use the stable gateway IP.

8. Combine approaches for hybrid scraping

You can use both: route most requests through rotating proxies, but send requests that need a consistent IP through the WireGuard tunnel. In Scrapy, you can conditionally set the proxy to None (which uses the system's default route, i.e., the WireGuard tunnel) for specific requests.

# Example: in a spider, for rank-tracking requests, don't set a proxy
# The default route (WireGuard) will be used.
def start_requests(self):
    # Rotating proxies for general pages
    yield scrapy.Request('https://example.com/general', meta={'proxy': None})
    # For rank tracking, bypass the rotating middleware
    yield scrapy.Request('https://example.com/serp', meta={'proxy': 'direct'})

Then modify the middleware to skip rotation when proxy == 'direct':

def process_request(self, request, spider):
    if request.meta.get('proxy') == 'direct':
        return None  # use default route
    # ... rest of rotation logic

Troubleshooting

  • Middleware not applied: Ensure DOWNLOADER_MIDDLEWARES includes the correct path and that the value (543) doesn't conflict with other middlewares. Check Scrapy's default middleware order.
  • All requests still use the same IP: Verify that ROTATING_PROXY_LIST is populated and that the middleware's process_request is running. Add a print or a debug log.
  • Frequent 403/429 errors: Your proxies may be blocked or too slow. Reduce CONCURRENT_REQUESTS and DOWNLOAD_DELAY, or switch to higher-quality proxies. Consider a dedicated WireGuard gateway if the target site blocks rotating IPs aggressively.
  • WireGuard tunnel fails to connect: Double-check the private key, public key, endpoint, and firewall rules (UDP port 51820 by default). Run sudo wg show to see handshake status.
  • Scrapy ignores the WireGuard tunnel: If you set AllowedIPs = 0.0.0.0/0, all traffic should route through the tunnel. Confirm with ip route and test with curl ifconfig.me.

Summary

You now know how to implement proxy rotation in Scrapy using custom middleware, add retry logic, and monitor proxy health. For projects that require a stable, dedicated IP—such as SEO rank tracking or team-based access—a WireGuard gateway like NordLayer can complement or replace rotating proxies. Choose the approach that matches your scraping goals: rotation for scale and anonymity, dedicated gateways for consistency and control.