Advanced ProxyScrape Rotation in Scrapy: Custom Downloader Middleware with Retry and Ban Detection
Build a custom Scrapy downloader middleware that rotates ProxyScrape residential and datacenter proxies, handles retries, and detects bans for reliable large-scale scraping.
Overview
Scrapy is a powerful Python framework for large-scale web scraping, but without proxy rotation you will quickly hit rate limits and IP bans. This tutorial shows you how to build a custom Scrapy downloader middleware that rotates ProxyScrape residential and datacenter proxies, handles retries automatically, and detects bans for reliable data collection.
ProxyScrape provides residential, datacenter, and mobile proxies with HTTP and SOCKS5 support, a pool of 97M+ IPs, and strong long-term uptime. You can choose pay-as-you-go or flat monthly billing depending on your scraping volume.
Prerequisites
- Python 3.8 or newer installed.
- Scrapy installed (
pip install scrapy). - A ProxyScrape account with residential or datacenter proxies enabled.
- Your proxy endpoint and credentials from the ProxyScrape dashboard.
- Optional: A list of individual proxy endpoints if you prefer static rotation over a rotating endpoint.
Step 1: Create a Scrapy Project
Open your terminal and create a new Scrapy project:
pip install scrapy
scrapy startproject proxyscrape_scraper
cd proxyscrape_scraper
Step 2: Store ProxyScrape Credentials Securely
Never hard-code credentials in your source files. Use environment variables and read them in settings.py.
Create a .env file or export variables in your shell:
export PROXYSCRAPE_USER='your_username'
export PROXYSCRAPE_PASS='your_password'
export PROXYSCRAPE_ENDPOINT='http://proxy.proxyscrape.com:8080'
Then in settings.py, add:
import os
PROXYSCRAPE_USER = os.getenv('PROXYSCRAPE_USER')
PROXYSCRAPE_PASS = os.getenv('PROXYSCRAPE_PASS')
PROXYSCRAPE_ENDPOINT = os.getenv('PROXYSCRAPE_ENDPOINT')
Step 3: Build the Rotation Middleware
Create a new file proxyscrape_scraper/middlewares.py and define a middleware that:
- Picks a proxy from a list or uses your rotating endpoint.
- Injects credentials into the proxy URL.
- Retries on ban-related HTTP status codes.
# proxyscrape_scraper/middlewares.py
import random
import logging
from scrapy import signals
from scrapy.exceptions import NotConfigured
class ProxyScrapeRotationMiddleware:
def __init__(self, proxy_list, credentials):
self.proxy_list = proxy_list
self.credentials = credentials
self.logger = logging.getLogger(__name__)
@classmethod
def from_crawler(cls, crawler):
proxy_list = crawler.settings.getlist('PROXY_LIST')
credentials = (
crawler.settings.get('PROXYSCRAPE_USER'),
crawler.settings.get('PROXYSCRAPE_PASS'),
)
if not proxy_list:
raise NotConfigured('PROXY_LIST is empty')
return cls(proxy_list, credentials)
def process_request(self, request, spider):
proxy = random.choice(self.proxy_list)
user, password = self.credentials
if user and password:
proxy = proxy.replace('://', f'://{user}:{password}@')
request.meta['proxy'] = proxy
self.logger.debug(f'Using proxy: {proxy}')
def process_response(self, request, response, spider):
if response.status in [403, 429, 503]:
self.logger.warning(
f'Ban detected with status {response.status} on {request.url}'
)
request.meta['proxy'] = None
return request.replace(dont_filter=True)
return response
Step 4: Configure Settings for Rotation and Retries
In settings.py, define your proxy list, enable the middleware, and tune retry behaviour.
# settings.py
PROXY_LIST = [
'http://proxy1.proxyscrape.com:8080',
'socks5://proxy2.proxyscrape.com:1080',
# Add more endpoints from your ProxyScrape dashboard
]
DOWNLOADER_MIDDLEWARES = {
'proxyscrape_scraper.middlewares.ProxyScrapeRotationMiddleware': 543,
'scrapy.downloadermiddlewares.retry.RetryMiddleware': 550,
}
RETRY_ENABLED = True
RETRY_TIMES = 3
RETRY_HTTP_CODES = [403, 429, 500, 502, 503, 504]
Step 5: Optional — Sticky Sessions with Session IDs
If your ProxyScrape plan supports sticky sessions via username parameters, you can append a session ID to the username. Check your ProxyScrape dashboard for the exact syntax (often something like username-session-abc123).
# Inside process_request, replace the username logic:
session_id = ''.join(random.choices('abcdef0123456789', k=8))
user = f'{self.credentials[0]}-session-{session_id}'
proxy = proxy.replace('://', f'://{user}:{self.credentials[1]}@')
Use sticky sessions when you need to maintain the same IP across multiple requests to the same domain, such as during login flows or multi-step checkouts.
Step 6: Test Your Rotating Scraper
Create a simple spider to verify that requests are going through different proxies.
# proxyscrape_scraper/spiders/ip_test.py
import scrapy
class IpTestSpider(scrapy.Spider):
name = 'ip_test'
start_urls = ['https://httpbin.org/ip']
def parse(self, response):
yield {'ip': response.json()['origin']}
Run the spider and save the results:
scrapy crawl ip_test -o ips.json
Open ips.json and confirm that the ip values are different. If you see the same IP repeatedly, check your PROXY_LIST or rotating endpoint configuration.
Choosing the Right ProxyScrape Proxy Type for Scrapy
| Proxy Type | Best For | Rotation | Sticky Sessions | Protocol Support |
|---|---|---|---|---|
| Residential | General web scraping, geo-targeted data | Rotating endpoint or list | Typically available via session IDs | HTTP, SOCKS5 |
| Datacenter | High-speed scraping of non-protected sites | Rotating endpoint or list | Often static | HTTP, SOCKS5 |
| Mobile | Mobile-only sites, social media, app testing | Rotating endpoint | Usually available | HTTP, SOCKS5 |
Troubleshooting
- Authentication errors (407 Proxy Authentication Required): Double-check your username and password. If using a rotating endpoint, ensure credentials are embedded correctly as
user:pass@host:port. - Connection timeouts: Try switching from HTTP to SOCKS5 or vice versa. Some sites block one protocol more aggressively. Also verify that your firewall allows outbound traffic on the proxy port.
- Bans not detected: If you still receive CAPTCHAs or empty data, extend your ban detection to look for CAPTCHA indicators in the response body (for example, the presence of
captchain the HTML). You can also reduce your request rate and increaseDOWNLOAD_DELAY. - SOCKS5 errors: Scrapy requires the
requestslibrary with SOCKS support for SOCKS5 proxies. Install it withpip install requests[socks]if you seeMissing dependencies for SOCKS support. - Sticky sessions not working: Confirm the exact session parameter format in your ProxyScrape dashboard. Some plans use
username-session-{id}while others use a separate password field.
Summary
You now have a production-ready Scrapy middleware that rotates ProxyScrape proxies, retries on bans, and optionally maintains sticky sessions. This setup lets you scale your scraping projects without worrying about rate limits or IP blocks. For best results, combine rotating residential proxies with a conservative download delay and monitor your success rates. Adjust your PROXY_LIST and retry logic as your scraping targets change.