How to Use ProxyScrape Residential Proxies with Scrapy in Python
Integrate ProxyScrape residential proxies into your Scrapy spiders for reliable Python web scraping. Includes rotation middleware, authentication, and SOCKS5 setup.
Overview
Scrapy is a powerful Python framework for web scraping, but many sites block requests from a single IP. Using proxies from ProxyScrape helps you distribute requests, avoid rate limits, and access geo-restricted content. ProxyScrape offers residential, datacenter, and mobile proxies with HTTP and SOCKS5 support, pay-as-you-go or flat monthly billing, and a pool of 97M+ IPs built for long-term uptime.
This tutorial shows you how to integrate ProxyScrape proxies into a Scrapy project. You will learn to configure authentication, rotate proxies per request, and handle common errors. The approach works on Windows, macOS, and Linux.
Prerequisites
- Python 3.8 or newer installed.
- Scrapy installed (
pip install scrapy). - A ProxyScrape account with residential or datacenter proxies. You can sign up for a plan or use pay-as-you-go.
- Your proxy credentials: host, port, username, and password from the ProxyScrape dashboard.
Understanding Proxy Authentication in Scrapy
ProxyScrape uses username/password authentication. In Scrapy, you pass the proxy URL in the proxy request meta key. The format is:
- HTTP/HTTPS:
http://username:password@host:port - SOCKS5:
socks5://username:password@host:port
Scrapy natively supports HTTP and HTTPS proxies. For SOCKS5, you need an additional package like scrapy-socks.
Step 1: Install Scrapy and Create a Project
Open a terminal and run:
pip install scrapy
scrapy startproject proxyscrape_scrapy
cd proxyscrape_scrapy
This creates a new Scrapy project with a default spider directory.
Step 2: Get Your Proxy Credentials from ProxyScrape
Log in to your ProxyScrape dashboard and create a proxy list. You will receive a host, port, username, and password. For residential proxies, you may get a gateway host that rotates automatically, or a list of endpoints. For datacenter proxies, you typically get a fixed IP and port.
Example credentials:
- Host:
proxy.proxyscrape.com - Port:
8080 - Username:
user123 - Password:
pass456
Your proxy URL becomes: http://user123:[email protected]:8080
If your password contains special characters, URL-encode them (e.g., @ becomes %40).
Step 3: Configure a Basic Proxy in a Spider
Create a simple spider that checks your IP address through the proxy. In proxyscrape_scrapy/spiders/example.py:
import scrapy
class ExampleSpider(scrapy.Spider):
name = 'example'
start_urls = ['https://httpbin.org/ip']
def start_requests(self):
proxy = 'http://user123:[email protected]:8080'
for url in self.start_urls:
yield scrapy.Request(
url,
meta={'proxy': proxy},
callback=self.parse
)
def parse(self, response):
yield {'ip': response.text}
Run the spider:
scrapy crawl example -O output.json
The output shows the IP address as seen by the target site. If it matches your proxy IP, the setup works.
Step 4: Implement a Rotating Proxy Middleware
Rotating proxies per request improves anonymity and reduces the chance of blocks. Create a middleware in proxyscrape_scrapy/middlewares.py:
import random
from scrapy import signals
class ProxyRotationMiddleware:
def __init__(self, proxies):
self.proxies = proxies
@classmethod
def from_crawler(cls, crawler):
return cls(crawler.settings.getlist('PROXY_LIST'))
def process_request(self, request, spider):
proxy = random.choice(self.proxies)
request.meta['proxy'] = proxy
spider.logger.debug(f'Using proxy: {proxy}')
Enable the middleware and define your proxy list in settings.py:
DOWNLOADER_MIDDLEWARES = {
'proxyscrape_scrapy.middlewares.ProxyRotationMiddleware': 543,
}
PROXY_LIST = [
'http://user1:pass1@host1:port1',
'http://user2:pass2@host2:port2',
'http://user3:pass3@host3:port3',
]
Now every request uses a random proxy from the list. You can populate the list with multiple ProxyScrape endpoints or use their rotating gateway (one endpoint that rotates automatically).
Step 5: Handle SOCKS5 Proxies (Optional)
Scrapy does not support SOCKS5 natively. Install scrapy-socks:
pip install scrapy-socks
Add the SOCKS5 middleware to settings.py:
DOWNLOADER_MIDDLEWARES = {
'scrapy_socks.Socks5DownloaderMiddleware': 543,
}
Then set the proxy in request meta as socks5://user:pass@host:port. Note that you cannot use both the rotation middleware and scrapy-socks simultaneously without custom logic. For most use cases, HTTP proxies are simpler and sufficient.
Step 6: Test and Verify Rotation
Run your spider and check the output:
scrapy crawl example -O output.json
Inspect output.json. If you see different IP addresses across requests, rotation is working. You can also add a delay between requests to be polite:
DOWNLOAD_DELAY = 2
Troubleshooting
- 407 Proxy Authentication Required: Check your username and password. Ensure they are URL-encoded. Verify that your ProxyScrape plan is active.
- Connection refused or timeout: Confirm the host and port. Check your firewall or VPN. Try a different proxy endpoint.
- Frequent blocks: Switch to residential proxies, increase
DOWNLOAD_DELAY, and enable rotation. Datacenter proxies are more likely to be blocked on strict sites. - SOCKS5 errors: Ensure
scrapy-socksis installed and configured. If you get import errors, reinstall the package. - SSL errors: Add
DOWNLOAD_HANDLERSor usehttpsin the proxy URL if the target site requires it.
Choosing the Right Proxy Type for Scrapy
| Proxy Type | Best For | Notes |
|---|---|---|
| Residential | Scraping sites with anti-bot protection | High anonymity, large pool, slower |
| Datacenter | High-speed scraping of less protected sites | Faster, cheaper, easier to block |
| Mobile | Scraping mobile-specific content or apps | Rarely blocked, higher cost |
ProxyScrape provides all three types with flexible billing. Start with residential for most scraping tasks.
Summary
You have learned how to integrate ProxyScrape proxies into Scrapy: install Scrapy, obtain credentials, configure a basic proxy, build a rotation middleware, and optionally use SOCKS5. Test your setup and adjust rotation and delays to avoid blocks. ProxyScrape’s reliable residential and datacenter proxies with 97M+ IPs and long-term uptime make them a solid choice for Python scraping projects.