Skip to content
intermediate

How to Use ProxyScrape Residential Proxies for Web Scraping in Python

A step-by-step guide to integrating ProxyScrape's residential and datacenter proxies into your Python web scraping scripts using HTTP and SOCKS5 protocols.

SOCKS5 HTTP(S) Web Scraping

Overview

ProxyScrape provides residential, datacenter, and mobile proxies with HTTP and SOCKS5 support, backed by a pool of 97M+ IPs. This tutorial shows how to use ProxyScrape proxies in Python for web scraping tasks. You'll learn to configure requests with both HTTP and SOCKS5, handle authentication, and rotate proxies. The steps apply to any proxy provider, but we use ProxyScrape as an example due to its reliable uptime and long-term stability.

Steps

  1. Obtain proxy credentials from ProxyScrape. Log in to your ProxyScrape dashboard and navigate to the proxy list or API section. You will find the proxy endpoint (IP:port), username, and password. ProxyScrape offers pay-as-you-go and flat monthly billing plans, so choose the one that fits your scraping volume.

  2. Install required Python libraries. Open your terminal and install requests and requests[socks] for SOCKS5 support:

    pip install requests requests[socks]
    
  3. Write a basic scraping script using HTTP proxies. Create a Python file and use the following code to send requests through ProxyScrape's HTTP proxy. Replace the placeholders with your actual credentials.

    import requests
    
    proxy_url = "http://username:password@proxy-endpoint:port"
    proxies = {
        "http": proxy_url,
        "https": proxy_url,
    }
    
    response = requests.get("http://httpbin.org/ip", proxies=proxies, timeout=10)
    print(response.json())
    
  4. Use SOCKS5 proxies for better anonymity. ProxyScrape supports SOCKS5. Modify the proxy URL scheme to socks5://:

    import requests
    
    proxy_url = "socks5://username:password@proxy-endpoint:port"
    proxies = {
        "http": proxy_url,
        "https": proxy_url,
    }
    
    response = requests.get("http://httpbin.org/ip", proxies=proxies, timeout=10)
    print(response.json())
    
  5. Rotate proxies to avoid rate limits. If your ProxyScrape plan provides a rotating endpoint, use that endpoint for each request. Otherwise, maintain a list of proxy endpoints from your dashboard and cycle through them. Example:

    import requests
    import random
    
    proxy_list = [
        "http://user:pass@proxy1:port",
        "http://user:pass@proxy2:port",
        # ... add more
    ]
    
    def get_random_proxy():
        return random.choice(proxy_list)
    
    for _ in range(10):
        proxy_url = get_random_proxy()
        proxies = {"http": proxy_url, "https": proxy_url}
        try:
            response = requests.get("http://httpbin.org/ip", proxies=proxies, timeout=10)
            print(response.json())
        except Exception as e:
            print(f"Proxy failed: {e}")
    
  6. Handle errors and retries. Proxy connections can fail. Wrap your requests in try/except blocks and implement a retry mechanism with a new proxy on failure. This ensures your scraper remains robust.

  7. Test and scale. Run your script and verify that the IP changes with each request. For large-scale scraping, consider using a session object and adjusting concurrency based on your ProxyScrape plan limits.

Note: Always respect website terms of service and robots.txt. ProxyScrape's residential and datacenter proxies are reliable, but ethical scraping practices are your responsibility.