ProxyScrape Python Tutorial: Build an Async Rotating Proxy Pool with httpx and SOCKS5
An advanced Python pattern for ProxyScrape: model your HTTP and SOCKS5 endpoints as a pool, rotate across them under a concurrency limit, back off on failures, and track per-endpoint health.
Overview
This is an advanced Python tutorial. Instead of sending one request through one proxy, you build a small fetcher that rotates across the ProxyScrape endpoints in your plan, caps concurrency, retries with jittered backoff, and parks failing endpoints in a cooldown.
You will:
- Install
httpxwith SOCKS5 support. - Keep ProxyScrape credentials in environment variables.
- Model your HTTP and SOCKS5 endpoints as a pool.
- Rotate, throttle, retry, and report per-endpoint health.
- Smoke test before scaling up.
ProxyScrape offers residential, datacenter, and mobile proxies, supports HTTP and SOCKS5, and bills either pay-as-you-go or flat monthly. Which you choose changes cost and latency, not the client code — the code below treats endpoints as configuration.
Choose a proxy type before you write code
| Workload | ProxyScrape type | Why it fits |
|---|---|---|
| High-volume, low-friction targets (docs, sitemaps, public APIs) | Datacenter | Fastest option; stable long-term uptime suits scheduled jobs |
| Bot-filtered pages or geo-specific content | Residential | Large pool (97M+ IPs) with real-user origins |
| Login flows and aggressive anti-bot targets | Mobile | Highest trust; use low concurrency |
Prerequisites
- Python 3.10 or newer.
- A ProxyScrape plan with credentials, plus the host and port for
httpand/orsocks5. - Outbound access from your machine to the ProxyScrape host and port.
Steps
Step 1 — Create a project and install dependencies
mkdir proxyscrape-async
cd proxyscrape-async
python -m venv .venv
source .venv/bin/activate
pip install "httpx[socks]"
The [socks] extra installs socksio. Without it, httpx raises an error the first time you configure a socks5:// endpoint. httpx 0.26 and later use the proxy= argument (older releases used proxies=), so check your version with pip show httpx if the examples do not match.
Step 2 — Keep credentials and endpoints in environment variables
# pool.py
import os
from dataclasses import dataclass
from urllib.parse import quote
@dataclass(frozen=True)
class ProxyEndpoint:
label: str
url: str
def endpoint_url(scheme: str, host: str, port: str, username: str, password: str) -> str:
user = quote(username, safe="")
secret = quote(password, safe="")
return f"{scheme}://{user}:{secret}@{host}:{port}"
HOST = os.environ["PROXYSCRAPE_HOST"]
USERNAME = os.environ["PROXYSCRAPE_USERNAME"]
PASSWORD = os.environ["PROXYSCRAPE_PASSWORD"]
POOL: list[ProxyEndpoint] = []
if http_port := os.environ.get("PROXYSCRAPE_HTTP_PORT"):
POOL.append(ProxyEndpoint("http-1", endpoint_url("http", HOST, http_port, USERNAME, PASSWORD)))
if socks_port := os.environ.get("PROXYSCRAPE_SOCKS5_PORT"):
POOL.append(ProxyEndpoint("socks5-1", endpoint_url("socks5", HOST, socks_port, USERNAME, PASSWORD)))
if not POOL:
raise SystemExit("Set PROXYSCRAPE_HTTP_PORT and/or PROXYSCRAPE_SOCKS5_PORT")
Export the values in your shell, or inject them from a secrets manager in CI:
export PROXYSCRAPE_HOST="host-from-your-plan"
export PROXYSCRAPE_USERNAME="your-username"
export PROXYSCRAPE_PASSWORD="your-password"
export PROXYSCRAPE_HTTP_PORT="port-from-your-plan"
export PROXYSCRAPE_SOCKS5_PORT="port-from-your-plan"
Percent-encoding via quote() matters: passwords containing @, :, or / otherwise break URL parsing. Never commit these values.
Step 3 — Verify every endpoint before scaling
Run this checklist in order. If a step fails, fix it before touching the async code.
- Check the HTTP/HTTPS endpoint.
- Check the SOCKS5 endpoint.
- Confirm the returned IP is not your own IP.
curl -sS -x "http://$PROXYSCRAPE_USERNAME:$PROXYSCRAPE_PASSWORD@$PROXYSCRAPE_HOST:$PROXYSCRAPE_HTTP_PORT" "https://api.ipify.org?format=json"
curl -sS -x "socks5h://$PROXYSCRAPE_USERNAME:$PROXYSCRAPE_PASSWORD@$PROXYSCRAPE_HOST:$PROXYSCRAPE_SOCKS5_PORT" "https://api.ipify.org?format=json"
socks5h:// tells curl to let the proxy resolve DNS. If your Python client offers both socks5:// and socks5h://, prefer the h variant when local DNS resolution fails or when you want DNS to leave your machine together with the request.
Step 4 — Build the rotating, concurrent fetcher
# fetch_async.py
from __future__ import annotations
import asyncio
import itertools
import random
import time
import httpx
from pool import POOL, ProxyEndpoint
class RotatingFetcher:
"""Round-robin fetcher with concurrency limits, backoff, and cooldowns."""
def __init__(
self,
pool: list[ProxyEndpoint],
concurrency: int = 8,
max_attempts: int = 4,
timeout: float = 20.0,
cooldown_seconds: float = 45.0,
) -> None:
if not pool:
raise ValueError("Proxy pool is empty")
self._pool = pool
self._cycle = itertools.cycle(pool)
self._semaphore = asyncio.Semaphore(concurrency)
self._max_attempts = max_attempts
self._cooldown_seconds = cooldown_seconds
self._failures: dict[str, int] = {}
self._cooldown_until: dict[str, float] = {}
self._clients = {
item.label: httpx.AsyncClient(
proxy=item.url,
timeout=timeout,
follow_redirects=True,
headers={"User-Agent": "Mozilla/5.0 (compatible; ProxyScoutTutorial/1.0)"},
)
for item in pool
}
def _pick(self) -> ProxyEndpoint:
now = time.monotonic()
for _ in range(len(self._pool) * 2):
candidate = next(self._cycle)
if self._cooldown_until.get(candidate.label, 0.0) <= now:
return candidate
return min(self._pool, key=lambda item: self._cooldown_until.get(item.label, 0.0))
def _record_failure(self, label: str) -> None:
self._failures[label] = self._failures.get(label, 0) + 1
penalty = self._cooldown_seconds * self._failures[label]
self._cooldown_until[label] = time.monotonic() + penalty
def _record_success(self, label: str) -> None:
self._failures[label] = 0
self._cooldown_until[label] = 0.0
@staticmethod
def _backoff(attempt: int) -> float:
return min((2 ** attempt) + random.uniform(0, 0.5), 30.0)
async def fetch(self, url: str) -> httpx.Response:
last_error: Exception | None = None
for attempt in range(1, self._max_attempts + 1):
item = self._pick()
try:
async with self._semaphore:
response = await self._clients[item.label].get(url)
if response.status_code in {407, 429} or response.status_code >= 500:
raise httpx.HTTPStatusError(
f"retryable status {response.status_code}",
request=response.request,
response=response,
)
self._record_success(item.label)
return response
except Exception as exc:
last_error = exc
self._record_failure(item.label)
if attempt < self._max_attempts:
await asyncio.sleep(self._backoff(attempt))
raise RuntimeError(f"All {self._max_attempts} attempts failed for {url}") from last_error
def report(self) -> str:
active = {label: count for label, count in self._failures.items() if count}
if not active:
return "no endpoints have failed yet"
return ", ".join(f"{label}={count}" for label, count in sorted(active.items()))
async def aclose(self) -> None:
await asyncio.gather(*(client.aclose() for client in self._clients.values()))
Design notes worth keeping:
- One
AsyncClientper endpoint. In httpx the proxy is bound to the client, not to an individual request, so a client per endpoint is the simplest correct rotation model. - The semaphore caps in-flight requests no matter how many URLs you submit.
407and429join5xxas retryable statuses, and they count as endpoint failures so a bad endpoint stops receiving traffic for a while.- Failure penalties grow linearly and reset on success, which prevents a flaky endpoint from being retried in a tight loop.
Step 5 — Smoke test and tune concurrency
# run_smoke_test.py
import asyncio
import httpx
from fetch_async import RotatingFetcher
from pool import POOL
async def main() -> None:
fetcher = RotatingFetcher(pool=POOL, concurrency=8)
urls = [f"https://example.com/?page={index}" for index in range(1, 21)]
try:
results = await asyncio.gather(*(fetcher.fetch(url) for url in urls), return_exceptions=True)
ok = sum(1 for item in results if isinstance(item, httpx.Response))
print(f"success={ok} failed={len(results) - ok}")
print(f"endpoint health: {fetcher.report()}")
finally:
await fetcher.aclose()
if __name__ == "__main__":
asyncio.run(main())
Then tune:
- Run it once. If
failedis 0, raiseconcurrencyby 4 and run again. - Stop increasing when latency or timeouts start climbing — that is your practical ceiling for the plan you bought.
- If failures cluster on a single endpoint, treat that endpoint as unhealthy and contact support instead of retrying harder.
Plan note: with pay-as-you-go billing you can push concurrency up for a burst; a flat monthly plan usually suits steady, scheduled jobs better.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
Import error mentioning socksio |
httpx[socks] was not installed |
pip install "httpx[socks]" |
| Every request fails with 407 | Wrong credentials, or a password with special characters left unencoded | Re-check the environment variables and keep the quote() calls |
| Timeouts climb as concurrency rises | Too many in-flight requests for the plan or the target | Lower concurrency, raise timeout |
| Requests keep landing in cooldown | Penalties compounding after repeated failures | Verify endpoints with curl first, and consider capping cooldown_seconds |
The socks5 endpoint fails while http works |
Local DNS resolution problem | Try the socks5h scheme if your library supports it |
| The target returns 403 consistently | That endpoint IP is filtered for this target | Rotate across more endpoints, or move to residential or mobile proxies |
Summary
- ProxyScrape supports HTTP and SOCKS5; represent each endpoint from your plan as a
ProxyEndpointand rotate across them. - Bind a proxy to an
httpx.AsyncClient, cap in-flight work with a semaphore, and retry with jittered backoff. - Treat 407, 429, and 5xx responses as endpoint failures rather than user errors, and cool the endpoint down.
- Verify endpoints with curl before scaling, then raise concurrency gradually until latency — not luck — becomes your limit.