ProxyScrape Web Scraping with Playwright: Route Headless Chromium Through HTTP and SOCKS5 Proxies
Configure Playwright so headless Chromium exits through ProxyScrape, verify the exit IP before scraping, rotate HTTP or SOCKS5 endpoints per browser context, and classify proxy errors correctly.
Overview
Headless browsers put the proxy configuration inside the browser, not inside your HTTP library. That changes the setup: you configure the proxy when you launch a browser or create a browser context, you have to handle proxy authentication, and per-context configuration is what makes rotation possible without restarting Chromium for every request.
This tutorial uses Playwright with Chromium and ProxyScrape endpoints (HTTP or SOCKS5). It covers a single-endpoint launch, an exit-IP check, per-context rotation across several endpoints, proxy error handling, and running it on Linux CI without exhausting memory.
When a headless browser is the right tool
- The target renders content client-side, so a plain HTTP request returns an empty shell.
- You need to interact: fill forms, click, scroll, or wait for network idle.
- You need browser-level behaviour such as cookies,
localStorage, or service workers. - You are validating how a site behaves for a specific region.
Stick with a plain HTTP client when the page is server-rendered; it is far cheaper and faster.
Prerequisites
- Node.js 18 or newer and npm.
- Playwright installed with Chromium.
- ProxyScrape credentials plus the host and port for
httpand/orsocks5. Residential, datacenter, and mobile proxies all work here; pick the type that matches the target's defences. - Linux, macOS, or Windows. The CI notes assume Debian or Ubuntu runners.
Steps
Step 1 — Install Playwright and Chromium
mkdir proxyscrape-playwright
cd proxyscrape-playwright
npm init -y
npm install playwright
npx playwright install --with-deps chromium
--with-deps installs the OS libraries Chromium needs on Linux images and CI runners.
Step 2 — Decide whether the proxy applies to the browser or the context
| Scope | Where it is set | Use it for | Caveats |
|---|---|---|---|
| Browser | chromium.launch({ proxy }) |
A single endpoint for the whole run | Every context shares one exit IP |
| Context | browser.newContext({ proxy }) |
Rotation: one endpoint per context, in parallel | Engine support differs; Chromium is the safe default — confirm per-context proxy support for your Playwright version before relying on it in Firefox or WebKit |
Two practical rules:
- Pass credentials as separate
usernameandpasswordfields. Chromium ignores credentials embedded in the proxy server URL. - The
servervalue takes a scheme:http://host:portfor ProxyScrape's HTTP endpoints andsocks5://host:portfor SOCKS5.
Step 3 — Launch through one endpoint and check the exit IP
// launch-proxy.mjs
import { chromium } from 'playwright';
const proxy = {
server: process.env.PROXYSCRAPE_SERVER, // http://HOST:PORT or socks5://HOST:PORT
username: process.env.PROXYSCRAPE_USERNAME,
password: process.env.PROXYSCRAPE_PASSWORD,
};
const browser = await chromium.launch({ proxy, headless: true });
try {
const context = await browser.newContext();
const page = await context.newPage();
await page.goto('https://api.ipify.org?format=json', { waitUntil: 'domcontentloaded' });
console.log('exit ip:', await page.textContent('body'));
} finally {
await browser.close();
}
export PROXYSCRAPE_SERVER="http://host-from-your-plan:port-from-your-plan"
export PROXYSCRAPE_USERNAME="your-username"
export PROXYSCRAPE_PASSWORD="your-password"
node launch-proxy.mjs
If this prints an IP, the proxy path works end to end. If it hangs or throws, jump to the troubleshooting table before writing more code.
Step 4 — Fail fast when traffic is not proxied
A silent misconfiguration — for example, a typo in the server string — can leave you scraping from your own IP. Capture your direct IP once, then assert against it.
// exit-ip.mjs
export async function getExitIp(page) {
await page.goto('https://api.ipify.org?format=json', { waitUntil: 'domcontentloaded' });
const body = await page.textContent('body');
return JSON.parse(body).ip;
}
export function assertProxied(exitIp, directIp) {
if (directIp && exitIp === directIp) {
throw new Error(`Traffic was not proxied: exit IP ${exitIp} equals the direct IP`);
}
}
Ordered procedure:
- Run a one-off script with no proxy configured and record the result as
DIRECT_IP. - Run the proxied script again.
- Call
assertProxied(exitIp, process.env.DIRECT_IP)before any real scraping, and fail the run when the assertion trips.
Step 5 — Rotate endpoints across browser contexts
One browser, many contexts, one endpoint each.
// rotate-contexts.mjs
import { chromium } from 'playwright';
const ENDPOINTS = (process.env.PROXYSCRAPE_ENDPOINTS ?? '')
.split(',')
.map((entry) => entry.trim())
.filter(Boolean)
.map((entry) => {
const [scheme, host, port, username, password] = entry.split('|');
return { server: `${scheme}://${host}:${port}`, username, password };
});
if (ENDPOINTS.length === 0) {
throw new Error('Set PROXYSCRAPE_ENDPOINTS to http|HOST|PORT|USER|PASS,socks5|HOST|PORT|USER|PASS');
}
const TARGETS = (process.env.TARGETS ?? 'https://example.com,https://example.org').split(',');
const CONCURRENCY = Number(process.env.CONCURRENCY ?? 3);
const USER_AGENT =
'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/124.0 Safari/537.36';
function isProxyError(error) {
const message = String(error && error.message ? error.message : '');
return /ERR_PROXY|ERR_TUNNEL|ERR_PROXY_AUTH/i.test(message);
}
async function runBatch(browser, proxy, urls) {
const context = await browser.newContext({ proxy, userAgent: USER_AGENT });
context.setDefaultTimeout(30000);
const results = [];
try {
const page = await context.newPage();
for (const url of urls) {
try {
const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
results.push({ url, endpoint: proxy.server, status: response ? response.status() : null });
} catch (error) {
results.push({
url,
endpoint: proxy.server,
error: isProxyError(error) ? `proxy: ${error.message}` : error.message,
});
}
}
} finally {
await context.close();
}
return results;
}
const browser = await chromium.launch({ headless: true });
try {
const queue = ENDPOINTS.map((proxy, index) => ({
proxy,
urls: TARGETS.filter((_, targetIndex) => targetIndex % ENDPOINTS.length === index),
}));
const workers = Array.from({ length: Math.min(CONCURRENCY, queue.length) }, async () => {
const collected = [];
while (queue.length > 0) {
const job = queue.shift();
collected.push(...(await runBatch(browser, job.proxy, job.urls)));
}
return collected;
});
const results = (await Promise.all(workers)).flat();
for (const result of results) {
console.log(JSON.stringify(result));
}
} finally {
await browser.close();
}
export PROXYSCRAPE_ENDPOINTS="http|HOST|PORT|USER|PASS,socks5|HOST|PORT|USER|PASS"
export TARGETS="https://example.com,https://example.org,https://example.net"
export CONCURRENCY=3
node rotate-contexts.mjs
Why this shape works:
- The browser starts once; contexts are cheap and each one carries its own endpoint.
- The worker pool caps how many contexts exist at the same time, which is the main memory lever in Playwright.
- Every result records the endpoint that served it, so failures can be attributed without guessing.
Step 6 — Treat proxy failures as endpoint failures, not page failures
net::ERR_PROXY_CONNECTION_FAILED— the proxy host or port is wrong or unreachable.net::ERR_TUNNEL_CONNECTION_FAILED— the CONNECT tunnel to an HTTPS target was refused; confirm the protocol you configured matches the endpoint.net::ERR_PROXY_AUTH_REQUESTED— credentials were rejected; re-check the username and password, and pass them as fields rather than inside the server URL.TimeoutError— the endpoint is slow or the target is heavy; raise the timeout and lower concurrency.
When one of these appears, remove that endpoint from the queue for the rest of the run and continue with the others. Retrying the same broken endpoint repeatedly only burns time.
Step 7 — Run it on Linux CI without exhausting resources
- Install browser dependencies in the pipeline with
npx playwright install --with-deps chromium. - Keep concurrency to two or three contexts per vCPU on a CI runner; each context carries a rendered page.
- Launch the browser once per process and close contexts in
finallyblocks. - Prefer datacenter proxies for high-volume, low-defence targets on scheduled jobs, and keep residential or mobile endpoints for the harder targets.
- Watch unit economics: pay-as-you-go billing makes bursty CI runs easy to start, while a flat monthly plan is usually cheaper for a fixed nightly job.
Troubleshooting
| Symptom | Likely cause | Fix |
|---|---|---|
| Pages load, but from your own IP | Proxy ignored (typo in server, or the option was not passed to the context) |
Log the exit IP with the Step 4 helper and raise when it matches the direct IP |
ERR_PROXY_CONNECTION_FAILED |
Wrong host or port, or CI egress blocked | Test the same endpoint with curl, then check network policy |
ERR_PROXY_AUTH_REQUESTED |
Credentials wrong, or embedded in the URL | Pass username and password fields, never user:pass@host |
ERR_TUNNEL_CONNECTION_FAILED on HTTPS |
Protocol mismatch between the endpoint and the scheme you configured | Try the HTTP endpoint, then the SOCKS5 endpoint, and keep the one that works |
| Every context reports the same exit IP | Browser-level proxy used instead of context-level, or the engine does not support per-context proxies | Use chromium.launch() and set proxy on each newContext() |
| The run kills the CI runner | Too many concurrent contexts | Lower CONCURRENCY to 2 and re-measure memory |
Summary
- In Playwright, the proxy belongs to the launch call or to the browser context, not to individual requests.
- Verify the exit IP before scraping; a misconfigured proxy can silently leave you on your own IP.
- Rotate by giving each context its own ProxyScrape HTTP or SOCKS5 endpoint and limiting how many contexts run at once.
- Classify
ERR_PROXY*andERR_TUNNEL*failures as endpoint problems, drop the endpoint, and keep the run going.