Advanced Web Scraping with ChangeMyIP Residential Proxies and Playwright
Combine ChangeMyIP's rotating residential proxies with Playwright to scrape dynamic websites at scale. This advanced tutorial covers per-request IP rotation, browser fingerprint evasion, and robust error handling.
Overview
Playwright is a powerful browser automation library that supports Chromium, Firefox, and WebKit. When combined with ChangeMyIP residential proxies, it becomes a robust tool for advanced web scraping, bypassing geo-restrictions and rate limits. This tutorial covers setting up Playwright with ChangeMyIP rotating residential proxies, implementing per-request IP rotation, managing browser fingerprints, and handling dynamic content.
Prerequisites
- Node.js 16+ and npm
- A ChangeMyIP account with residential proxy access (or datacenter if you prefer)
- Basic knowledge of JavaScript/Node.js
Understanding ChangeMyIP Residential Proxy Endpoints
ChangeMyIP provides residential proxies in over 40 countries. You can authenticate using username/password. The endpoint format for HTTP proxies is:
http://<username>:<password>@<proxy-host>:<port>
For SOCKS5:
socks5://<username>:<password>@<proxy-host>:<port>
Replace placeholders with credentials from your ChangeMyIP dashboard. Rotating residential proxies assign a new IP per request by default, or you can maintain a sticky session using a session ID in the username (consult ChangeMyIP docs). For this tutorial, we assume rotating.
Step 1: Set Up a Playwright Project
- Create a new directory and initialise npm:
mkdir changeMyIP-playwright-scraper
cd changeMyIP-playwright-scraper
npm init -y
- Install Playwright and its dependencies:
npm install playwright
npx playwright install
Step 2: Configure Playwright to Use a ChangeMyIP Proxy
Create a file scraper.js and add the following code. This launches a Chromium browser with your ChangeMyIP proxy:
const { chromium } = require('playwright');
const proxyUrl = 'http://username:password@proxy-host:port'; // Replace with your ChangeMyIP details
(async () => {
const browser = await chromium.launch({
proxy: {
server: proxyUrl,
},
});
const page = await browser.newPage();
await page.goto('https://httpbin.org/ip');
const content = await page.textContent('body');
console.log(content);
await browser.close();
})();
Run it with node scraper.js. You should see the proxy IP, not your real IP.
Step 3: Implement Rotating Proxies per Request
For advanced scraping, you might want a fresh IP for each page or request. Since ChangeMyIP residential proxies rotate on each request by default, you can simply create a new browser context for each scrape job. If you need to force a new IP, launch a new browser instance or use a new proxy session.
Here's a function that scrapes a URL with a new IP each time:
async function scrapeWithNewIP(url) {
const browser = await chromium.launch({
proxy: { server: proxyUrl },
});
const page = await browser.newPage();
await page.goto(url, { waitUntil: 'networkidle' });
const data = await page.evaluate(() => document.body.innerText);
await browser.close();
return data;
}
Call this function for each target URL. For efficiency, reuse the browser and create a new context with a different proxy (if your plan allows multiple endpoints). Otherwise, launching a new browser is acceptable for moderate scale.
Step 4: Manage Browser Fingerprints to Avoid Blocks
Anti-bot systems detect automation. Randomise browser properties per session. Playwright allows setting user agent, viewport, timezone, and more. Use a library like playwright-extra with puppeteer-extra-plugin-stealth (compatible with Playwright via playwright-extra). Or manually set headers.
Example: launch with custom user agent and viewport:
const browser = await chromium.launch({
proxy: { server: proxyUrl },
});
const context = await browser.newContext({
userAgent: 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/120.0.0.0 Safari/537.36',
viewport: { width: 1920, height: 1080 },
locale: 'en-US',
timezoneId: 'America/New_York',
});
Rotate user agents from a list for each session.
Step 5: Scrape Dynamic Content and Handle Pagination
Playwright auto-waits for elements. Use page.waitForSelector for dynamic content. For infinite scroll, scroll until no new items load.
Example: scrape product titles from an e-commerce page with infinite scroll:
async function scrapeProducts(page) {
await page.goto('https://example.com/products', { waitUntil: 'networkidle' });
let previousHeight = 0;
while (true) {
const currentHeight = await page.evaluate(() => document.body.scrollHeight);
if (currentHeight === previousHeight) break;
previousHeight = currentHeight;
await page.evaluate(() => window.scrollTo(0, document.body.scrollHeight));
await page.waitForTimeout(2000);
}
const products = await page.evaluate(() => {
return Array.from(document.querySelectorAll('.product-title')).map(el => el.innerText);
});
return products;
}
Step 6: Implement Retry Logic with Fresh IPs
When a request fails (e.g., CAPTCHA, timeout), retry with a new IP. Since rotating proxies change IP per request, simply retrying might suffice. For sticky sessions, you can change the session ID.
async function scrapeWithRetry(url, maxRetries = 3) {
for (let i = 0; i < maxRetries; i++) {
const browser = await chromium.launch({ proxy: { server: proxyUrl } });
try {
const page = await browser.newPage();
await page.goto(url, { timeout: 30000 });
const data = await page.content();
await browser.close();
return data;
} catch (error) {
console.error(`Attempt ${i + 1} failed: ${error.message}`);
await browser.close();
}
}
throw new Error(`Failed after ${maxRetries} retries`);
}
Troubleshooting
- Authentication error: Double-check username and password. Ensure they are URL-encoded if they contain special characters.
- Proxy connection refused: Verify the proxy host and port. Check that your IP is whitelisted if using IP authentication (ChangeMyIP supports username/password, so use that).
- CAPTCHA or block: Rotate user agents and introduce random delays. Use residential proxies for better success rates.
- Slow performance: Reduce browser instances, use headless mode, and block unnecessary resources (images, CSS) with
page.route. - Sticky session not working: Consult ChangeMyIP documentation for the correct session ID format in the username.
Summary
You have learned how to integrate ChangeMyIP residential proxies with Playwright for advanced web scraping. By rotating IPs, managing browser fingerprints, and handling dynamic content, you can scrape at scale while avoiding blocks. Remember to respect websites' terms of service and robots.txt. ChangeMyIP's residential proxy network offers 40+ countries and multiple protocols, making it a suitable choice for these tasks.