Skip to content
Intermediate / 3 min read

How to Set Up Smartproxy Datacenter Proxies on Linux for Python Price Monitoring

Route scheduled Python price-monitoring requests through Smartproxy datacenter proxies on Linux, with credentials kept out of your code, retries and backoff, rotation guidance, CSV logging, and a cron schedule.

Linux SOCKS5 HTTP(S) Web Scraping

Overview

Smartproxy is a beginner-friendly proxy network with residential and datacenter pools reachable over HTTP and SOCKS5, billed on a flat monthly model. This tutorial points its datacenter pool at a concrete, common job on Linux: scheduled price monitoring in Python.

Price monitoring suits datacenter IPs well. Requests are fast, the targets are public product pages, and the workload is read-heavy. If a particular retailer starts blocking datacenter ranges, the same script runs against the residential pool — only the endpoint and the credentials change.

By the end you will have:

  • Credentials stored outside your source code
  • A proxy request verified from the command line
  • A Python script that fetches product pages through the proxy with retries and backoff
  • Prices appended to a CSV file with timestamps
  • A cron entry that runs the job on a schedule

You do not need a GUI for any of this. Smartproxy also offers browser extensions, but a headless Linux job only needs the host, port, username, and password from your dashboard.

Prerequisites

  • A Linux host — the commands below work on the Debian/Ubuntu and Fedora/RHEL families
  • Python 3.9 or newer
  • A Smartproxy account with datacenter proxy access
  • The proxy host, port, username, and password listed in your Smartproxy dashboard

Datacenter or Residential?

Both pool types are configured identically in code, so the decision is about the target site, not about tooling.

Factor Datacenter Residential
Speed per request Fast Moderate
Cost per GB Lower Higher
Block risk on large retailers Higher Lower
Best fit High-volume reads of tolerant sites Protected sites and geo-accurate pricing

Rules of thumb:

  • Start with datacenter, because it is the cheaper and faster way to prove your parser works.
  • Move a single stubborn target to residential rather than migrating the whole job.
  • Keep the proxy scheme and authentication identical between pools so switching is a config change, not a rewrite.

Steps

Step 1 — Export your credentials and verify the endpoint

Set the four values from your dashboard as environment variables for the current shell session. Replace the placeholders with your own values.

export SMART_HOST="<host-from-dashboard>"
export SMART_PORT="<port-from-dashboard>"
export SMART_USER="<username>"
export SMART_PASS="<password>"

Run a single request through the proxy. Passing the credentials with -U avoids problems with special characters in the password.

curl -sS --max-time 20 -x "http://${SMART_HOST}:${SMART_PORT}" -U "${SMART_USER}:${SMART_PASS}" https://api.ipify.org

The response should be an IP address that is not your server's own IP. If you see a 407 error instead, jump to the troubleshooting table at the end.

Step 2 — Create the project and move the secrets into a file

  1. Create a directory and a virtual environment.
  2. Install requests plus python-dotenv, which loads the .env file into the environment at runtime.
mkdir -p ~/price-monitor && cd ~/price-monitor
python3 -m venv .venv && source .venv/bin/activate
pip install requests python-dotenv

Create the secrets file and lock down its permissions. This file should never be committed.

# ~/price-monitor/.env
SMART_HOST=<host-from-dashboard>
SMART_PORT=<port-from-dashboard>
SMART_USER=<username>
SMART_PASS=<password>
chmod 600 .env
cat > .gitignore <<'EOF'
.env
.venv/
EOF

Step 3 — Build the proxy URL in Python

Build the URL once and reuse it. Quoting the username and password with quote() protects against characters such as @, : or # that would otherwise break URL parsing.

# ~/price-monitor/proxy_config.py
import os
from urllib.parse import quote
from dotenv import load_dotenv

load_dotenv()

def proxy_url(scheme='http'):
    user = quote(os.environ['SMART_USER'], safe='')
    password = quote(os.environ['SMART_PASS'], safe='')
    host = os.environ['SMART_HOST']
    port = os.environ['SMART_PORT']
    return f'{scheme}://{user}:{password}@{host}:{port}'

PROXIES = {'http': proxy_url(), 'https': proxy_url()}

If your endpoint uses a different port for HTTPS traffic, set both keys separately rather than assuming the port is shared.

Step 4 — Fetch a page through the proxy with retries

Wrap the request in a small retry loop. Exponential backoff keeps you from hammering a target that is rate-limiting you.

# ~/price-monitor/fetch_price.py
import time
import requests
from proxy_config import PROXIES

HEADERS = {
    'User-Agent': 'Mozilla/5.0 (X11; Linux x86_64) AppleWebKit/537.36 '
                  '(KHTML, like Gecko) Chrome/124.0 Safari/537.36',
    'Accept-Language': 'en-US,en;q=0.9',
}

def fetch(url, retries=3, timeout=25):
    last_error = None
    for attempt in range(1, retries + 1):
        try:
            response = requests.get(url, headers=HEADERS, proxies=PROXIES, timeout=timeout)
            if response.status_code == 200:
                return response.text
            last_error = 'HTTP %s' % response.status_code
        except requests.RequestException as exc:
            last_error = str(exc)
        time.sleep(2 ** attempt)
    raise RuntimeError('Failed to fetch %s: %s' % (url, last_error))

if __name__ == '__main__':
    html = fetch('https://example.com/product/123')
    print(html[:500])

Step 5 — Extract the price and append it to CSV

HTML parsers such as BeautifulSoup or selectolax give more resilient extraction, but a regex is enough to validate the pipeline before you invest in selectors.

# ~/price-monitor/parse_price.py
import csv
import datetime
import re
import sys
from fetch_price import fetch

def parse_price(html):
    # Replace this pattern with a selector that matches the target site's markup.
    match = re.search(r'data-price=([0-9.]+)', html)
    return float(match.group(1)) if match else None

def main(url):
    price = parse_price(fetch(url))
    with open('prices.csv', 'a', newline='', encoding='utf-8') as handle:
        writer = csv.writer(handle)
        writer.writerow([datetime.datetime.now(datetime.UTC).isoformat(), url, price])

if __name__ == '__main__':
    main(sys.argv[1] if len(sys.argv) > 1 else 'https://example.com/product/123')

Step 6 — Decide how rotation should work for your endpoint

Rotation behaviour depends on the endpoint you were given, so check your dashboard before assuming anything:

  • If the port rotates the exit IP on every request, simply calling fetch() again is enough — each new connection can leave from a different IP.
  • If the port holds a sticky session, cycle through the ports listed in your dashboard and keep one port per worker.
  • Keep reusing the same requests.Session() only when you want a sticky session; a fresh session per request is the safer default for monitoring.

To confirm rotation, request the IP endpoint a few times in a row.

def check_ip():
    return fetch('https://api.ipify.org').strip()

if __name__ == '__main__':
    for _ in range(5):
        print(check_ip())

Step 7 — Add modest concurrency

A thread pool shortens a long list of URLs, but concurrency is also the fastest way to get blocked. Start small and only raise the worker count while error rates stay flat.

Concurrent workers When it is reasonable
1 to 3 First run against a new target
4 to 6 Stable datacenter run, no rate-limit responses
8 or more Only after testing, and usually with residential IPs
from concurrent.futures import ThreadPoolExecutor, as_completed
from fetch_price import fetch

urls = [f'https://example.com/product/{sku}' for sku in range(1, 6)]

with ThreadPoolExecutor(max_workers=5) as pool:
    futures = {pool.submit(fetch, url): url for url in urls}
    for future in as_completed(futures):
        url = futures[future]
        try:
            print(url, len(future.result()))
        except Exception as exc:
            print(url, 'failed:', exc)

Step 8 — Schedule the job with cron

Edit your crontab and add one line. The cd matters because the script relies on its .env file being in the working directory.

crontab -e
0 */6 * * * cd /home/<user>/price-monitor && .venv/bin/python parse_price.py https://example.com/product/123 >> run.log 2>&1

Use the absolute path to the virtual environment's Python. A bare python under cron often resolves to the system interpreter and will not see your installed packages.

Using SOCKS5 Instead of HTTP

Smartproxy supports SOCKS5 alongside HTTP, and switching is a one-line change. Confirm the SOCKS5 port in your dashboard, since it may differ from the HTTP port.

pip install 'requests[socks]'
PROXIES = {'http': proxy_url('socks5'), 'https': proxy_url('socks5')}

SOCKS5 is useful when you also want DNS resolution to travel through the proxy, or when another tool in your pipeline only speaks SOCKS5.

Troubleshooting

Symptom Likely cause What to try
407 Proxy Authentication Required Wrong or stale credentials, or a password with unescaped special characters Re-copy the values from the dashboard, verify with the curl -U command from step 1, and rely on quote() in Python
403 or a CAPTCHA page Target blocks datacenter ranges, or headers look automated Slow the run down, send realistic browser headers, and move that target to the residential pool
429 Too Many Requests Too many concurrent workers Reduce max_workers, keep the exponential backoff, and add a pause between batches
Connection timeout on every request Wrong host or port, or outbound traffic on that port is blocked Re-run the curl check; review ufw, iptables or any corporate egress policy for the proxy port
Works with an IP checker but fails on the target Anti-bot protection specific to that site Test one request in a normal browser session, adjust headers, and expect some targets to need residential IPs
CERTIFICATE_VERIFY_FAILED Outdated CA bundle on the host Update ca-certificates and the virtual environment's packages rather than disabling verification

Summary

  • Smartproxy's datacenter pool is a practical default for high-volume, read-only price monitoring on Linux.
  • Keep host, port, username, and password in a .env file with 600 permissions, and load them with python-dotenv.
  • Quote credentials before putting them in a proxy URL, or special characters will break parsing.
  • Always set a request timeout, retry with backoff, and keep concurrency low until error rates prove it is safe to increase.
  • Check your dashboard to learn whether your endpoint rotates per request or holds a sticky session, then design the run around that.
  • Switch a single blocked target to the residential pool instead of rebuilding the job, and use SOCKS5 only when you need it for DNS or tooling reasons.