Skip to content
Advanced / 3 min read

Scrapy 代理轮换:如何在 Scrapy 中轮换代理(以及何时更适合使用专用 WireGuard 网关)

使用自定义中间件在 Scrapy 中实现代理轮换,优雅地处理封禁,并了解何时像 NordLayer 这样的稳定 WireGuard 网关是 SEO 监控和网页抓取的合适选择。

Linux WireGuard 网页抓取 SEO 监测

概述

大规模网页抓取经常会遇到速率限制、IP 封禁或地理限制。Scrapy 是一个流行的 Python 框架,但它并不开箱即用地支持代理轮换。本教程将向你展示如何为 Scrapy 构建轮换代理中间件、配置重试,并判断何时专用 WireGuard VPN 网关(如 NordLayer)比一组轮换代理更合适。

你将学习两种方法:

  • 轮换代理:每个会话使用多个 IP 来分散请求并避免被封禁。
  • 专用 WireGuard 网关:对于需要一致性的任务,例如排名跟踪、广告验证或访问会惩罚 IP 变化的地理特定内容,使用单一稳定的 IP。

NordLayer 提供企业级 WireGuard 隧道,具有专用网关和固定月费。它不是轮换代理服务,但当一致性很重要时,它可以作为你的 Scrapy 爬虫的稳定出口 IP。

先决条件

  • Python 3.7 或更高版本
  • 基本熟悉 Scrapy
  • 一份 HTTP/HTTPS 或 SOCKS5 代理列表(用于轮换方法)
  • 可选:一个带有专用 WireGuard 网关的 NordLayer 账户(用于稳定 IP 方法)
  • 一台 Linux 服务器或本地机器(命令以 Debian/Ubuntu 为例;其他系统请自行调整)

理解 Scrapy 中的代理轮换

Scrapy 使用下载器中间件来处理请求和响应。要轮换代理,你需要拦截每个请求,从代理池中分配一个代理,并可选择在请求失败时用新代理重试。

关键概念:

  • 代理池:代理 URL 列表(如需要则包含凭据)。
  • 轮换策略:随机选择、轮询,或根据代理健康状况加权。
  • 重试逻辑:使用不同代理重新排队失败的请求。

何时应考虑改用专用 WireGuard 网关:

  • 你需要一个一致的 IP,用于跨多个位置进行 SEO 排名跟踪。
  • 你的目标网站对频繁 IP 变化的封禁比静态 IP 更激进。
  • 你希望获得加密的、团队管理的访问方式,而无需维护代理列表。

步骤

1. 安装 Scrapy 并创建项目

如果尚未安装,请安装 Scrapy 并搭建一个新项目。

pip install scrapy
scrapy startproject proxy_scraper
cd proxy_scraper
scrapy genspider example example.com

2. 创建轮换代理中间件

打开 proxy_scraper/middlewares.py 并添加以下中间件。它会为每个请求随机选择一个代理,并在常见封禁状态码时重试。

# proxy_scraper/middlewares.py
import random
from scrapy import signals
from scrapy.exceptions import NotConfigured

class RotatingProxyMiddleware:
    def __init__(self, proxy_list):
        self.proxy_list = proxy_list
        self.current_proxy = None

    @classmethod
    def from_crawler(cls, crawler):
        proxy_list = crawler.settings.getlist('ROTATING_PROXY_LIST')
        if not proxy_list:
            raise NotConfigured('ROTATING_PROXY_LIST is empty')
        return cls(proxy_list)

    def process_request(self, request, spider):
        self.current_proxy = random.choice(self.proxy_list)
        request.meta['proxy'] = self.current_proxy
        spider.logger.debug(f'Using proxy: {self.current_proxy}')

    def process_response(self, request, response, spider):
        if response.status in [403, 429]:
            spider.logger.warning(f'Blocked with proxy {self.current_proxy}. Retrying...')
            new_request = request.copy()
            new_request.dont_filter = True
            return new_request
        return response

3. 配置设置以启用中间件并定义代理

编辑 proxy_scraper/settings.py 以激活中间件并提供你的代理列表。将示例代理替换为你自己的代理。

# proxy_scraper/settings.py

DOWNLOADER_MIDDLEWARES = {
    'proxy_scraper.middlewares.RotatingProxyMiddleware': 543,
}

ROTATING_PROXY_LIST = [
    'http://user:[email protected]:8080',
    'http://user:[email protected]:8080',
    'socks5://user:[email protected]:1080',
]

# Optional: retry settings
RETRY_TIMES = 5
RETRY_HTTP_CODES = [500, 502, 503, 504, 522, 524, 408, 429, 403]

4. 用简单的 spider 测试轮换

运行你的 spider 并查看 DEBUG 日志,确认使用了不同的代理。

scrapy crawl example -L DEBUG

如果你看到类似 Using proxy: http://... 的行,说明中间件正在工作。如果所有请求都使用同一个代理,请检查中间件顺序,并确认 ROTATING_PROXY_LIST 不为空。

5. 通过健康检查和退避提高可靠性

对于生产环境抓取,添加基本的代理健康跟踪。此示例在代理被反复封禁后将其标记为失败,并暂时移除它。

# proxy_scraper/middlewares.py (enhanced)
import random
from collections import defaultdict

class RotatingProxyMiddleware:
    def __init__(self, proxy_list):
        self.proxy_list = proxy_list
        self.failure_count = defaultdict(int)
        self.max_failures = 3

    @classmethod
    def from_crawler(cls, crawler):
        proxy_list = crawler.settings.getlist('ROTATING_PROXY_LIST')
        return cls(proxy_list)

    def process_request(self, request, spider):
        available = [p for p in self.proxy_list if self.failure_count[p] < self.max_failures]
        if not available:
            spider.logger.error('All proxies exhausted')
            return
        proxy = random.choice(available)
        request.meta['proxy'] = proxy

    def process_response(self, request, response, spider):
        proxy = request.meta.get('proxy')
        if response.status in [403, 429]:
            self.failure_count[proxy] += 1
            new_request = request.copy()
            new_request.dont_filter = True
            return new_request
        else:
            self.failure_count[proxy] = 0
        return response

6. 何时切换到专用 WireGuard 网关

轮换代理非常适合大规模、无状态的抓取。但有些任务需要稳定、专用的 IP:

  • 需要一致位置数据的 SEO 排名跟踪
  • 访问将会话绑定到 IP 的 API
  • 多个用户共享同一出口 IP 的团队协作抓取

NordLayer WireGuard 隧道为你提供专用网关,并采用固定月费。你可以将所有 Scrapy 流量通过该隧道路由,从而有效使用单个静态 IP。

7. 在 Linux 上设置 WireGuard 并通过它路由 Scrapy

安装 WireGuard,并使用你的 NordLayer 网关凭据配置隧道。

sudo apt update
sudo apt install wireguard

在 /etc/wireguard/wg0.conf 创建配置文件(将占位符替换为你的 NordLayer 值)。

[Interface]
PrivateKey = <your-private-key>
Address = 10.0.0.2/32
DNS = 1.1.1.1

[Peer]
PublicKey = <nordlayer-public-key>
Endpoint = <gateway-ip>:51820
AllowedIPs = 0.0.0.0/0
PersistentKeepalive = 25

启动隧道:

sudo wg-quick up wg0

现在所有 Scrapy 流量(以及任何其他流量)都会通过 NordLayer 网关出口。验证你的公网 IP 是否已更改:

curl ifconfig.me

像往常一样运行你的 Scrapy spider。它将使用稳定的网关 IP。

8. 结合两种方法进行混合抓取

你可以同时使用两者:将大多数请求通过轮换代理路由,但将需要一致 IP 的请求通过 WireGuard 隧道发送。在 Scrapy 中,你可以针对特定请求有条件地将代理设置为 None(这会使用系统默认路由,即 WireGuard 隧道)。

# Example: in a spider, for rank-tracking requests, don't set a proxy
# The default route (WireGuard) will be used.
def start_requests(self):
    # Rotating proxies for general pages
    yield scrapy.Request('https://example.com/general', meta={'proxy': None})
    # For rank tracking, bypass the rotating middleware
    yield scrapy.Request('https://example.com/serp', meta={'proxy': 'direct'})

然后修改中间件,当 proxy == 'direct' 时跳过轮换:

def process_request(self, request, spider):
    if request.meta.get('proxy') == 'direct':
        return None  # use default route
    # ... rest of rotation logic

故障排除

  • 中间件未生效:确保 DOWNLOADER_MIDDLEWARES 包含正确的路径,并且值 (543) 与其他中间件不冲突。检查 Scrapy 的默认中间件顺序。
  • 所有请求仍然使用同一个 IP:验证 ROTATING_PROXY_LIST 已填充,并且中间件的 process_request 正在运行。添加 print 或调试日志。
  • 频繁出现 403/429 错误:你的代理可能已被封禁或速度太慢。降低 CONCURRENT_REQUESTS 和 DOWNLOAD_DELAY,或切换到更高质量的代理。如果目标网站激进地封禁轮换 IP,请考虑使用专用 WireGuard 网关。
  • WireGuard 隧道连接失败:仔细检查私钥、公钥、端点和防火墙规则(默认 UDP 端口 51820)。运行 sudo wg show 查看握手状态。
  • Scrapy 忽略 WireGuard 隧道:如果你设置了 AllowedIPs = 0.0.0.0/0,所有流量都应通过隧道路由。用 ip route 确认,并用 curl ifconfig.me 测试。

总结

你现在知道如何使用自定义中间件在 Scrapy 中实现代理轮换、添加重试逻辑并监控代理健康。对于需要稳定、专用 IP 的项目——例如 SEO 排名跟踪或团队访问——像 NordLayer 这样的 WireGuard 网关可以补充或替代轮换代理。选择与你的抓取目标相匹配的方法:轮换用于规模和匿名性,专用网关用于一致性和控制。