首页
/ Scrapling 自适应 Web 抓取框架指南:Fetcher、Spider 与 CLI 全解析

Scrapling 自适应 Web 抓取框架指南:Fetcher、Spider 与 CLI 全解析

2026-09-04 18:47:39作者:董斯意

本文基于 Scrapling 仓库中的项目说明文档(docs/README_KR.md)整理而成,系统讲解这个“从单请求到大规模爬取通吃”的自适应 Web 抓取框架:你将学会如何选用三类 Fetcher(HTTP / 隐身浏览器 / 浏览器自动化)、搭建带检查点与多会话路由的 Spider 爬虫、使用 CLI 交互式 Shell 与 extract 命令做零代码提取,并理解安装依赖分层与性能基准背后的实现依据。

框架概览:一个库覆盖从单请求到完整爬取

Scrapling 的定位是自适应 Web Scraping 框架。它由三块能力组成:

  • 自适应解析器:解析器会“学习”网站变更,页面更新后能自动重新定位元素;
  • 抗反爬 Fetcher:无需额外配置即可绕过 Cloudflare Turnstile 一类反爬系统;
  • Spider 框架:支持暂停/恢复、自动代理轮换、并发多会话爬取——全部只需几行 Python 代码。

文档开篇给出的最小示例浓缩了前两块能力:

from scrapling.fetchers import Fetcher, AsyncFetcher, StealthyFetcher, DynamicFetcher

StealthyFetcher.adaptive = True
p = StealthyFetcher.fetch('https://example.com', headless=True, network_idle=True)  # 规避检测地抓取网站
products = p.css('.product', auto_save=True)   # 即使网站改版也能存活的数据提取
products = p.css('.product', adaptive=True)    # 网站结构变化后,用 adaptive=True 重新找回元素

而“从单请求扩展到完整爬取”只差一个 Spider 子类:

from scrapling.spiders import Spider, Response

class MySpider(Spider):
    name = "demo"
    start_urls = ["https://example.com/"]

    async def parse(self, response: Response):
        for item in response.css('.product'):
            yield {"title": item.css('h2::text').get()}

MySpider().start()

从源码结构看,scrapling 包按职责拆分为 scrapling/fetchers/(抓取入口)、scrapling/engines/(底层引擎与浏览器控制)、scrapling/spiders/(爬取框架)、scrapling/core/(解析核心与 Shell、存储)与 scrapling/cli.py(命令行入口),这与文档“一个库、分层可选依赖”的设计一致。

核心能力总览

Spider:完整爬虫框架

文档为 Spider 列出的能力,均能在 scrapling/spiders/spider.py 中找到对应实现载体:

  • Scrapy 风格 APIstart_urls、异步 parse 回调、Request/Response 对象;
  • 并发控制:类属性 concurrent_requests(默认 4)、concurrent_requests_per_domaindownload_delay,可在 Spider 中直接覆盖;
  • 多会话路由:HTTP 请求与隐身无头浏览器共享同一接口,通过请求 sid 把流量路由到不同会话;
  • 暂停与恢复:检查点持久化,Ctrl+C 优雅停止后重启即从断点继续。对应 scrapling/spiders/spider.py#L106-L111:构造参数 crawldir 用于存放检查点文件,interval(默认 300 秒)控制周期性落盘间隔;
  • 流式模式async for item in spider.stream() 实时接收带统计的 item,适合 UI、管道与长时爬取;
  • 拦截检测与重试:可自定义逻辑识别被拦截请求并自动重试(max_blocked_retries,默认 3);
  • AutoThrottle:自动按站点响应速度调整每域延迟,遇到拦截/限速时加倍延迟或按 Retry-After 等待,恢复后再提速。源码中对应 autothrottle_enabledautothrottle_start_delay(默认 5 秒)、autothrottle_max_delay(默认 60 秒)、autothrottle_block_backoff 等类属性(见 scrapling/spiders/spider.py#L88-L93);
  • robots.txt 遵守:可选 robots_txt_obey 标志(默认 False),遵守 DisallowCrawl-delayRequest-rate 指令并按域缓存;
  • 开发模式development_mode 首次运行时把响应缓存到磁盘(development_cache_dir),后续运行回放缓存,不再请求目标服务器,方便反复调试 parse() 逻辑;
  • 即用模板CrawlSpider(规则化链接跟踪)、SitemapSpider(sitemap/robots.txt 驱动)、XMLFeedSpider/CSVFeedSpider(XML/RSS、CSV 源站)、ShopifySpider(通过 JSON API 抓取任意 Shopify 商店全部商品,每个变体一个条目),实现见 scrapling/spiders/templates/
  • 链接提取:独立 LinkExtractor,支持 allow/deny 模式、域名过滤、CSS/XPath 作用域、扩展名过滤与 URL 规范化;
  • 内建导出result.items.to_json()to_jsonl()to_csv()to_xml() 直接导出结果,无需自建管道。

会话化的高级网站抓取

Fetcher 层提供四种入口(全部在 scrapling/fetchers/init.py 中延迟导出):

组件 定位 适用场景
Fetcher / AsyncFetcher 快速 HTTP 请求 普通页面,模拟浏览器 TLS 指纹与头部,支持 HTTP/3
DynamicFetcher Playwright 浏览器自动化(Chromium / 系统 Chrome) 动态渲染页面
StealthyFetcher 隐身 + 指纹伪装 Cloudflare Turnstile / 反爬拦截页
FetcherSession / DynamicSession / StealthySessionAsync* 版本 持久会话 跨请求保留 Cookie 与状态

其他关键能力:

  • 代理轮换:内置 ProxyRotator(循环或自定义策略),支持逐请求代理覆盖;
  • 域名与广告拦截:浏览器型 Fetcher 可拦截指定域名(含子域),或启用内建广告拦截——约 3,500 个已知广告/追踪域名,名单定义于 scrapling/engines/toolbelt/ad_domains.py
  • DNS 防泄漏:使用代理时可选经 Cloudflare DoH 路由 DNS 查询;
  • 远程浏览器cdp_url 连接已运行的浏览器(本地、远端或托管服务),executable_path 指定自编译 Chromium;
  • XHR 捕获:传入 capture_xhr URL 模式后,页面加载期间匹配的 XHR/fetch 响应会全部收集为 Response 对象存放在 response.captured_xhr,无需逆向接口即可拿到站点 API 数据;
  • 全异步支持:所有 Fetcher 均有异步版本与专属异步会话类。

自适应抓取与 AI 集成

  • 智能元素跟踪:基于相似度算法在改版后重定位元素(auto_save=True 保存基线、adaptive=True 恢复查找);
  • 灵活选择:CSS、XPath、条件过滤、文本匹配、正则匹配;
  • 相似元素发现find_similar() 自动找出与目标相似的元素;
  • MCP 服务器:内建 MCP 服务供 Claude/Cursor 等 AI 通过 Scrapling 先提取目标内容再交给模型,降低 token 消耗;还能跨调用保持浏览器会话、截图、经 CDP 控制远程浏览器。实现与说明见 scrapling/core/ai.pydocs/ai/mcp-server.mddocs/api-reference/mcp-server.md
  • Agent Skill:仓库内置即用型 Agent Skill,把整个库的 API 教给编码代理,使其生成的代码贴合当前 API 而非凭空猜测。

性能与工程化特性

官方说明强调:优化后的解析速度超过多数 Python 抓取库;内存占用经数据结构与延迟加载优化;JSON 序列化基于 orjson(见 pyproject.toml 核心依赖 orjson>=3.11.8),比标准库快约 10 倍;官方文档称其具备 92% 测试覆盖率与完整类型提示,并持续由 PyRight 与 MyPy 校验(pyproject.toml 中同时配置了 [tool.mypy][tool.pyright])。此外还提供:内建 IPython 交互式 Shell(可把 curl 请求转成 Scrapling 请求)、免代码 CLI 抓取、完整的 DOM 遍历 API(父/兄弟/子节点)、自动选择器生成、与 Scrapy/BeautifulSoup 风格一致的伪元素 API,以及 scrapling_response 装饰器——给 Scrapy 回调加一行装饰即可用 Scrapling 解析器解析已有响应(集成实现见 scrapling/integrations/scrapy.py)。

快速开始

基础 HTTP 请求

from scrapling.fetchers import Fetcher, FetcherSession

with FetcherSession(impersonate='chrome') as session:  # 使用 Chrome 最新 TLS 指纹
    page = session.get('https://quotes.toscrape.com/', stealthy_headers=True)
    quotes = page.css('.quote .text::text').getall()

# 或一次性请求
page = Fetcher.get('https://quotes.toscrape.com/')
quotes = page.css('.quote .text::text').getall()

隐身模式

from scrapling.fetchers import StealthyFetcher, StealthySession

with StealthySession(headless=True, solve_cloudflare=True) as session:  # 保持浏览器直到工作完成
    page = session.fetch('https://nopecha.com/demo/cloudflare', google_search=False)
    data = page.css('#padded_content a').getall()

# 或一次性请求 —— 为该请求开浏览器,完成后关闭
page = StealthyFetcher.fetch('https://nopecha.com/demo/cloudflare')
data = page.css('#padded_content a').getall()

完整浏览器自动化

from scrapling.fetchers import DynamicFetcher, DynamicSession

with DynamicSession(headless=True, disable_resources=False, network_idle=True) as session:
    page = session.fetch('https://quotes.toscrape.com/', load_dom=False)
    data = page.xpath('//span[@class="text"]/text()').getall()  # 同样支持 XPath

# 或一次性请求风格
page = DynamicFetcher.fetch('https://quotes.toscrape.com/')
data = page.css('.quote .text::text').getall()

Spider 实战

并发爬取带翻页的站点:

from scrapling.spiders import Spider, Request, Response

class QuotesSpider(Spider):
    name = "quotes"
    start_urls = ["https://quotes.toscrape.com/"]
    concurrent_requests = 10

    async def parse(self, response: Response):
        for quote in response.css('.quote'):
            yield {
                "text": quote.css('.text::text').get(),
                "author": quote.css('.author::text').get(),
            }

        next_page = response.css('.next a')
        if next_page:
            yield response.follow(next_page[0].attrib['href'])

result = QuotesSpider().start()
print(f"共抓取 {len(result.items)} 条引语")
result.items.to_json("quotes.json")

一个 Spider 内混用多种会话类型——把受保护页面路由到隐身会话:

from scrapling.spiders import Spider, Request, Response
from scrapling.fetchers import FetcherSession, AsyncStealthySession

class MultiSessionSpider(Spider):
    name = "multi"
    start_urls = ["https://example.com/"]

    def configure_sessions(self, manager):
        manager.add("fast", FetcherSession(impersonate="chrome"))
        manager.add("stealth", AsyncStealthySession(headless=True), lazy=True)

    async def parse(self, response: Response):
        for link in response.css('a::attr(href)').getall():
            if "protected" in link:
                yield Request(link, sid="stealth")
            else:
                yield Request(link, sid="fast", callback=self.parse)  # 显式回调

Scrapling Spider 架构示意图

长时爬取的暂停与恢复只需传入 crawldir

QuotesSpider(crawldir="./crawl_data").start()

按 Ctrl+C 即优雅暂停并自动保存进度;下次启动传入相同 crawldir 便从断点恢复。

如果完全不想写爬取逻辑,可直接继承模板,例如抓取 Shopify 商店全目录:

from scrapling.spiders import ShopifySpider

class MyStore(ShopifySpider):
    target_website = "example.com"

result = MyStore().start()  # 商店全部商品,每个变体一个条目

高级解析与导航

from scrapling.fetchers import Fetcher

page = Fetcher.get('https://quotes.toscrape.com/')

# 多种选择方式
quotes = page.css('.quote')                          # CSS
quotes = page.xpath('//div[@class="quote"]')         # XPath
quotes = page.find_all('div', {'class': 'quote'})    # BeautifulSoup 风格
quotes = page.find_all('div', class_='quote')
quotes = page.find_all(['div'], class_='quote')
quotes = page.find_all(class_='quote')
quotes = page.find_by_text('quote', tag='div')       # 按文本内容查找

# 高级导航
quote_text = page.css('.quote')[0].css('.text::text').get()
quote_text = page.css('.quote').css('.text::text').getall()  # 链式选择
first_quote = page.css('.quote')[0]
author = first_quote.next_sibling.css('.author::text')
parent_container = first_quote.parent

# 元素关系与相似度
similar_elements = first_quote.find_similar()
below_elements = first_quote.below_elements()

不抓取网页也能直接使用解析器,用法完全一致:

from scrapling.parser import Selector

page = Selector("<html>...</html>")

异步会话管理

import asyncio
from scrapling.fetchers import FetcherSession, AsyncStealthySession, AsyncDynamicSession

async with FetcherSession(http3=True) as session:  # 上下文管理器,同步/异步模式均可
    page1 = session.get('https://quotes.toscrape.com/')
    page2 = session.get('https://quotes.toscrape.com/', impersonate='firefox135')

async with AsyncStealthySession(max_pages=2) as session:
    tasks = [session.fetch(url) for url in
             ['https://example.com/page1', 'https://example.com/page2']]
    print(session.get_pool_stats())   # 可选:浏览器标签池状态(使用中/空闲/出错)
    results = await asyncio.gather(*tasks)
    print(session.get_pool_stats())

CLI 与交互式 Shell

Scrapling 提供完整命令行接口(入口定义在 scrapling/cli.py):

scrapling shell

启动交互式 Web Scraping Shell(带 Scrapling 预置对象、快捷键、curl 转 Scrapling 请求等工具;对应文档 docs/cli/interactive-shell.md)。

免编程直接把页面导出为文件——默认提取 body 内容,输出格式由扩展名决定:.txt 为纯文本、.md 为 Markdown 化内容、.html 为原始 HTML:

scrapling extract get 'https://example.com' content.md
scrapling extract get 'https://example.com' content.txt --css-selector '#fromSkipToProducts' --impersonate 'chrome'
scrapling extract fetch 'https://example.com' content.md --css-selector '#fromSkipToProducts' --no-headless
scrapling extract stealthy-fetch 'https://nopecha.com/demo/cloudflare' captchas.html --css-selector '#padded_content a' --solve-cloudflare

scrapling/cli.py 的实现可以看出各子命令的完整参数面:

  • extract get/post/put/delete(HTTP 类):--impersonate(单个浏览器名,或逗号分隔列表随机选取)、--stealthy-headers(默认开启)、--css-selector/-s(返回全部匹配)、--proxy--timeout(默认 30 秒)、--cookies--headers/-H--params/-p--verify--follow-redirects
  • extract fetch / stealthy-fetch(浏览器类):在共享参数之外还有 --headless/--no-headless(默认无头)、--wait-selector--wait(加载后额外等待毫秒)、--network-idle--disable-resources(丢弃非必要资源提速)、--solve-cloudflare--block-ads(拦截已知广告/追踪域)、--dns-over-https(经 Cloudflare DoH 防 DNS 泄漏)、--real-chrome(使用本机 Chrome)、--locale
  • 所有 extract 命令均支持 --ai-targeted,只提取主体内容并清理隐藏元素,便于喂给 AI。

完整的 Shell 与 extract 命令细节分别见 docs/cli/overview.mddocs/cli/extract-commands.md

性能基准

文档给出的“5000 个嵌套元素文本提取”对比(单位 ms,vs Scrapling 为相对倍数):

# 耗时 (ms) 相对 Scrapling
1 Scrapling 1.99 1.0x
2 Parsel/Scrapy 2.06 1.035x
3 原生 Lxml 2.56 1.286x
4 PyQuery 23.98 ~12x
5 Selectolax 197.02 ~99x
6 MechanicalSoup 1545.15 ~776.5x
7 BS4 + lxml 1562.10 ~785.0x
8 BS4 + html5lib 3412.73 ~1714.9x

元素相似度与文本查找对比:Scrapling 2.3 ms vs AutoScraper 12.58 ms(约 5.47 倍)。

测量方法可在 benchmarks.py 中复现:对含 5000 个 div.item 的 HTML 文档,先用 timeit 做 2 轮预热,再以 time.process_time 计时、repeat=100 取平均(benchmarks.py#L22-L43);其中 Scrapling 项直接执行 Selector(large_html, adaptive=False).css(".item::text").getall(),而 Lxml 对照组也刻意使用与 Parsel/Scrapling 相同的 HTML 解析器以保证公平。

安装与依赖分层

Scrapling 要求 Python 3.10+pyproject.tomlrequires-python = ">=3.10",当前版本 0.4.13):

pip install scrapling

注意:基础安装只包含解析器引擎及其依赖(lxmlcssselectorjsontldw3lib),不含 Fetcher 与 CLI 相关依赖。此时 from scrapling.fetchers import ... 会抛 ModuleNotFoundErrorscrapling/fetchers/init.py 采用模块级 __getattr__ 延迟导入映射表,真正需要某 Fetcher 时才去导入其实现模块——缺少依赖时该导入即失败,这正是文档中该警告的来源。

需要 Fetcher 与 Spider 时,安装可选依赖并下载浏览器:

pip install "scrapling[fetchers]"

scrapling install            # 常规安装
scrapling install --force    # 强制重装

这会将浏览器、系统依赖与指纹伪装依赖一并下载;也可以直接用代码安装:

from scrapling.cli import install

install([], standalone_mode=False)           # 常规安装
install(["--force"], standalone_mode=False)  # 强制重装

其余可选功能(pyproject.toml 定义的 extras):

pip install "scrapling[ai]"      # MCP 服务器(mcp、markdownify + fetchers)
pip install "scrapling[shell]"   # 交互式 Shell 与 extract 命令(IPython + fetchers)
pip install "scrapling[all]"     # 全部功能

安装任意外挂功能后,若尚未执行过,仍需 scrapling install 补装浏览器依赖。

Docker

每个发布版都会自动构建并推送含全部功能与浏览器的镜像:

docker pull pyd4vinci/scrapling
# 或
docker pull ghcr.io/d4vinci/scrapling:latest

镜像构建脚本见仓库根目录 Dockerfile

使用注意与许可

官方免责说明:本库仅供教育与研究目的,使用者须自行遵守所在司法辖区的爬虫与隐私法律,并尊重目标网站的条款与 robots.txt——这一点与 Spider 内建的 robots_txt_obey 能力相呼应(实现见 scrapling/spiders/robotstxt.py)。

该项目以 BSD-3-Clause 许可发布(LICENSE)。代码致谢:scrapling/core/translator.py 中的选择器翻译子模块借鉴了 BSD 许可的 Parsel 项目。贡献者请先阅读 CONTRIBUTING.md;测试套件覆盖 CLI、解析器、Spider、Fetcher(同步/异步)、Scrapy 集成等模块(tests/),可作为行为验证的参照。

登录后查看全文
热门项目推荐
相关项目推荐

项目优选

收起
kernelkernel
deepin linux kernel
C
33
18
ops-transformerops-transformer
本项目是CANN提供的transformer类大模型算子库,实现网络在NPU上加速计算。
C++
1.12 K
2.72 K
ops-nnops-nn
本项目是CANN提供的神经网络类计算算子库,实现网络在NPU上加速计算。
C++
904
1.82 K
docsdocs
暂无描述
Markdown
889
5.78 K
pytorchpytorch
作为 Ascend for PyTorch 社区的核心组件,TorchNPU 是昇腾专为 PyTorch 打造的深度学习适配插件,使 PyTorch 框架能够直接调用昇腾 NPU,为开发者提供昇腾 AI 处理器的超强算力。
Python
854
1.34 K
kernelkernel
openEuler内核是openEuler操作系统的核心,既是系统性能与稳定性的基石,也是连接处理器、设备与服务的桥梁。
C
527
590
jiuwenswarmjiuwenswarm
JiuwenSwarm 是一款基于openJiuwen开发的智能AI Agent,它能够将大语言模型的强大能力,通过你日常使用的各类通讯应用,直接延伸至你的指尖。
Python
3.52 K
1.01 K
ops-mathops-math
本项目是CANN提供的数学类基础计算算子库,实现网络在NPU上加速计算。
C++
1.33 K
1.45 K
AscendNPU-IRAscendNPU-IR
AscendNPU-IR是基于MLIR(Multi-Level Intermediate Representation)构建的,面向昇腾亲和算子编译时使用的中间表示,提供昇腾完备表达能力,通过编译优化提升昇腾AI处理器计算效率,支持通过生态框架使能昇腾AI处理器与深度调优
C++
540
384
flutter_flutterflutter_flutter
本仓库是 Flutter SDK 与 Flutter Engine 的 OpenHarmony 适配版本,由 CPF-Flutter 团队维护。开发者可使用熟悉的 Flutter 技术栈开发 OpenHarmony 应用,3.35.7 及以后的适配版本可基于本仓库源码构建支持 OpenHarmony 的 Flutter Engine。
Dart
1.17 K
341