Xberg PHP 绑定中 Markdown 输出格式的契约测试与实现原理
Xberg PHP 绑定中 Markdown 输出格式的契约测试与实现原理
Xberg 是一个以 Rust 为核心的多语言文档智能提取引擎,PHP 是其官方绑定之一。output_format: markdown 是 Xberg 提取管线中最常用的输出格式开关之一:它把任意输入文档(PDF、Office、HTML、图片等)的正文转换为结构化的 Markdown,供 RAG 索引、LLM 上下文注入与文档归档直接消费。本文以仓库中 PHP 契约测试(contract test)为线索,讲解如何在 PHP 中开启 Markdown 输出、如何通过 URI 与字节两种输入路径验证结果,并深入 Rust 核心解释该格式的枚举定义、渲染器分发与 MCP 配置覆盖机制,最后给出可复制运行的完整 PHP 示例与断言清单。
一、测试片段的定位:它验证什么
关联文档是 docs-site/src/snippets-generated/php/contract/output_format_markdown.md,它是一段由 alef 工具自动生成的 PHP 契约测试代码(文件名中的 contract 表示它属于跨语言契约测试集合)。其核心逻辑只有三步:
<?php
declare(strict_types=1);
require_once __DIR__ . '/vendor/autoload.php';
use Xberg\Xberg;
use Xberg\ExtractInput;
$input = \Xberg\ExtractInput::from_json(json_encode(["kind" => "uri", "uri" => "https://example.com/pdf/fake_memo.pdf"]));
$result = Xberg::extract($input, ["output_format" => "markdown"]);
var_dump($result->getResults()[0]->mimeType);
var_dump($result->getResults()[0]->content);
var_dump($result->getResults()[0]->getMetadata()->outputFormat);
这段代码验证了三件事实:
- 输入路径:通过
ExtractInput::from_json()构造一个uri类型的输入(指向fake_memo.pdf),走的是「从 URL 拉取文档」的提取路径; - 配置传递:在
Xberg::extract()的第二个参数中传入["output_format" => "markdown"],即声明输出格式为 Markdown; - 结果读取:从
getResults()[0]读取第一个提取结果,断言其mimeType为application/pdf、content为非空字符串(长度 ≥ 10)、getMetadata()->outputFormat为markdown。
对应地,e2e/php/tests/ContractTest.php 中的 test_output_format_markdown()(第 519 行起)就是这段代码的正式 PHPUnit 版本:它使用 mock server 提供 fake_memo.pdf,通过 ExtractInput::from_json 构造 URI 输入,再用 ExtractionConfig::from_json(json_encode(["outputFormat" => "markdown"])) 构造配置,最终断言 mimeType == "application/pdf"、strlen(content) >= 10、metadata.outputFormat == "markdown" 三个条件。
二、两种输入路径:URI 与 bytes
契约测试同时覆盖了两种输入方式,这是理解该测试价值的关键:
2.1 URI 路径(文档中的示例)
关联文档展示的是 URI 路径:ExtractInput::from_json 接受一个 JSON 字符串,其中 kind 字段决定输入类型。当 kind 为 uri 时,只需提供 uri 字段,Xberg 会自行抓取该地址的文档并按 MIME 类型路由到对应提取器。对应的 fixture 是 fixtures/contract/output_format_markdown.json,它用 mock server 返回 application/octet-stream 头部的 PDF 内容(body_file 指向 test_documents/pdf/fake_memo.pdf),并在 config 中声明 output_format: markdown。
2.2 Bytes 路径(配套测试)
同一主题的 bytes 版本在 docs-site/src/snippets-generated/php/contract/output_format_bytes_markdown.md 与 e2e/php/tests/ContractTest.php 的 test_output_format_bytes_markdown()(第 504 行起)中:
<?php
declare(strict_types=1);
require_once __DIR__ . '/vendor/autoload.php';
use Xberg\Xberg;
use Xberg\ExtractInput;
$input = \Xberg\ExtractInput::from_json(json_encode([
"bytes" => "pdf/fake_memo.pdf", // 简化写法;真实测试中为 PDF 文件的字节数组
"config" => ["outputFormat" => "markdown"],
"filename" => "fake_memo.pdf",
"kind" => "bytes",
"mimeType" => "application/pdf"
]));
$result = Xberg::extract($input, ["output_format" => "markdown"]);
var_dump($result->getResults()[0]->mimeType);
var_dump($result->getResults()[0]->content);
var_dump($result->getResults()[0]->getMetadata()->outputFormat);
bytes 路径的关键差异在于:输入 JSON 必须包含 kind: "bytes"、bytes(原始字节数组)、mimeType 与可选的 filename。对应 fixture fixtures/contract/output_format_bytes_markdown.json 中内嵌了完整的一页 PDF 字节数组(约 12 KB,含 FlateDecode 流、Helvetica 字体与 fake-memo 标题),测试期望与 URI 版本完全一致。两条路径共享同一套 Rust 提取核心,因此断言(mimeType、content 长度、outputFormat 元数据)完全相同。
三、PHP API 层的实现:从静态方法到原生扩展
从 PHP 侧看,入口是 Xberg\Xberg::extract()。该类的源码位于 crates/xberg-php/src/Xberg.php,它是一个由 alef 生成的静态门面类,全部方法都委托给原生扩展类 \Xberg\XbergApi:
public static function extract(ExtractInput $input, ExtractionConfig $config): ExtractionResult
{
return \Xberg\XbergApi::extract($input, $config); // delegate to native extension class
}
配套的批量入口 extractBatch(array $inputs, ExtractionConfig $config) 也在同一文件中(第 67 行起),契约测试的 PHPUnit 版本实际就是通过 XbergApi::extractBatch(<a href="https://link.gitcode.com/i/966ff538af8db0558d62f3a33beaf3e4" target="_blank">$input], $config) 调用的。真正的实现位于 Rust 侧 [crates/xberg-php/src/lib.rs,它用 ext-php-rs 把 Xberg 核心类型逐一导出为 PHP 类:
ExtractInput:kind、uri、bytes、mimeType、filename、内联config等字段都通过#[php(prop)]映射为 PHP 属性,并支持from_json与from_json风格的构造函数;ExtractionConfig:暴露outputFormat属性(见 lib.rs 第 2303 行附近的#[php(prop, name = "outputFormat")]),同时接受output_format蛇形命名(通过#[serde(alias = "outputFormat")]双向兼容);ExtractionResult/ExtractedDocument:提供getResults()、getMetadata()、getPages()、getTables()等读取器,元数据中的outputFormat即测试断言的字段。
因此,两种配置写法是等价的:
// 蛇形命名(JSON 配置键)
$config = \Xberg\ExtractionConfig::from_json(json_encode(["output_format" => "markdown"]));
// 驼峰命名(PHP 属性名)
$config = \Xberg\ExtractionConfig::from_json(json_encode(["outputFormat" => "markdown"]));
两者都会解析为 Rust 核心的 OutputFormat::Markdown。
四、Rust 核心:OutputFormat 枚举与渲染器分发
Markdown 输出格式的权威定义在 crates/xberg/src/core/config/formats.rs(第 17~35 行):
pub enum OutputFormat {
Plain, // 纯文本(默认)
Markdown, // Markdown 格式
Djot, // Djot 轻量标记
Html, // HTML 格式
Json, // 以标题驱动分节的 JSON 树
DocTags, // Docling DocTags(表格渲染为 OTSL)
Custom(String), // 通过 RendererRegistry 注册的自定义渲染器名称
}
几点值得注意的语义(均有源码可查):
- 默认值是
Plain(#[default]),所以不传output_format时返回的是原始提取文本;显式声明markdown才会触发 Markdown 渲染。 FromStr解析是大小写不敏感的(第 68~81 行):"markdown"、"MARKDOWN"、"md"、"MD"都会解析为Markdown;"plain"与"text"都映射到Plain。该模块自带的单元测试(第 196~200 行)明确覆盖了markdown/MARKDOWN/md/MD四种写法。- 渲染器分发:
Markdown、Djot、Html、DocTags以及Custom(name)都通过渲染器注册表(RendererRegistry)查找对应名称的渲染器(renderer_name()返回"markdown"等,第 42~51 行);Plain与Json不走渲染器注册表,由核心直接处理。这意味着markdown渲染器本身是一个可注册、可替换的组件,第三方渲染器可通过Custom("docx")、Custom("latex")等名称接入。
OutputFormat 枚举值最终控制 ExtractedDocument.content 字段的生成方式(见该文件第 10~14 行的模块注释),这也是测试断言 strlen(content) >= 10 与 metadata.outputFormat == "markdown" 的底层依据。
五、配置解析、校验与 MCP 覆盖
5.1 配置解析
output_format 是 ExtractionConfig 的一个字段,其解析、环境变量映射与文件配置合并分别在:
- crates/xberg/src/core/config/extraction/core.rs
- crates/xberg/src/core/config/extraction/env.rs(环境变量映射)
- crates/xberg/src/core/config/extraction/file_config.rs(配置文件解析)
- crates/xberg/src/core/config/extraction/types.rs(配置结构体定义)
5.2 配置校验
crates/xberg/src/core/config_validation/mod.rs 与 crates/xberg/src/core/config_validation/sections.rs 负责对配置做校验;例如禁止 extract 输入中同时携带互相冲突的 OCR 配置(error_extract_input_conflicting_ocr),输出格式相关的不合法组合也会在这里被拦截。
5.3 MCP 配置覆盖(值得注意的机制)
crates/xberg/src/mcp/format.rs 展示了 MCP(Model Context Protocol)服务中 output_format 的强制覆盖逻辑:当 MCP 请求中携带 "output_format": "markdown" 时,服务会把请求级配置的 output_format 强制改写为 markdown(该文件第 121~130 行与第 171~180 行的测试分别验证了 config.output_format 被覆写为 markdown 的行为)。这说明:在 MCP 场景下,输出格式可由调用方在请求内声明并覆盖默认配置,而普通 PHP API 调用则需要像契约测试那样显式传入 output_format。
六、可复制的完整 PHP 实战示例
将契约测试展开为一段可直接在项目中运行的完整脚本(环境要求:PHP 8.2+、已通过 Composer 安装 xberg-io/xberg 并加载扩展,参见 packages/php/README.md 的安装说明):
<?php
declare(strict_types=1);
require_once __DIR__ . '/vendor/autoload.php';
use Xberg\ExtractInput;
use Xberg\ExtractionConfig;
use Xberg\Xberg;
// 1) URI 输入 + Markdown 输出:从远端 URL 提取 PDF 并转为 Markdown
$uriInput = ExtractInput::from_json(json_encode([
'kind' => 'uri',
'uri' => 'https://example.com/pdf/fake_memo.pdf',
]));
$config = ExtractionConfig::from_json(json_encode([
'outputFormat' => 'markdown', // 或 'output_format' => 'markdown'
]));
$result = Xberg::extract($uriInput, $config);
$doc = $result->getResults()[0];
echo 'MIME: ' . $doc->mimeType . PHP_EOL; // application/pdf
echo 'Chars: ' . strlen($doc->content) . PHP_EOL; // >= 10
echo 'Output format: ' . $doc->getMetadata()->outputFormat . PHP_EOL; // markdown
// 2) Bytes 输入 + Markdown 输出:直接处理内存中的 PDF 字节
$pdfBytes = file_get_contents('fake_memo.pdf'); // 本地文件
$bytesInput = ExtractInput::from_json(json_encode([
'kind' => 'bytes',
'bytes' => array_values(unpack('C*', $pdfBytes)), // 转为字节数组
'mimeType' => 'application/pdf',
'filename' => 'fake_memo.pdf',
]));
$result = Xberg::extract($bytesInput, $config);
echo 'Bytes path output format: ' . $result->getResults()[0]->getMetadata()->outputFormat . PHP_EOL;
// 3) 批量提取(等价于契约测试实际调用的 extractBatch)
$batchResult = Xberg::extractBatch([$uriInput, $bytesInput], $config);
foreach ($batchResult->getResults() as $i => $r) {
printf("Doc %d: mime=%s, format=%s, chars=%d\n",
$i, $r->mimeType, $r->getMetadata()->outputFormat, strlen($r->content));
}
要点提醒:
- 字节数组的构造在 PHP 中可用
unpack('C*', $bytes)完成,也可以直接沿用契约 fixture 中预置的整数数组; filename与mimeType在 bytes 路径中建议显式给出,便于核心做格式检测与元数据填充;- 若输入是扫描件,可叠加 OCR 配置(
OcrConfig,如backend: 'tesseract'、language: 'eng')后再开启 Markdown 输出,此时 Markdown 内容来自 OCR 识别结果; - 断言清单与契约一致:
mimeType == "application/pdf"、strlen(content) >= 10、metadata.outputFormat == "markdown"。
七、小结
从本文可以看出,output_format: markdown 的契约测试虽短,却完整覆盖了 Xberg PHP 绑定的一条关键路径:输入构造(URI/bytes)→ 配置传递(outputFormat/output_format)→ Rust 核心格式分发(OutputFormat::Markdown)→ 结果读取(content + metadata)。配套的 fixture 断言(fixtures/contract/output_format_markdown.json、fixtures/contract/output_format_bytes_markdown.json)与 e2e 测试(e2e/php/tests/ContractTest.php)保证了该行为在跨语言、跨版本演进中保持一致。开发者如果要在自己的 PHP 服务中接入文档智能提取,直接以本文的示例为模板、按需叠加 OCR、表格或分块配置即可,Markdown 输出开箱即用。