用 Go regexp 包实现文本匹配、过滤与抓取:build-web-application-with-golang 第 7.3 节精讲

原创2026-10-04 14:10:381,034 阅读
文章标签:文档教程

用 Go regexp 包实现文本匹配、过滤与抓取:build-web-application-with-golang 第 7.3 节精讲

本篇围绕《Build Web Application with Golang》第 7.3 节(pt-br/07.3.md)展开,系统讲解 Go 标准库 regexp 包的三大能力:匹配(Match)、查找(Find)与替换(Replace/Expand),并给出可用于爬虫清洗 HTML、表单输入校验的完整实战代码。读者学完后,将能独立用正则表达式完成 Web 开发中的文本采集、数据提取与用户输入合法性校验。

为什么需要正则表达式:strings 与 regexp 的取舍

Go 的 strings 包已经能完成很多文本工作,例如查找(Contains、Index)、替换(Replace)、解析(Split、Join)等,而且它的性能优于正则表达式。但这些都属于"琐碎"的简单操作。当需求升级为大小写不敏感的搜索、按模式提取数据或批量清洗文本时,regexp 才是最佳选择。

在实际开发中应遵循一个朴素原则:

  • 如果 strings 包足以满足需求,优先使用它——代码易写、易读、性能好;
  • 如果需要更高级的模式操作,再引入 regexp。

正则表达式(Regexp)是复杂但强大的模式匹配与文本操作工具。虽然在性能上不如纯文本匹配,但它足够灵活:基于其语法,几乎可以从源文本中过滤出任何内容。在 Web 开发中需要采集数据时,用正则表达式提取有效信息并不困难。

Go regexp 包概览:RE2 标准与 UTF-8

Go 通过标准库 regexp 提供官方正则支持。如果你在其他语言中使用过正则表达式,上手会非常容易。需要特别注意的是:Go 实现了 RE2 标准(\C 除外)。

RE2 是 Google 设计的一套正则引擎规范,其核心设计目标是保证线性时间匹配——即匹配耗时与输入长度成正比,不存在灾难性回溯(catastrophic backtracking)问题。因此 Go 的正则不支持反向引用(backreference)与部分环视(lookaround)语法,这是与 Perl 风格正则的最大差异。对 Web 场景中的用户输入、网络抓取数据做匹配时,这种确定性带来了安全性上的好处。

另一个要点:Go 中的正则默认按 UTF-8 处理字符。中文、日文等多字节字符都可以直接纳入字符类,例如第 4.2 节表单验证中用 ^<a href="https://link.gitcode.com/i/e466b4a6d435029b46f8e56bf298805f" target="_blank">\x{4e00}-\x{9fa5}]+$ 校验纯中文字符串(见 [pt-br/04.2.md)。

Match:三个顶层匹配函数

regexp 包提供 3 个包级匹配函数:若模式与输入匹配则返回 true,否则返回 false;若正则本身有语法错误,则返回 error。

func Match(pattern string, b []byte) (matched bool, error error)
func MatchReader(pattern string, r io.RuneReader) (matched bool, error error)
func MatchString(pattern string, s string) (matched bool, error error)

三个函数的行为一致,区别仅在于输入源类型:[]byte 字节切片、io.RuneReader 与 string。适合一次性校验、无需复用正则的场景。

示例一:校验 IP 地址

func IsIP(ip string) (b bool) {
	if m, _ := regexp.MatchString("^[0-9]{1,3}\\.[0-9]{1,3}\\.[0-9]{1,3}\\.[0-9]{1,3}$", ip); !m {
		return false
	}
	return true
}

该模式要求文本从 ^ 到 $ 恰好是四段 1~3 位数字、以 . 分隔,可用于 Web 后端对用户提交的 IP 字段做初步过滤。

示例二:判断命令行输入是否为数字

func main() {
	if len(os.Args) == 1 {
		fmt.Println("Usage: regexp [string]")
		os.Exit(1)
	} else if m, _ := regexp.MatchString("^[0-9]+$", os.Args[1]); m {
		fmt.Println("Number")
	} else {
		fmt.Println("Not number")
	}
}

模式 ^[0-9]+$ 要求整个字符串由一位或多位数字组成,以此区分"数字"与"非数字"输入。

Filter:从校验到采集——爬虫实战

Match 模式只能验证内容是否合法,无法切割、过滤或采集数据。要完成这些任务,必须使用"复杂模式":先通过 Compile 系列把正则编译成 *Regexp 对象,再调用其方法。

下面是一个经典爬虫示例:抓取网页 HTML,逐步清洗掉标签、样式与脚本,最后输出纯文本。

package main

import (
	"fmt"
	"io/ioutil"
	"net/http"
	"regexp"
	"strings"
)

func main() {
	resp, err := http.Get("http://www.baidu.com")
	if err != nil {
		fmt.Println("http get error.")
	}
	defer resp.Body.Close()
	body, err := ioutil.ReadAll(resp.Body)
	if err != nil {
		fmt.Println("http read error")
		return
	}

	src := string(body)

	// Convert HTML tags to lower case.
	re, _ := regexp.Compile("\\<[\\S\\s]+?\\>")
	src = re.ReplaceAllStringFunc(src, strings.ToLower)

	// Remove STYLE.
	re, _ = regexp.Compile("\\<style[\\S\\s]+?\\</style\\>")
	src = re.ReplaceAllString(src, "")

	// Remove SCRIPT.
	re, _ = regexp.Compile("\\<script[\\S\\s]+?\\</script\\>")
	src = re.ReplaceAllString(src, "")

	// Remove all HTML code in angle brackets, and replace with newline.
	re, _ = regexp.Compile("\\<[\\S\\s]+?\\>")
	src = re.ReplaceAllString(src, "\n")

	// Remove continuous newline.
	re, _ = regexp.Compile("\\s{2,}")
	src = re.ReplaceAllString(src, "\n")

	fmt.Println(strings.TrimSpace(src))
}

逐段解析这个"清洗流水线":

  1. \\<[\\S\\s]+?\\>:[\S\s] 匹配任意非空白或空白字符(等价于"任意字符"),+? 为非贪婪量词,整体匹配"从 < 到 > 的最短内容",即单个 HTML 标签。第一步用 ReplaceAllStringFunc 配合 strings.ToLower 把所有标签名统一为小写,便于后续精确匹配;
  2. \\<style[\\S\\s]+?\\</style\\>:非贪婪匹配整段 <style>...</style>,直接替换为空串删除;
  3. \\<script[\\S\\s]+?\\</script\\>:同样删除 <script>...</script> 及其内容;
  4. 再次用标签模式把所有剩余标签替换为换行符 \n;
  5. \\s{2,}:把连续两个以上的空白字符压缩为单个换行,收尾用 strings.TrimSpace 去掉首尾空白。

需要说明的是:示例代码基于早期 Go 版本编写(使用 ioutil.ReadAll),在现代 Go 中可改用 io.ReadAll;且爬取外部站点时应遵守目标站点的 robots 协议与合规要求。

Compile 系列:复杂模式的入口

示例中的第一步 regexp.Compile 会先校验正则语法是否正确,成功则返回可复用的 *Regexp 对象,失败则返回错误:

func Compile(expr string) (*Regexp, error)
func CompilePOSIX(expr string) (*Regexp, error)
func MustCompile(str string) *Regexp
func MustCompilePOSIX(str string) *Regexp

CompilePOSIX 与 Compile 的区别在于搜索语义:

  • CompilePOSIX 遵循 POSIX 语法,采用 leftmost-longest(最左最长) 匹配;
  • Compile 采用 leftmost(最左) 匹配。

例如对正则 [a-z]{2,4} 与内容 "aa09aaa88aaaa":CompilePOSIX 返回最长的 aaaa,而 Compile 返回最先匹配到的 aa。带 Must 前缀的版本在语法错误时会直接 panic(适合在初始化阶段使用,把错误提前暴露),否则返回 *Regexp。

Find:18 个查找方法

编译得到的 *Regexp 提供 18 个查找方法,覆盖三种输入源(字节切片、字符串、io.RuneReader):

func (re *Regexp) Find(b []byte) []byte
func (re *Regexp) FindAll(b []byte, n int) [][]byte
func (re *Regexp) FindAllIndex(b []byte, n int) [][]int
func (re *Regexp) FindAllString(s string, n int) []string
func (re *Regexp) FindAllStringIndex(s string, n int) [][]int
func (re *Regexp) FindAllStringSubmatch(s string, n int) [][]string
func (re *Regexp) FindAllStringSubmatchIndex(s string, n int) [][]int
func (re *Regexp) FindAllSubmatch(b []byte, n int) [][][]byte
func (re *Regexp) FindAllSubmatchIndex(b []byte, n int) [][]int
func (re *Regexp) FindIndex(b []byte) (loc []int)
func (re *Regexp) FindReaderIndex(r io.RuneReader) (loc []int)
func (re *Regexp) FindReaderSubmatchIndex(r io.RuneReader) []int
func (re *Regexp) FindString(s string) string
func (re *Regexp) FindStringIndex(s string) (loc []int)
func (re *Regexp) FindStringSubmatch(s string) []string
func (re *Regexp) FindStringSubmatchIndex(s string) []int
func (re *Regexp) FindSubmatch(b []byte) [][]byte
func (re *Regexp) FindSubmatchIndex(b []byte) []int

忽略输入源差异后,核心方法可归纳为 8 个:

func (re *Regexp) Find(b []byte) []byte
func (re *Regexp) FindAll(b []byte, n int) [][]byte
func (re *Regexp) FindAllIndex(b []byte, n int) [][]int
func (re *Regexp) FindAllSubmatch(b []byte, n int) [][][]byte
func (re *Regexp) FindAllSubmatchIndex(b []byte, n int) [][]int
func (re *Regexp) FindIndex(b []byte) (loc []int)
func (re *Regexp) FindSubmatch(b []byte) [][]byte
func (re *Regexp) FindSubmatchIndex(b []byte) []int

其中 Find 返回第一个匹配内容,FindAll 返回全部匹配(n < 0 表示全部,n > 0 限制返回数量,n == 0 返回空),带 Index 的变体返回匹配的起止位置,带 Submatch 的变体额外返回括号分组捕获的内容。

综合示例:逐一观察每个方法的输出

package main

import (
	"fmt"
	"regexp"
)

func main() {
	a := "I am learning Go language"

	re, _ := regexp.Compile("[a-z]{2,4}")

	// Find the first match.
	one := re.Find([]byte(a))
	fmt.Println("Find:", string(one))

	// Find all matches and save to a slice, n less than 0 means return all matches, indicates length of slice if it's greater than 0.
	all := re.FindAll([]byte(a), -1)
	fmt.Println("FindAll", all)

	// Find index of first match, start and end position.
	index := re.FindIndex([]byte(a))
	fmt.Println("FindIndex", index)

	// Find index of all matches, the n does same job as above.
	allindex := re.FindAllIndex([]byte(a), -1)
	fmt.Println("FindAllIndex", allindex)

	re2, _ := regexp.Compile("am(.*)lang(.*)")

	// Find first submatch and return array, the first element contains all elements, the second element contains the result of first (), the third element contains the result of second ().
	// Output:
	// the first element: "am learning Go language"
	// the second element: " learning Go ", notice spaces will be outputed as well.
	// the third element: "uage"
	submatch := re2.FindSubmatch([]byte(a))
	fmt.Println("FindSubmatch", submatch)
	for _, v := range submatch {
		fmt.Println(string(v))
	}

	// Same thing like FindIndex().
	submatchindex := re2.FindSubmatchIndex([]byte(a))
	fmt.Println(submatchindex)

	// FindAllSubmatch, find all submatches.
	submatchall := re2.FindAllSubmatch([]byte(a), -1)
	fmt.Println(submatchall)

	// FindAllSubmatchIndex,find index of all submatches.
	submatchallindex := re2.FindAllSubmatchIndex([]byte(a), -1)
	fmt.Println(submatchallindex)
}

对输入 "I am learning Go language" 与模式 [a-z]{2,4}:

  • Find 返回第一个匹配 "am";
  • FindAll(..., -1) 返回全部匹配切片,如 ["am" "learning" "Go" "language"](注意大小写敏感,I 大写不匹配);
  • FindIndex 返回 [2 4] 这样的起止下标,FindAllIndex 返回每个匹配的起止下标列表;
  • 对模式 am(.*)lang(.*),FindSubmatch 返回 [][]byte:第 0 个元素是整体匹配 "am learning Go lang"(此处示例注释为 "am learning Go language"),第 1 个元素是第一个分组 (.*) 捕获的内容(含空格),第 2 个元素是第二个分组捕获的内容;
  • FindSubmatchIndex、FindAllSubmatch、FindAllSubmatchIndex 则分别返回分组捕获的索引与全部匹配。

这段示例完整展示了"查找 + 分组捕获 + 索引定位"的组合用法,是数据提取类任务(如解析日志、抽取字段)的核心。

Match 方法族

除包级函数外,*Regexp 自身也提供 3 个匹配方法,行为与包级函数完全一致。事实上,包级导出函数底层就是调用这些方法:

func (re *Regexp) Match(b []byte) bool
func (re *Regexp) MatchReader(r io.RuneReader) bool
func (re *Regexp) MatchString(s string) bool

在实际项目中,更推荐先 Compile 得到 *Regexp 再调用 MatchString,因为正则只编译一次,可反复复用,性能更好。

Replace:字符串替换六件套

*Regexp 提供 6 个替换方法:

func (re *Regexp) ReplaceAll(src, repl []byte) []byte
func (re *Regexp) ReplaceAllFunc(src []byte, repl func([]byte) []byte) []byte
func (re *Regexp) ReplaceAllLiteral(src, repl []byte) []byte
func (re *Regexp) ReplaceAllLiteralString(src, repl string) string
func (re *Regexp) ReplaceAllString(src, repl string) string
func (re *Regexp) ReplaceAllStringFunc(src string, repl func(string) string) string

它们在上面的爬虫示例中已经充分体现:

  • ReplaceAllString(src, "") 用于删除(替换为空串)style 与 script 块;
  • ReplaceAllString(src, "\n") 用于把标签替换为换行符;
  • ReplaceAllStringFunc(src, strings.ToLower) 对每个匹配的标签调用回调函数做小写转换;
  • 带 Literal 的版本不解析替换串中的 $1、$name 等反向引用,把替换串当作纯字面量处理。

Expand:基于命名分组的模板展开

Expand 系列方法用于把匹配结果按模板展开到目标切片,非常适合"批量格式化提取结果"的场景:

func (re *Regexp) Expand(dst []byte, template []byte, src []byte, match []int) []byte
func (re *Regexp) ExpandString(dst []byte, template string, src string, match []int) []byte

示例:把多行命令文本中的参数按命名分组重新组装输出。

func main() {
	src := []byte(`
		call hello alice
		hello bob
		call hello eve
	`)
	pat := regexp.MustCompile(`(?m)(call)\s+(?P<cmd>\w+)\s+(?P<arg>.+)\s*$`)
	res := []byte{}
	for _, s := range pat.FindAllSubmatchIndex(src, -1) {
		res = pat.Expand(res, []byte("$cmd('$arg')\n"), src, s)
	}
	fmt.Println(string(res))
}

关键点解读:

  • (?m) 开启多行模式,使 ^/$ 匹配每行的行首行尾;
  • (?P<cmd>\w+)、(?P<arg>.+) 是命名分组语法,分组捕获的内容可通过 $cmd、$arg 在模板中引用;
  • FindAllSubmatchIndex 返回全部匹配及其分组索引,Expand 依据这些索引把模板中的 $cmd('$arg')\n 展开成 hello('alice') 等格式。

这种"命名分组 + Expand 模板"的组合,是文本协议解析、代码生成、日志格式化等场景的利器。

从源码看 regexp 在表单验证中的落地

本书第 4.2 节(pt-br/04.2.md)已经强调:Web 开发的首要原则是绝不信任来自客户端的表单数据,所有入参都必须在服务端校验。而正则表达式正是其中最强力的工具之一。第 4.2 节给出了几类经典校验模式:

  • 数字:^[0-9]+$;
  • 中文姓名:^[\x{4e00}-\x{9fa5}]+$(Unicode 中日文字符区段);
  • 英文字母:^[a-zA-Z]+$;
  • 邮箱:^([\w\.\_]{2,10})@(\w{1,}).([a-z]{2,4})$。

这些校验逻辑在仓库示例代码中有完整实现:文件 pt-br/code/src/apps/ch.4.2/validator/main.go 中的 checkChineseName 用 regexp.MatchString("^<a href="https://link.gitcode.com/i/65a85c546c083752c72462ed1eb87f77" target="_blank">\\x{4e00}-\\x{9fa5}]+$", strings.Trim(str, " ")) 校验中文姓名,checkEmail 用 `^[^@]+@[^@]+$` 校验邮箱格式。这些校验器通过 stringValidator 映射表注册到 ProfilePage.GetErrors() 中,由 [pt-br/code/src/apps/ch.4.2/main.go 的 checkProfile 处理器在表单提交时统一调用——这正是第 7.3 节正则知识在真实 Web 应用中的直接落地。

从源码结构可以推断出本书的编排思路:第 7 章(pt-br/07.0.md)系统讲解文本处理(XML、JSON、正则、模板、文件、字符串),而正则一章为后续的输入过滤(第 9.2 节)和 Web 框架开发中的表单处理(第 14 章)提供了基础工具。

小结

通过本章,你已完整掌握 Go regexp 包的三大能力:

能力 核心 API 典型场景
匹配 Match / MatchReader / MatchString 校验 IP、数字、邮箱、中文等输入合法性
查找 Find / FindAll / FindSubmatch 系列 从网页、日志中提取字段与分组数据
替换与展开 ReplaceAll 系列 / Expand 系列 清洗 HTML、格式化输出、模板展开

使用建议:简单的查找替换优先用 strings 包;需要模式匹配、分组提取或批量清洗时使用 regexp。记住 Go 采用 RE2 标准(\C 除外),字符按 UTF-8 处理;在 Web 开发中,正则表达式是服务端输入校验与数据采集不可或缺的组成部分。

相关链接

登录后查看全文
build-web-application-with-golang