Go 正则表达式实战:深入解析 Build Web Application with Golang 第 7.3 节 regexp 包

原创2026-09-29 23:11:291,359 阅读
文章标签:文档教程

Go 正则表达式实战:深入解析 Build Web Application with Golang 第 7.3 节 regexp 包

本篇技术指南以开源项目 Build Web Application with Golang(de/07.3.md)中第 7.3 节"Regexp"为骨架,系统讲解 Go 标准库 regexp 包在 Web 开发中的应用:从 Match 系列函数做输入校验,到 Compile 复杂模式完成抓取、过滤与文本清洗,再到 Find、ReplaceAll、Expand 等方法的完整用法。读完本文,你将掌握 Go 正则表达式的 RE2 语法约束、三类匹配入口、18 个查找方法的选型,以及如何在真实表单校验与爬虫场景中落地这些能力。

为什么 Web 开发需要正则表达式

正则表达式(Regular Expression)是一种强大但也复杂的模式匹配与文本处理工具。它的执行性能虽然不如纯文本匹配,却极其灵活:只要给出合适的语法,几乎可以从任意源文本中过滤出你想要的内容。在 Web 开发中收集数据时,用正则表达式提取有意义的信息几乎是常规操作。

Go 在标准库中提供了 regexp 包作为官方支持。如果你在其他语言中用过正则表达式,上手会非常容易。这里需要特别留意一个标准层面的差异:Go 实现的是 RE2 标准,除了 \C 之外遵循 RE2 的完整语法约定。RE2 是 Google 设计的一种线性时间复杂度的正则引擎,它刻意不支持回溯(backtracking),因此不存在灾难性回溯(catastrophic backtracking)导致的服务阻塞问题,这是 Go 正则适合服务端处理用户输入的重要原因。

另外要明确一个选型原则:strings 包其实能完成许多工作,比如搜索(Contains、Index)、替换(Replace)、解析(Split、Join)等,而且速度比正则更快。但这些都属于"琐碎"操作;当你需要大小写不敏感搜索等更高级的行为时,正则才是最佳选择。结论很简单:strings 能满足就用 strings(易用、易读、更快);需要更高级操作再上 regexp。

如果你还记得前面章节的表单验证(第 4 章),我们当时就是用正则来校验用户输入的有效性,并且要意识到所有字符都是 UTF-8 编码的。现在就来深入学习 Go 的 regexp 包。

Match:三函数快速匹配验证

regexp 包提供了 3 个包级函数用于匹配判断:命中模式返回 true,否则返回 false。

func Match(pattern string, b []byte) (matched bool, error error)
func MatchReader(pattern string, r io.RuneReader) (matched bool, error error)
func MatchString(pattern string, s string) (matched bool, error error)

这三个函数做的事情完全一致——检查 pattern 是否匹配输入源,匹配返回 true;唯一的差异在于输入源的三种类型:字节切片([]byte)、io.RuneReader(以 rune 为单位流式读取的 reader)和字符串(string)。注意:如果正则本身存在语法错误,函数会返回 error,因此实际使用时通常用 m, _ := regexp.MatchString(...) 或显式处理第二个返回值。

示例一:校验 IP 地址

下面用 MatchString 校验一个 IPv4 格式的字符串:

func IsIP(ip string) (b bool) {
	if m, _ := regexp.MatchString("^[0-9]{1,3}\\.[0-9]{1,3}\\.[0-9]{1,3}\\.[0-9]{1,3}$", ip); !m {
		return false
	}
	return true
}

模式 ^[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}$ 的含义是:从开头(^)到结尾($)必须是由点号分隔的四段数字,每段 1~3 位([0-9]{1,3}),点号需要转义为 \.。注意该模式只校验"形如 IPv4",并不校验每段是否在 0~255 范围内,如需严格校验可进一步改写。

示例二:校验命令行输入是否为数字

func main() {
	if len(os.Args) == 1 {
		fmt.Println("Usage: regexp [string]")
		os.Exit(1)
	} else if m, _ := regexp.MatchString("^[0-9]+$", os.Args[1]); m {
		fmt.Println("Number")
	} else {
		fmt.Println("Not number")
	}
}

这里通过 ^[0-9]+$ 判断用户传入的第一个命令行参数是否全部由数字组成。运行方式:go run main.go 123 输出 Number;go run main.go abc 输出 Not number;不带参数则打印用法提示并退出。

以上 Match/MatchReader/MatchString 只解决"是否匹配"的验证问题,使用起来都很简单。

仓库印证:表单验证器中的正则实践

第 4 章的表单验证示例正是这一模式的真实落地。在 validator 源码 中,checkChineseName 用正则校验中文姓名只能包含中文字符:

// Checks if all the characters are chinese characters. Won't check if empty.'
func checkChineseName(str string) error {
	if str != "" {
		if m, _ := regexp.MatchString("^[\\x{4e00}-\\x{9fa5}]+$", strings.Trim(str, " ")); !m {
			return errors.New("Please make sure that the chinese name only contains chinese characters.")
		}
	}
	return nil
}

其中 \x{4e00}-\x{9fa5} 是 UTF-8 编码下的中文 Unicode 码点区间,这正是文档强调"所有字符都是 UTF-8"的实际体现。同文件里的 checkEmail 也借助正则做基础格式校验:

func checkEmail(str string) error {
	if m, err := regexp.MatchString(`^[^@]+@[^@]+$`, str); !m {
		fmt.Println("err = ", err)
		return errors.New("Please enter a valid email address.")
	}
	return nil
}

这两个校验函数分别挂在 stringValidator 映射表(de/code/src/apps/ch.4.2/validator/main.go#L50-L59)的 chineseName 与 email 键上,通过 GetErrors() 在 ch.4.2 主程序 中被 checkProfile 调用。你可以运行 go run main.go(在 de/code/src/apps/ch.4.2 目录下,并将 GOPATH 指向代码目录,参见 de/code/readme.md)后访问 http://localhost:9090/profile 实际体验。

Filter:进入复杂模式,用 Compile 过滤与切割数据

匹配模式只能"验证"内容,却无法切割、过滤或收集数据。要完成这些任务,必须使用正则的复杂模式(先编译出 *Regexp 对象,再调用其方法)。

爬虫实战:清洗 HTML 页面

假设我们要写一个简单的爬虫。下面的示例演示了何时必须用正则来过滤与切割数据——它抓取百度首页,然后逐步清洗掉 HTML 标签、样式与脚本,最终只保留纯文本:

package main

import (
	"fmt"
	"io/ioutil"
	"net/http"
	"regexp"
	"strings"
)

func main() {
	resp, err := http.Get("http://www.baidu.com")
	if err != nil {
		fmt.Println("http get error.")
	}
	defer resp.Body.Close()
	body, err := ioutil.ReadAll(resp.Body)
	if err != nil {
		fmt.Println("http read error")
		return
	}

	src := string(body)

	// Convert HTML tags to lower case.
	re, _ := regexp.Compile("\\<[\\S\\s]+?\\>")
	src = re.ReplaceAllStringFunc(src, strings.ToLower)

	// Remove STYLE.
	re, _ = regexp.Compile("\\<style[\\S\\s]+?\\</style\\>")
	src = re.ReplaceAllString(src, "")

	// Remove SCRIPT.
	re, _ = regexp.Compile("\\<script[\\S\\s]+?\\</script\\>")
	src = re.ReplaceAllString(src, "")

	// Remove all HTML code in angle brackets, and replace with newline.
	re, _ = regexp.Compile("\\<[\\S\\s]+?\\>")
	src = re.ReplaceAllString(src, "\n")

	// Remove continuous newline.
	re, _ = regexp.Compile("\\s{2,}")
	src = re.ReplaceAllString(src, "\n")

	fmt.Println(strings.TrimSpace(src))
}

清洗流程共四步,每一步都对应一个正则:

  1. \\<[\\S\\s]+?\\>:匹配尖括号内的标签,配合 ReplaceAllStringFunc(src, strings.ToLower) 将标签统一转为小写(HTML 标签大小写不敏感,统一小写便于后续匹配);
  2. \\<style[\\S\\s]+?\\</style\\>:把 <style>...</style> 整块删除(替换为空串);
  3. \\<script[\\S\\s]+?\\</script\\>:把 <script>...</script> 整块删除;
  4. 再次用 \\<[\\S\\s]+?\\> 把所有剩余标签替换为换行符,然后用 \\s{2,} 把连续空白压缩为单个换行,最后 strings.TrimSpace 去掉首尾空白。

其中 [\\S\\s] 表示"任意字符"(\S 非空白 + \s 空白),+? 是非贪婪量词,保证匹配到最短的标签内容。这里 Compile 是复杂模式的第一步:它会校验正则语法是否正确,然后返回一个可复用的 *Regexp 对象供后续操作使用。

Compile 家族:四种编译入口

func Compile(expr string) (*Regexp, error)
func CompilePOSIX(expr string) (*Regexp, error)
func MustCompile(str string) *Regexp
func MustCompilePOSIX(str string) *Regexp

CompilePOSIX 与 Compile 的区别在于匹配语义:前者遵循 POSIX 语法,采用最左最长匹配(leftmost longest search);后者采用最左匹配(leftmost search)。举例说明:对模式 [a-z]{2,4} 匹配内容 "aa09aaa88aaaa" 时,CompilePOSIX 返回 aaaa(取最长的连续匹配),而 Compile 返回 aa(取最左的第一个匹配)。带 Must 前缀的版本在语法错误时会直接 panic,否则返回 *Regexp,适用于模式在编写期就确定的场景(例如上面 Expand 示例中的 MustCompile)。

Find 家族:18 个方法背后的 8 种核心能力

有了 *Regexp 对象,我们就可以用它操作内容了。regexp 包提供如下 18 个查找方法:

func (re *Regexp) Find(b []byte) []byte
func (re *Regexp) FindAll(b []byte, n int) [][]byte
func (re *Regexp) FindAllIndex(b []byte, n int) [][]int
func (re *Regexp) FindAllString(s string, n int) []string
func (re *Regexp) FindAllStringIndex(s string, n int) [][]int
func (re *Regexp) FindAllStringSubmatch(s string, n int) [][]string
func (re *Regexp) FindAllStringSubmatchIndex(s string, n int) [][]int
func (re *Regexp) FindAllSubmatch(b []byte, n int) [][][]byte
func (re *Regexp) FindAllSubmatchIndex(b []byte, n int) [][]int
func (re *Regexp) FindIndex(b []byte) (loc []int)
func (re *Regexp) FindReaderIndex(r io.RuneReader) (loc []int)
func (re *Regexp) FindReaderSubmatchIndex(r io.RuneReader) []int
func (re *Regexp) FindString(s string) string
func (re *Regexp) FindStringIndex(s string) (loc []int)
func (re *Regexp) FindStringSubmatch(s string) []string
func (re *Regexp) FindStringSubmatchIndex(s string) []int
func (re *Regexp) FindSubmatch(b []byte) [][]byte
func (re *Regexp) FindSubmatchIndex(b []byte) []int

仔细看会发现,这 18 个方法其实是同一批能力针对不同输入源(字节切片、字符串、io.RuneReader)的重复展开。忽略输入源差异后,核心只有 8 种:

func (re *Regexp) Find(b []byte) []byte
func (re *Regexp) FindAll(b []byte, n int) [][]byte
func (re *Regexp) FindAllIndex(b []byte, n int) [][]int
func (re *Regexp) FindAllSubmatch(b []byte, n int) [][][]byte
func (re *Regexp) FindAllSubmatchIndex(b []byte, n int) [][]int
func (re *Regexp) FindIndex(b []byte) (loc []int)
func (re *Regexp) FindSubmatch(b []byte) [][]byte
func (re *Regexp) FindSubmatchIndex(b []byte) []int

方法命名遵循固定规则:Find 返回匹配内容本身;FindIndex 返回匹配区间的起止下标;Submatch 额外返回括号分组捕获的内容;All 返回全部匹配而非第一个;n 参数控制返回数量——n < 0 返回所有匹配,n > 0 表示最多返回前 n 个。

完整示例:逐一验证 Find 系列行为

package main

import (
	"fmt"
	"regexp"
)

func main() {
	a := "I am learning Go language"

	re, _ := regexp.Compile("[a-z]{2,4}")

	// Find the first match.
	one := re.Find([]byte(a))
	fmt.Println("Find:", string(one))

	// Find all matches and save to a slice, n less than 0 means return all matches, indicates length of slice if it's greater than 0.
	all := re.FindAll([]byte(a), -1)
	fmt.Println("FindAll", all)

	// Find index of first match, start and end position.
	index := re.FindIndex([]byte(a))
	fmt.Println("FindIndex", index)

	// Find index of all matches, the n does same job as above.
	allindex := re.FindAllIndex([]byte(a), -1)
	fmt.Println("FindAllIndex", allindex)

	re2, _ := regexp.Compile("am(.*)lang(.*)")

	// Find first submatch and return array, the first element contains all elements, the second element contains the result of first (), the third element contains the result of second ().
	// Output:
	// the first element: "am learning Go language"
	// the second element: " learning Go ", notice spaces will be outputed as well.
	// the third element: "uage"
	submatch := re2.FindSubmatch([]byte(a))
	fmt.Println("FindSubmatch", submatch)
	for _, v := range submatch {
		fmt.Println(string(v))
	}

	// Same thing like FindIndex().
	submatchindex := re2.FindSubmatchIndex([]byte(a))
	fmt.Println(submatchindex)

	// FindAllSubmatch, find all submatches.
	submatchall := re2.FindAllSubmatch([]byte(a), -1)
	fmt.Println(submatchall)

	// FindAllSubmatchIndex,find index of all submatches.
	submatchallindex := re2.FindAllSubmatchIndex([]byte(a), -1)
	fmt.Println(submatchallindex)
}

对结果做几点解读:

  • Find 只返回第一个匹配,FindAll(..., -1) 返回全部;
  • FindIndex 返回形如 [start, end] 的起止下标对(半开区间,即不含 end 位置),FindAllIndex 返回所有下标对的二维切片;
  • re2 := regexp.Compile("am(.*)lang(.*)") 中包含两个捕获组。FindSubmatch 返回的切片中:第一个元素是完整匹配 "am learning Go language",第二个元素是第一个括号捕获的 " learning Go "(注意空格也会被捕获输出),第三个元素是第二个括号捕获的 "uage";
  • FindSubmatchIndex 返回各分组的起止下标;FindAllSubmatch / FindAllSubmatchIndex 则是其"全部匹配"版本,n 的作用与 FindAll 一致。

这些方法在爬虫、日志解析、数据抽取场景中是最常用的工具。

Match 方法:包级函数的底层实现

如前所述,regexp 包还有 3 个实例方法用于匹配,它们与包级函数做的是完全相同的事——事实上,包级导出函数在底层就是调用这些方法:

func (re *Regexp) Match(b []byte) bool
func (re *Regexp) MatchReader(r io.RuneReader) bool
func (re *Regexp) MatchString(s string) bool

与包级函数相比,实例方法少了一个"编译"环节(编译已在 Compile 时完成),因此同一 *Regexp 在循环中反复匹配时性能更好;且返回只有 bool,不携带语法错误(错误在编译期就已经暴露了)。在需要频繁校验的场景(比如逐条校验请求参数),先用 MustCompile 编译一次,再循环调用 MatchString 是更高效、更安全的写法。

ReplaceAll 家族:替换与清洗

接下来看正则替换方法:

func (re *Regexp) ReplaceAll(src, repl []byte) []byte
func (re *Regexp) ReplaceAllFunc(src []byte, repl func([]byte) []byte) []byte
func (re *Regexp) ReplaceAllLiteral(src, repl []byte) []byte
func (re *Regexp) ReplaceAllLiteralString(src, repl string) string
func (re *Regexp) ReplaceAllString(src, repl string) string
func (re *Regexp) ReplaceAllStringFunc(src string, repl func(string) string) string

用法要点:

  • ReplaceAllString / ReplaceAll:把匹配到的内容替换为固定文本(如爬虫示例中把标签替换为 "" 或 "\n");
  • ReplaceAllStringFunc / ReplaceAllFunc:替换文本由回调函数动态生成,回调接收每次匹配到的原文、返回替换结果。爬虫示例第一步就是用 ReplaceAllStringFunc(src, strings.ToLower) 把每个标签转成小写;
  • ReplaceAllLiteralString / ReplaceAllLiteral:字面替换版本,替换文本中的 $ 不会被解析为分组引用,适合替换内容里恰好含 $ 字符的场景(如模板片段)。

这些方法在上面的爬虫示例中已经完整演示过(删除 style/script、标签转小写、连续空白压缩),这里不再重复。

Expand:用命名分组重组文本

Expand 与 ExpandString 用于基于捕获组索引和模板进行文本重建:

func (re *Regexp) Expand(dst []byte, template []byte, src []byte, match []int) []byte
func (re *Regexp) ExpandString(dst []byte, template string, src string, match []int) []byte

Expand 的四个参数依次是:目标缓冲区 dst(结果会追加到其后)、模板 template(可引用 $name 或 $1 形式的分组名)、原始输入 src、以及由 FindAllSubmatchIndex 得到的分组下标 match。它通常与命名分组((?P<name>...))配合使用。

看下面的例子——把 call hello alice 这类命令文本改写成函数调用形式:

func main() {
	src := []byte(`
		call hello alice
		hello bob
		call hello eve
	`)
	pat := regexp.MustCompile(`(?m)(call)\s+(?P<cmd>\w+)\s+(?P<arg>.+)\s*$`)
	res := []byte{}
	for _, s := range pat.FindAllSubmatchIndex(src, -1) {
		res = pat.Expand(res, []byte("$cmd('$arg')\n"), src, s)
	}
	fmt.Println(string(res))
}

逐段拆解:

  • 模式 (?m)(call)\s+(?P<cmd>\w+)\s+(?P<arg>.+)\s*$:(?m) 开启多行模式,使 ^/$ 匹配每行边界;(?P<cmd>...) 与 (?P<arg>...) 是命名捕获组,分别捕获命令名与参数;
  • FindAllSubmatchIndex(src, -1) 找出所有匹配行及其分组下标;
  • Expand 按照模板 "$cmd('$arg')\n" 将每个匹配重建为 hello('alice') 形式的调用语句,并不断追加到 res 缓冲区;
  • 由于模式要求行首有 call,hello bob 这一行不会被匹配,最终输出只包含两行改写后的调用。

Expand 适合日志格式化、协议报文重组、代码生成等"提取后再组装"的批处理场景。

小结:regexp 包的完整工具链

至此,Go regexp 包的核心能力已经完整覆盖:匹配验证(包级 Match 三函数与实例 Match 三方法)、编译入口(Compile / CompilePOSIX / MustCompile / MustCompilePOSIX,注意 POSIX 的最左最长语义差异)、查找抽取(Find 系列 18 个方法,可归纳为 8 种核心能力)、替换清洗(ReplaceAll 系列,含回调与字面量版本)、重组生成(Expand / ExpandString 配合命名分组)。

在实际 Web 项目中,建议遵循本节确立的选型顺序:strings 包能解决的琐碎操作优先用 strings;需要大小写不敏感匹配、分组捕获、批量过滤时才引入 regexp;正式代码中尽量用 MustCompile 在初始化阶段完成编译(模式错误直接 panic 暴露),业务循环内只调用 MatchString / FindAll 等方法以获得最佳性能。表单验证(参见 de/code/src/apps/ch.4.2/validator/main.go)与 HTML 清洗爬虫就是这两个方向的典型实践,你可以在此基础上继续探索更多用法。

本章为第 7 章"文本文件处理"(de/07.0.md)的组成部分,前后衔接紧密:上一节介绍了 JSON 处理,下一节将进入 模板引擎 Templates;完整目录见 de/preface.md。

登录后查看全文
build-web-application-with-golang