Go 正则表达式实战:深入解析 Build Web Application with Golang 第 7.3 节 regexp 包
Go 正则表达式实战:深入解析 Build Web Application with Golang 第 7.3 节 regexp 包
本篇技术指南以开源项目 Build Web Application with Golang(de/07.3.md)中第 7.3 节"Regexp"为骨架,系统讲解 Go 标准库 regexp 包在 Web 开发中的应用:从 Match 系列函数做输入校验,到 Compile 复杂模式完成抓取、过滤与文本清洗,再到 Find、ReplaceAll、Expand 等方法的完整用法。读完本文,你将掌握 Go 正则表达式的 RE2 语法约束、三类匹配入口、18 个查找方法的选型,以及如何在真实表单校验与爬虫场景中落地这些能力。
为什么 Web 开发需要正则表达式
正则表达式(Regular Expression)是一种强大但也复杂的模式匹配与文本处理工具。它的执行性能虽然不如纯文本匹配,却极其灵活:只要给出合适的语法,几乎可以从任意源文本中过滤出你想要的内容。在 Web 开发中收集数据时,用正则表达式提取有意义的信息几乎是常规操作。
Go 在标准库中提供了 regexp 包作为官方支持。如果你在其他语言中用过正则表达式,上手会非常容易。这里需要特别留意一个标准层面的差异:Go 实现的是 RE2 标准,除了 \C 之外遵循 RE2 的完整语法约定。RE2 是 Google 设计的一种线性时间复杂度的正则引擎,它刻意不支持回溯(backtracking),因此不存在灾难性回溯(catastrophic backtracking)导致的服务阻塞问题,这是 Go 正则适合服务端处理用户输入的重要原因。
另外要明确一个选型原则:strings 包其实能完成许多工作,比如搜索(Contains、Index)、替换(Replace)、解析(Split、Join)等,而且速度比正则更快。但这些都属于"琐碎"操作;当你需要大小写不敏感搜索等更高级的行为时,正则才是最佳选择。结论很简单:strings 能满足就用 strings(易用、易读、更快);需要更高级操作再上 regexp。
如果你还记得前面章节的表单验证(第 4 章),我们当时就是用正则来校验用户输入的有效性,并且要意识到所有字符都是 UTF-8 编码的。现在就来深入学习 Go 的 regexp 包。
Match:三函数快速匹配验证
regexp 包提供了 3 个包级函数用于匹配判断:命中模式返回 true,否则返回 false。
func Match(pattern string, b []byte) (matched bool, error error)
func MatchReader(pattern string, r io.RuneReader) (matched bool, error error)
func MatchString(pattern string, s string) (matched bool, error error)
这三个函数做的事情完全一致——检查 pattern 是否匹配输入源,匹配返回 true;唯一的差异在于输入源的三种类型:字节切片([]byte)、io.RuneReader(以 rune 为单位流式读取的 reader)和字符串(string)。注意:如果正则本身存在语法错误,函数会返回 error,因此实际使用时通常用 m, _ := regexp.MatchString(...) 或显式处理第二个返回值。
示例一:校验 IP 地址
下面用 MatchString 校验一个 IPv4 格式的字符串:
func IsIP(ip string) (b bool) {
if m, _ := regexp.MatchString("^[0-9]{1,3}\\.[0-9]{1,3}\\.[0-9]{1,3}\\.[0-9]{1,3}$", ip); !m {
return false
}
return true
}
模式 ^[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}\.[0-9]{1,3}$ 的含义是:从开头(^)到结尾($)必须是由点号分隔的四段数字,每段 1~3 位([0-9]{1,3}),点号需要转义为 \.。注意该模式只校验"形如 IPv4",并不校验每段是否在 0~255 范围内,如需严格校验可进一步改写。
示例二:校验命令行输入是否为数字
func main() {
if len(os.Args) == 1 {
fmt.Println("Usage: regexp [string]")
os.Exit(1)
} else if m, _ := regexp.MatchString("^[0-9]+$", os.Args[1]); m {
fmt.Println("Number")
} else {
fmt.Println("Not number")
}
}
这里通过 ^[0-9]+$ 判断用户传入的第一个命令行参数是否全部由数字组成。运行方式:go run main.go 123 输出 Number;go run main.go abc 输出 Not number;不带参数则打印用法提示并退出。
以上 Match/MatchReader/MatchString 只解决"是否匹配"的验证问题,使用起来都很简单。
仓库印证:表单验证器中的正则实践
第 4 章的表单验证示例正是这一模式的真实落地。在 validator 源码 中,checkChineseName 用正则校验中文姓名只能包含中文字符:
// Checks if all the characters are chinese characters. Won't check if empty.'
func checkChineseName(str string) error {
if str != "" {
if m, _ := regexp.MatchString("^[\\x{4e00}-\\x{9fa5}]+$", strings.Trim(str, " ")); !m {
return errors.New("Please make sure that the chinese name only contains chinese characters.")
}
}
return nil
}
其中 \x{4e00}-\x{9fa5} 是 UTF-8 编码下的中文 Unicode 码点区间,这正是文档强调"所有字符都是 UTF-8"的实际体现。同文件里的 checkEmail 也借助正则做基础格式校验:
func checkEmail(str string) error {
if m, err := regexp.MatchString(`^[^@]+@[^@]+$`, str); !m {
fmt.Println("err = ", err)
return errors.New("Please enter a valid email address.")
}
return nil
}
这两个校验函数分别挂在 stringValidator 映射表(de/code/src/apps/ch.4.2/validator/main.go#L50-L59)的 chineseName 与 email 键上,通过 GetErrors() 在 ch.4.2 主程序 中被 checkProfile 调用。你可以运行 go run main.go(在 de/code/src/apps/ch.4.2 目录下,并将 GOPATH 指向代码目录,参见 de/code/readme.md)后访问 http://localhost:9090/profile 实际体验。
Filter:进入复杂模式,用 Compile 过滤与切割数据
匹配模式只能"验证"内容,却无法切割、过滤或收集数据。要完成这些任务,必须使用正则的复杂模式(先编译出 *Regexp 对象,再调用其方法)。
爬虫实战:清洗 HTML 页面
假设我们要写一个简单的爬虫。下面的示例演示了何时必须用正则来过滤与切割数据——它抓取百度首页,然后逐步清洗掉 HTML 标签、样式与脚本,最终只保留纯文本:
package main
import (
"fmt"
"io/ioutil"
"net/http"
"regexp"
"strings"
)
func main() {
resp, err := http.Get("http://www.baidu.com")
if err != nil {
fmt.Println("http get error.")
}
defer resp.Body.Close()
body, err := ioutil.ReadAll(resp.Body)
if err != nil {
fmt.Println("http read error")
return
}
src := string(body)
// Convert HTML tags to lower case.
re, _ := regexp.Compile("\\<[\\S\\s]+?\\>")
src = re.ReplaceAllStringFunc(src, strings.ToLower)
// Remove STYLE.
re, _ = regexp.Compile("\\<style[\\S\\s]+?\\</style\\>")
src = re.ReplaceAllString(src, "")
// Remove SCRIPT.
re, _ = regexp.Compile("\\<script[\\S\\s]+?\\</script\\>")
src = re.ReplaceAllString(src, "")
// Remove all HTML code in angle brackets, and replace with newline.
re, _ = regexp.Compile("\\<[\\S\\s]+?\\>")
src = re.ReplaceAllString(src, "\n")
// Remove continuous newline.
re, _ = regexp.Compile("\\s{2,}")
src = re.ReplaceAllString(src, "\n")
fmt.Println(strings.TrimSpace(src))
}
清洗流程共四步,每一步都对应一个正则:
\\<[\\S\\s]+?\\>:匹配尖括号内的标签,配合ReplaceAllStringFunc(src, strings.ToLower)将标签统一转为小写(HTML 标签大小写不敏感,统一小写便于后续匹配);\\<style[\\S\\s]+?\\</style\\>:把<style>...</style>整块删除(替换为空串);\\<script[\\S\\s]+?\\</script\\>:把<script>...</script>整块删除;- 再次用
\\<[\\S\\s]+?\\>把所有剩余标签替换为换行符,然后用\\s{2,}把连续空白压缩为单个换行,最后strings.TrimSpace去掉首尾空白。
其中 [\\S\\s] 表示"任意字符"(\S 非空白 + \s 空白),+? 是非贪婪量词,保证匹配到最短的标签内容。这里 Compile 是复杂模式的第一步:它会校验正则语法是否正确,然后返回一个可复用的 *Regexp 对象供后续操作使用。
Compile 家族:四种编译入口
func Compile(expr string) (*Regexp, error)
func CompilePOSIX(expr string) (*Regexp, error)
func MustCompile(str string) *Regexp
func MustCompilePOSIX(str string) *Regexp
CompilePOSIX 与 Compile 的区别在于匹配语义:前者遵循 POSIX 语法,采用最左最长匹配(leftmost longest search);后者采用最左匹配(leftmost search)。举例说明:对模式 [a-z]{2,4} 匹配内容 "aa09aaa88aaaa" 时,CompilePOSIX 返回 aaaa(取最长的连续匹配),而 Compile 返回 aa(取最左的第一个匹配)。带 Must 前缀的版本在语法错误时会直接 panic,否则返回 *Regexp,适用于模式在编写期就确定的场景(例如上面 Expand 示例中的 MustCompile)。
Find 家族:18 个方法背后的 8 种核心能力
有了 *Regexp 对象,我们就可以用它操作内容了。regexp 包提供如下 18 个查找方法:
func (re *Regexp) Find(b []byte) []byte
func (re *Regexp) FindAll(b []byte, n int) [][]byte
func (re *Regexp) FindAllIndex(b []byte, n int) [][]int
func (re *Regexp) FindAllString(s string, n int) []string
func (re *Regexp) FindAllStringIndex(s string, n int) [][]int
func (re *Regexp) FindAllStringSubmatch(s string, n int) [][]string
func (re *Regexp) FindAllStringSubmatchIndex(s string, n int) [][]int
func (re *Regexp) FindAllSubmatch(b []byte, n int) [][][]byte
func (re *Regexp) FindAllSubmatchIndex(b []byte, n int) [][]int
func (re *Regexp) FindIndex(b []byte) (loc []int)
func (re *Regexp) FindReaderIndex(r io.RuneReader) (loc []int)
func (re *Regexp) FindReaderSubmatchIndex(r io.RuneReader) []int
func (re *Regexp) FindString(s string) string
func (re *Regexp) FindStringIndex(s string) (loc []int)
func (re *Regexp) FindStringSubmatch(s string) []string
func (re *Regexp) FindStringSubmatchIndex(s string) []int
func (re *Regexp) FindSubmatch(b []byte) [][]byte
func (re *Regexp) FindSubmatchIndex(b []byte) []int
仔细看会发现,这 18 个方法其实是同一批能力针对不同输入源(字节切片、字符串、io.RuneReader)的重复展开。忽略输入源差异后,核心只有 8 种:
func (re *Regexp) Find(b []byte) []byte
func (re *Regexp) FindAll(b []byte, n int) [][]byte
func (re *Regexp) FindAllIndex(b []byte, n int) [][]int
func (re *Regexp) FindAllSubmatch(b []byte, n int) [][][]byte
func (re *Regexp) FindAllSubmatchIndex(b []byte, n int) [][]int
func (re *Regexp) FindIndex(b []byte) (loc []int)
func (re *Regexp) FindSubmatch(b []byte) [][]byte
func (re *Regexp) FindSubmatchIndex(b []byte) []int
方法命名遵循固定规则:Find 返回匹配内容本身;FindIndex 返回匹配区间的起止下标;Submatch 额外返回括号分组捕获的内容;All 返回全部匹配而非第一个;n 参数控制返回数量——n < 0 返回所有匹配,n > 0 表示最多返回前 n 个。
完整示例:逐一验证 Find 系列行为
package main
import (
"fmt"
"regexp"
)
func main() {
a := "I am learning Go language"
re, _ := regexp.Compile("[a-z]{2,4}")
// Find the first match.
one := re.Find([]byte(a))
fmt.Println("Find:", string(one))
// Find all matches and save to a slice, n less than 0 means return all matches, indicates length of slice if it's greater than 0.
all := re.FindAll([]byte(a), -1)
fmt.Println("FindAll", all)
// Find index of first match, start and end position.
index := re.FindIndex([]byte(a))
fmt.Println("FindIndex", index)
// Find index of all matches, the n does same job as above.
allindex := re.FindAllIndex([]byte(a), -1)
fmt.Println("FindAllIndex", allindex)
re2, _ := regexp.Compile("am(.*)lang(.*)")
// Find first submatch and return array, the first element contains all elements, the second element contains the result of first (), the third element contains the result of second ().
// Output:
// the first element: "am learning Go language"
// the second element: " learning Go ", notice spaces will be outputed as well.
// the third element: "uage"
submatch := re2.FindSubmatch([]byte(a))
fmt.Println("FindSubmatch", submatch)
for _, v := range submatch {
fmt.Println(string(v))
}
// Same thing like FindIndex().
submatchindex := re2.FindSubmatchIndex([]byte(a))
fmt.Println(submatchindex)
// FindAllSubmatch, find all submatches.
submatchall := re2.FindAllSubmatch([]byte(a), -1)
fmt.Println(submatchall)
// FindAllSubmatchIndex,find index of all submatches.
submatchallindex := re2.FindAllSubmatchIndex([]byte(a), -1)
fmt.Println(submatchallindex)
}
对结果做几点解读:
Find只返回第一个匹配,FindAll(..., -1)返回全部;FindIndex返回形如[start, end]的起止下标对(半开区间,即不含end位置),FindAllIndex返回所有下标对的二维切片;re2 := regexp.Compile("am(.*)lang(.*)")中包含两个捕获组。FindSubmatch返回的切片中:第一个元素是完整匹配"am learning Go language",第二个元素是第一个括号捕获的" learning Go "(注意空格也会被捕获输出),第三个元素是第二个括号捕获的"uage";FindSubmatchIndex返回各分组的起止下标;FindAllSubmatch/FindAllSubmatchIndex则是其"全部匹配"版本,n的作用与FindAll一致。
这些方法在爬虫、日志解析、数据抽取场景中是最常用的工具。
Match 方法:包级函数的底层实现
如前所述,regexp 包还有 3 个实例方法用于匹配,它们与包级函数做的是完全相同的事——事实上,包级导出函数在底层就是调用这些方法:
func (re *Regexp) Match(b []byte) bool
func (re *Regexp) MatchReader(r io.RuneReader) bool
func (re *Regexp) MatchString(s string) bool
与包级函数相比,实例方法少了一个"编译"环节(编译已在 Compile 时完成),因此同一 *Regexp 在循环中反复匹配时性能更好;且返回只有 bool,不携带语法错误(错误在编译期就已经暴露了)。在需要频繁校验的场景(比如逐条校验请求参数),先用 MustCompile 编译一次,再循环调用 MatchString 是更高效、更安全的写法。
ReplaceAll 家族:替换与清洗
接下来看正则替换方法:
func (re *Regexp) ReplaceAll(src, repl []byte) []byte
func (re *Regexp) ReplaceAllFunc(src []byte, repl func([]byte) []byte) []byte
func (re *Regexp) ReplaceAllLiteral(src, repl []byte) []byte
func (re *Regexp) ReplaceAllLiteralString(src, repl string) string
func (re *Regexp) ReplaceAllString(src, repl string) string
func (re *Regexp) ReplaceAllStringFunc(src string, repl func(string) string) string
用法要点:
ReplaceAllString/ReplaceAll:把匹配到的内容替换为固定文本(如爬虫示例中把标签替换为""或"\n");ReplaceAllStringFunc/ReplaceAllFunc:替换文本由回调函数动态生成,回调接收每次匹配到的原文、返回替换结果。爬虫示例第一步就是用ReplaceAllStringFunc(src, strings.ToLower)把每个标签转成小写;ReplaceAllLiteralString/ReplaceAllLiteral:字面替换版本,替换文本中的$不会被解析为分组引用,适合替换内容里恰好含$字符的场景(如模板片段)。
这些方法在上面的爬虫示例中已经完整演示过(删除 style/script、标签转小写、连续空白压缩),这里不再重复。
Expand:用命名分组重组文本
Expand 与 ExpandString 用于基于捕获组索引和模板进行文本重建:
func (re *Regexp) Expand(dst []byte, template []byte, src []byte, match []int) []byte
func (re *Regexp) ExpandString(dst []byte, template string, src string, match []int) []byte
Expand 的四个参数依次是:目标缓冲区 dst(结果会追加到其后)、模板 template(可引用 $name 或 $1 形式的分组名)、原始输入 src、以及由 FindAllSubmatchIndex 得到的分组下标 match。它通常与命名分组((?P<name>...))配合使用。
看下面的例子——把 call hello alice 这类命令文本改写成函数调用形式:
func main() {
src := []byte(`
call hello alice
hello bob
call hello eve
`)
pat := regexp.MustCompile(`(?m)(call)\s+(?P<cmd>\w+)\s+(?P<arg>.+)\s*$`)
res := []byte{}
for _, s := range pat.FindAllSubmatchIndex(src, -1) {
res = pat.Expand(res, []byte("$cmd('$arg')\n"), src, s)
}
fmt.Println(string(res))
}
逐段拆解:
- 模式
(?m)(call)\s+(?P<cmd>\w+)\s+(?P<arg>.+)\s*$:(?m)开启多行模式,使^/$匹配每行边界;(?P<cmd>...)与(?P<arg>...)是命名捕获组,分别捕获命令名与参数; FindAllSubmatchIndex(src, -1)找出所有匹配行及其分组下标;Expand按照模板"$cmd('$arg')\n"将每个匹配重建为hello('alice')形式的调用语句,并不断追加到res缓冲区;- 由于模式要求行首有
call,hello bob这一行不会被匹配,最终输出只包含两行改写后的调用。
Expand 适合日志格式化、协议报文重组、代码生成等"提取后再组装"的批处理场景。
小结:regexp 包的完整工具链
至此,Go regexp 包的核心能力已经完整覆盖:匹配验证(包级 Match 三函数与实例 Match 三方法)、编译入口(Compile / CompilePOSIX / MustCompile / MustCompilePOSIX,注意 POSIX 的最左最长语义差异)、查找抽取(Find 系列 18 个方法,可归纳为 8 种核心能力)、替换清洗(ReplaceAll 系列,含回调与字面量版本)、重组生成(Expand / ExpandString 配合命名分组)。
在实际 Web 项目中,建议遵循本节确立的选型顺序:strings 包能解决的琐碎操作优先用 strings;需要大小写不敏感匹配、分组捕获、批量过滤时才引入 regexp;正式代码中尽量用 MustCompile 在初始化阶段完成编译(模式错误直接 panic 暴露),业务循环内只调用 MatchString / FindAll 等方法以获得最佳性能。表单验证(参见 de/code/src/apps/ch.4.2/validator/main.go)与 HTML 清洗爬虫就是这两个方向的典型实践,你可以在此基础上继续探索更多用法。
本章为第 7 章"文本文件处理"(de/07.0.md)的组成部分,前后衔接紧密:上一节介绍了 JSON 处理,下一节将进入 模板引擎 Templates;完整目录见 de/preface.md。