在 C# 里搞正则表达式,Regex 类是老朋友了------Regex.IsMatch()、Regex.Match()、Regex.Replace(),编译选项、捕获组、命名分组一应俱全。写起来规规矩矩,强类型加持,IntelliSense 随时待命,稳得像老司机开车。
在 Python 里搞正则表达式,内置 re 模块就是你的武器------re.match()、re.search()、re.findall()、re.sub(),函数式风格,几行代码搞定一切。没有对象创建的仪式感,直接上手就是干。
两者底层都是 PCRE 风格的正则引擎,语法几乎一模一样。但 API 风格差异很大:C# 是面向对象的 Regex 类 ,Python 是函数式的 re 模块。就像同样是切菜,一个用菜刀,一个用料理机------结果一样,姿势不同。
同样是正则表达式,一个面向对象,一个函数式------姿势不同,殊途同归。
基础语法对比
先来个最基础的:匹配、提取、替换三连。
C# 版本:
using System.Text.RegularExpressions;
// 匹配:字符串里有没有数字?
bool isMatch = Regex.IsMatch("hello123", @"\d+"); // true
// 提取:把数字和字母分开
Match match = Regex.Match("hello123 world", @"(\w+)(\d+)");
if (match.Success)
{
Console.WriteLine(match.Groups[1].Value); // hello
Console.WriteLine(match.Groups[2].Value); // 123
}
// 替换:把数字换成 NUM
string result = Regex.Replace("hello123", @"\d+", "NUM");
Console.WriteLine(result); // helloNUM
Python 版本:
import re
# 匹配:字符串里有没有数字?
is_match = bool(re.search(r"\d+", "hello123")) # True
# 提取:把数字和字母分开
match = re.search(r"(\w+)(\d+)", "hello123 world")
if match:
print(match.group(1)) # hello
print(match.group(2)) # 123
# 替换:把数字换成 NUM
result = re.sub(r"\d+", "NUM", "hello123")
print(result) # helloNUM
对比一下:
| 操作 | C# | Python |
|---|---|---|
| 匹配(是否存在) | Regex.IsMatch() |
bool(re.search()) |
| 提取(第一个) | Regex.Match() |
re.search() |
| 全局查找 | Regex.Matches() |
re.findall() / re.finditer() |
| 替换 | Regex.Replace() |
re.sub() |
| 编译 | new Regex(pattern) |
re.compile(pattern) |
| 分割 | Regex.Split() |
re.split() |
一句话记住:C# 是"类.方法(字符串, 模式)",Python 是"re.方法(模式, 字符串)"------参数顺序都不一样,迁移的时候别搞混了。
match vs search------最容易踩的坑
很多从 C# 转 Python 的同学,第一脚就踩在 re.match() 和 re.search() 的区别上。
C# 版本:
// Regex.Match() 会扫描整个字符串,找到第一个匹配
Match match1 = Regex.Match("abc123def", @"\d+");
Console.WriteLine(match1.Value); // 123 ------ 找到了!
Match match2 = Regex.Match("123abc", @"\d+");
Console.WriteLine(match2.Value); // 123 ------ 也找到了!
Python 版本:
# re.match() 只匹配字符串开头!
match1 = re.match(r"\d+", "abc123")
print(match1) # None ------ 开头不是数字,直接放弃了
match2 = re.match(r"\d+", "123abc")
print(match2.group()) # 123 ------ 开头是数字,OK
# re.search() 扫描整个字符串,和 C# 的 Regex.Match() 一样
match3 = re.search(r"\d+", "abc123def")
print(match3.group()) # 123 ------ 找到了!
C# 的 Regex.Match() 等价于 Python 的 re.search(),不等于 re.match()! 这是最常见的坑,没有之一。
全局查找对比:
// C# 全局查找
MatchCollection matches = Regex.Matches("abc123def456ghi789", @"\d+");
foreach (Match m in matches)
{
Console.Write($"{m.Value} "); // 123 456 789
}
# Python 全局查找
matches = re.findall(r"\d+", "abc123def456ghi789")
print(matches) # ['123', '456', '789']
# 带捕获组的 findall ------ 注意返回的是元组列表!
matches = re.findall(r"(\w+)(\d+)", "hello123 world456")
print(matches) # [('hello', '123'), ('world', '456')]
# 用 finditer 获取 Match 对象(类似 C# 的 Matches)
for m in re.finditer(r"\d+", "abc123def456"):
print(f"{m.group()} at position {m.start()}-{m.end()}")
# 123 at position 3-6
# 456 at position 9-12
常用模式
写正则最头疼的不是语法,而是"那个验证手机号的正则怎么写来着"。来,直接上模板。
邮箱验证
C# 版本:
// 邮箱验证
string emailPattern = @"^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$";
Console.WriteLine(Regex.IsMatch("user@example.com", emailPattern)); // true
Console.WriteLine(Regex.IsMatch("invalid-email", emailPattern)); // false
Console.WriteLine(Regex.IsMatch("test.name+tag@sub.domain.org", emailPattern)); // true
Python 版本:
import re
email_pattern = r"^[a-zA-Z0-9._%+-]+@[a-zA-Z0-9.-]+\.[a-zA-Z]{2,}$"
print(bool(re.match(email_pattern, "user@example.com"))) # True
print(bool(re.match(email_pattern, "invalid-email"))) # False
print(bool(re.match(email_pattern, "test.name+tag@sub.domain.org"))) # True
手机号验证(中国)
C# 版本:
// 中国大陆手机号:1开头,第二位3-9,后面9位数字
string phonePattern = @"^1[3-9]\d{9}$";
Console.WriteLine(Regex.IsMatch("13812345678", phonePattern)); // true
Console.WriteLine(Regex.IsMatch("12345678", phonePattern)); // false
Console.WriteLine(Regex.IsMatch("23812345678", phonePattern)); // false
Python 版本:
phone_pattern = r"^1[3-9]\d{9}$"
print(bool(re.match(phone_pattern, "13812345678"))) # True
print(bool(re.match(phone_pattern, "12345678"))) # False
print(bool(re.match(phone_pattern, "23812345678"))) # False
URL 验证
C# 版本:
// URL 验证(简化版)
string urlPattern = @"^https?://[\w\-]+(\.[\w\-]+)+[\w\-.,@?^=%&:/~+#]*$";
Console.WriteLine(Regex.IsMatch("https://example.com/path?q=1", urlPattern)); // true
Console.WriteLine(Regex.IsMatch("http://sub.domain.org/api", urlPattern)); // true
Console.WriteLine(Regex.IsMatch("ftp://files.example.com", urlPattern)); // false
Python 版本:
url_pattern = r"^https?://[\w\-]+(\.[\w\-]+)+[\w\-.,@?^=%&:/~+#]*$"
print(bool(re.match(url_pattern, "https://example.com/path?q=1"))) # True
print(bool(re.match(url_pattern, "http://sub.domain.org/api"))) # True
print(bool(re.match(url_pattern, "ftp://files.example.com"))) # False
提取日期
C# 版本:
// 提取 yyyy-MM-dd 格式的日期
string datePattern = @"(\d{4})-(\d{2})-(\d{2})";
Match match = Regex.Match("今天是 2024-01-15,明天是 2024-01-16", datePattern);
Console.WriteLine(match.Groups[0].Value); // 2024-01-15(完整匹配)
Console.WriteLine(match.Groups[1].Value); // 2024
Console.WriteLine(match.Groups[2].Value); // 01
Console.WriteLine(match.Groups[3].Value); // 15
// 全部提取
MatchCollection dates = Regex.Matches("今天是 2024-01-15,明天是 2024-01-16", datePattern);
foreach (Match d in dates)
{
Console.Write($"{d.Value} "); // 2024-01-15 2024-01-16
}
Python 版本:
date_pattern = r"(\d{4})-(\d{2})-(\d{2})"
match = re.search(date_pattern, "今天是 2024-01-15,明天是 2024-01-16")
print(match.group(0)) # 2024-01-15(完整匹配)
print(match.group(1)) # 2024
print(match.group(2)) # 01
print(match.group(3)) # 15
# 全部提取
dates = re.findall(date_pattern, "今天是 2024-01-15,明天是 2024-01-16")
print(dates) # [('2024', '01', '15'), ('2024', '01', '16')]
# 如果只要完整匹配
dates = [m.group() for m in re.finditer(date_pattern, "今天是 2024-01-15,明天是 2024-01-16")]
print(dates) # ['2024-01-15', '2024-01-16']
分组与捕获
分组是正则表达式的灵魂------没有分组,你只能知道"匹配了",但不知道"匹配了啥"。
命名分组
C# 用 (?<name>...) 命名,Python 用 (?P<name>...) 命名。
C# 版本:
var regex = new Regex(@"(?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})");
Match match = regex.Match("生日:2024-01-15,入职:2020-06-20");
if (match.Success)
{
Console.WriteLine($"年: {match.Groups["year"].Value}"); // 2024
Console.WriteLine($"月: {match.Groups["month"].Value}"); // 01
Console.WriteLine($"日: {match.Groups["day"].Value}"); // 15
}
// 获取所有命名分组
foreach (string name in match.Groups.Keys)
{
if (name != "0") // 跳过整个匹配
Console.WriteLine($"{name} = {match.Groups[name].Value}");
}
Python 版本:
pattern = re.compile(r"(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})")
match = pattern.search("生日:2024-01-15,入职:2020-06-20")
if match:
print(f"年: {match.group('year')}") # 2024
print(f"月: {match.group('month')}") # 01
print(f"日: {match.group('day')}") # 15
# groupdict() 一次拿到所有命名分组------Python 的杀手锏
info = match.groupdict()
print(info) # {'year': '2024', 'month': '01', 'day': '15'}
Python 的
groupdict()比 C# 的Groups.Keys遍历方便多了------一行代码拿到字典,直接用。
非捕获分组
有时候你只需要分组,不需要捕获内容------用 (?:...)。
C# 版本:
// 非捕获分组:(?:...)
var regex = new Regex(@"(?:https?|ftp)://([\w\-]+\.[\w\-]+)");
Match match = regex.Match("访问 https://example.com/path");
Console.WriteLine(match.Groups[0].Value); // https://example.com(完整匹配)
Console.WriteLine(match.Groups[1].Value); // example.com(只有捕获组)
// 注意:没有 Groups[2],(?:...) 不产生捕获组
Python 版本:
# 非捕获分组:(?:...)
pattern = re.compile(r"(?:https?|ftp)://([\w\-]+\.[\w\-]+)")
match = pattern.search("访问 https://example.com/path")
print(match.group(0)) # https://example.com(完整匹配)
print(match.group(1)) # example.com(只有捕获组)
# (?:...) 不产生捕获组,不会干扰编号
反向引用
反向引用让你匹配"前面出现过的内容"------比如找重复的单词。
C# 版本:
// 找重复的单词
var regex = new Regex(@"\b(\w+)\s+\1\b", RegexOptions.IgnoreCase);
MatchCollection matches = regex.Matches("this is is a test test string");
foreach (Match m in matches)
{
Console.WriteLine($"重复词: {m.Value}"); // is is, test test
}
// 命名反向引用
var namedRegex = new Regex(@"(?<word>\w+)\s+\k<word>\b");
// 和上面效果一样,但用了命名反向引用 \k<word>
Python 版本:
# 找重复的单词
pattern = re.compile(r"\b(\w+)\s+\1\b", re.IGNORECASE)
matches = pattern.findall("this is is a test test string")
print(matches) # ['is', 'test']
# 命名反向引用
named_pattern = re.compile(r"(?P<word>\w+)\s+(?P=word)\b")
matches = named_pattern.findall("this is is a test test string")
print(matches) # ['is', 'test']
# 用 finditer 获取完整匹配信息
for m in re.finditer(r"\b(\w+)\s+\1\b", "this is is a test test string", re.IGNORECASE):
print(f"重复: '{m.group(1)}' 在位置 {m.start()}-{m.end()}")
C# 的反向引用用
\k<name>,Python 用(?P=name)------写法不一样,但功能相同。
替换与分割
正则替换是批量处理文本的利器------比手动写循环优雅一万倍。
基本替换
C# 版本:
// 基本替换
string result1 = Regex.Replace("hello123world456", @"\d+", "#");
Console.WriteLine(result1); // hello#world#
// 带捕获组的替换(用 $1、$2 引用)
string result2 = Regex.Replace("2024-01-15", @"(\d{4})-(\d{2})-(\d{2})", "$3/$2/$1");
Console.WriteLine(result2); // 15/01/2024
// 命名分组替换
string result3 = Regex.Replace("2024-01-15",
@"(?<year>\d{4})-(?<month>\d{2})-(?<day>\d{2})",
"${day}/${month}/${year}");
Console.WriteLine(result3); // 15/01/2024
Python 版本:
# 基本替换
result1 = re.sub(r"\d+", "#", "hello123world456")
print(result1) # hello#world#
# 带捕获组的替换(用 \1、\2 引用)
result2 = re.sub(r"(\d{4})-(\d{2})-(\d{2})", r"\3/\2/\1", "2024-01-15")
print(result2) # 15/01/2024
# 命名分组替换
result3 = re.sub(
r"(?P<year>\d{4})-(?P<month>\d{2})-(?P<day>\d{2})",
r"\g<day>/\g<month>/\g<year>",
"2024-01-15"
)
print(result3) # 15/01/2024
用函数替换------高级玩法
有时候替换逻辑比较复杂,传一个函数进去。
C# 版本:
// 用 MatchEvaluator 进行复杂替换
string input = "价格:100元、200元、300元";
string result = Regex.Replace(input, @"(\d+)元", match =>
{
int price = int.Parse(match.Groups[1].Value);
return $"${price * 7}"; // 人民币转美元(假设汇率 7)
});
Console.WriteLine(result); // 价格:$700、$1400、$2100
Python 版本:
# 用函数替换
def convert_to_usd(match):
price = int(match.group(1))
return f"${price * 7}" # 人民币转美元(假设汇率 7)
result = re.sub(r"(\d+)元", convert_to_usd, "价格:100元、200元、300元")
print(result) # 价格:$700、$1400、$2100
Python 的函数替换更简洁------直接传函数名,不需要像 C# 那样写 lambda(当然也可以写 lambda)。
分割
C# 版本:
// 按多种分隔符分割
string input = "hello, world; python csharp";
string[] parts = Regex.Split(input, @"[,;\s]+");
foreach (var part in parts)
{
Console.Write($"[{part}] "); // [hello] [world] [python] [csharp]
}
// 带捕获组的分割(保留分隔符)
string input2 = "123abc456def789";
string[] parts2 = Regex.Split(input2, @"(\d+)");
foreach (var part in parts2)
{
Console.Write($"[{part}] "); // [] [123] [abc] [456] [def] [789] []
}
Python 版本:
# 按多种分隔符分割
parts = re.split(r"[,;\s]+", "hello, world; python csharp")
print(parts) # ['hello', 'world', 'python', 'csharp']
# 带捕获组的分割(保留分隔符)
parts2 = re.split(r"(\d+)", "123abc456def789")
print(parts2) # ['', '123', 'abc', '456', 'def', '789', '']
# 限制分割次数
parts3 = re.split(r"[,;\s]+", "hello, world; python csharp", maxsplit=2)
print(parts3) # ['hello', 'world', 'python csharp']
编译与性能
正则编译是一个值得了解的优化点------在循环里反复用同一个正则时特别明显。
C# 版本:
// C# 的 Regex 默认就是编译的
var regex = new Regex(@"\d{4}-\d{2}-\d{2}");
// 显式指定编译(对长期运行的程序有帮助)
var compiled = new Regex(@"\d{4}-\d{2}-\d{2}", RegexOptions.Compiled);
// 在循环中使用编译后的正则
var sw = System.Diagnostics.Stopwatch.StartNew();
for (int i = 0; i < 100_000; i++)
{
compiled.IsMatch("2024-01-15");
}
sw.Stop();
Console.WriteLine($"编译正则耗时: {sw.ElapsedMilliseconds}ms");
// 静态方法 vs 实例方法
// 静态方法(内部缓存,推荐偶尔使用)
bool a = Regex.IsMatch("hello123", @"\d+");
// 实例方法(推荐频繁使用)
var r = new Regex(@"\d+");
bool b = r.IsMatch("hello123");
Python 版本:
import re
import time
# 不编译:每次调用都要重新解析正则
start = time.time()
for i in range(100_000):
re.search(r"\d{4}-\d{2}-\d{2}", "2024-01-15")
print(f"不编译耗时: {time.time() - start:.3f}s")
# 编译:正则只解析一次
pattern = re.compile(r"\d{4}-\d{2}-\d{2}")
start = time.time()
for i in range(100_000):
pattern.search("2024-01-15")
print(f"编译后耗时: {time.time() - start:.3f}s")
# Python 的 re 模块内部也有缓存(LRU,最多 512 个)
# 所以偶尔用几次不编译也没问题
# 但频繁使用(循环、高频函数)建议显式 re.compile()
| 特性 | C# Regex |
Python re |
|---|---|---|
| 默认行为 | 已编译 | 未编译(有缓存) |
| 显式编译 | RegexOptions.Compiled |
re.compile() |
| 缓存机制 | 静态方法内部缓存 | LRU 缓存(512个) |
| 循环使用建议 | 用实例方法 | 用 re.compile() |
| 偶尔用一次 | 静态方法就行 | 直接 re.search() 就行 |
经验之谈 :如果正则只用一次,直接调用
re.search()就行,不用纠结编译不编译。如果要在循环里反复用,先re.compile()再调用,性能提升 10%-30%。C# 的Regex类天然就是编译好的,不用操心。
实际场景
光说不练假把式,来看两个真实场景。
场景一:日志解析
从一堆日志里提取时间、级别和消息。
C# 版本:
var logPattern = @"^(?<time>\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}) \[(?<level>\w+)\] (?<message>.+)$";
string[] logs = {
"2024-01-15 10:30:45 [ERROR] 数据库连接失败",
"2024-01-15 10:30:46 [INFO] 用户登录成功",
"2024-01-15 10:30:47 [WARNING] 磁盘空间不足",
"这是一条没有格式的日志"
};
foreach (var log in logs)
{
Match match = Regex.Match(log, logPattern);
if (match.Success)
{
Console.WriteLine($"[{match.Groups["level"].Value}] {match.Groups["time"].Value}: {match.Groups["message"].Value}");
}
}
// [ERROR] 2024-01-15 10:30:45: 数据库连接失败
// [INFO] 2024-01-15 10:30:46: 用户登录成功
// [WARNING] 2024-01-15 10:30:47: 磁盘空间不足
Python 版本:
import re
log_pattern = r"^(?P<time>\d{4}-\d{2}-\d{2} \d{2}:\d{2}:\d{2}) \[(?P<level>\w+)\] (?P<message>.+)$"
logs = [
"2024-01-15 10:30:45 [ERROR] 数据库连接失败",
"2024-01-15 10:30:46 [INFO] 用户登录成功",
"2024-01-15 10:30:47 [WARNING] 磁盘空间不足",
"这是一条没有格式的日志"
]
for log in logs:
match = re.search(log_pattern, log)
if match:
info = match.groupdict()
print(f"[{info['level']}] {info['time']}: {info['message']}")
# [ERROR] 2024-01-15 10:30:45: 数据库连接失败
# [INFO] 2024-01-15 10:30:46: 用户登录成功
# [WARNING] 2024-01-15 10:30:47: 磁盘空间不足
Python 的 groupdict() 在这种场景下简直是神器------直接拿到字典,不需要逐个取 Groups["xxx"]。
场景二:从 HTML 中提取数据
从网页内容里提取链接和图片地址。
C# 版本:
string html = @"
<a href='https://example.com'>链接1</a>
<a href='https://test.org'>链接2</a>
<img src='https://img.cdn.com/photo.jpg'>
<img src='https://img.cdn.com/logo.png'>
";
// 提取所有链接
var linkPattern = new Regex(@"<a\s+href=['""]([^'""]+)['""]>");
MatchCollection links = linkPattern.Matches(html);
foreach (Match link in links)
{
Console.WriteLine($"链接: {link.Groups[1].Value}");
}
// 提取所有图片
var imgPattern = new Regex(@"<img\s+src=['""]([^'""]+)['""]>");
MatchCollection imgs = imgPattern.Matches(html);
foreach (Match img in imgs)
{
Console.WriteLine($"图片: {img.Groups[1].Value}");
}
Python 版本:
import re
html = """
<a href='https://example.com'>链接1</a>
<a href='https://test.org'>链接2</a>
<img src='https://img.cdn.com/photo.jpg'>
<img src='https://img.cdn.com/logo.png'>
"""
# 提取所有链接
links = re.findall(r"""<a\s+href=['"]([^'"]+)['"]>""", html)
print("链接:", links) # ['https://example.com', 'https://test.org']
# 提取所有图片
imgs = re.findall(r"""<img\s+src=['"]([^'"]+)['"]>""", html)
print("图片:", imgs) # ['https://img.cdn.com/photo.jpg', 'https://img.cdn.com/logo.png']
友情提示:正则不是解析 HTML 的最佳工具------遇到复杂的嵌套标签会翻车。真正搞爬虫还是用 BeautifulSoup 或 lxml。但在简单场景下,正则又快又方便。
设计哲学
C# 的正则哲学 :面向对象 + 类型安全。
Regex类封装了所有操作,编译是默认行为。Match对象有丰富的属性和方法------.Success、.Value、.Groups、.Index、.Length,强类型让 IDE 提示无处不在。Python 的正则哲学 :函数式 + 简洁。
re.search()一行搞定,不需要创建对象。编译是可选的优化手段,缓存帮你兜底。.groupdict()返回字典,直接用 key 取值。共同点:底层都是 PCRE 风格的正则引擎,正则语法(模式本身)完全一致------区别只在 API 调用方式。
一句话总结设计哲学: C# 的正则是"先编译,再使用"------像 C# 里的所有东西一样,规矩先行。 Python 的正则是"先用起来,再优化"------像 Python 里的所有东西一样,够用就行。
迁移指南
从 C# 迁移到 Python,正则表达式的核心语法不变,变的是 API 调用方式。
| C# 写法 | Python 写法 | 说明 |
|---|---|---|
Regex.IsMatch(s, p) |
bool(re.search(p, s)) |
Python 没有直接的 IsMatch |
Regex.Match(s, p) |
re.search(p, s) |
返回 Match 对象 |
Regex.Matches(s, p) |
re.findall(p, s) 或 re.finditer(p, s) |
findall 返回列表,finditer 返回迭代器 |
Regex.Replace(s, p, r) |
re.sub(p, r, s) |
参数顺序不同! |
Regex.Split(s, p) |
re.split(p, s) |
参数顺序不同! |
new Regex(p) |
re.compile(p) |
编译正则 |
RegexOptions.Compiled |
不需要 | Python 默认就缓存了 |
RegexOptions.IgnoreCase |
re.IGNORECASE 或 re.I |
不区分大小写 |
RegexOptions.Multiline |
re.MULTILINE 或 re.M |
多行模式 |
match.Groups["name"] |
match.group("name") |
命名分组 |
match.Groups[1].Value |
match.group(1) |
编号分组 |
$1, $2(替换引用) |
\1, \2(替换引用) |
反斜杠 vs 美元符号! |
\k<name>(命名反向引用) |
(?P=name)(命名反向引用) |
写法完全不同! |
(?<name>...)(命名分组) |
(?P<name>...)(命名分组) |
Python 多了个 P |
RegexOptions.RightToLeft |
无内置支持 | Python 需要第三方库 |
最需要注意的三个差异:
参数顺序:C# 是
(字符串, 模式),Python 是(模式, 字符串)替换引用:C# 用
$1,Python 用\1命名分组:C# 是
(?<name>...),Python 是(?P<name>...)
坑点提醒
坑1:re.match() 只匹配开头
# 这是新手最常犯的错误
match = re.match(r"\d+", "abc123")
print(match) # None ------ 以为能匹配到 123,结果什么都没有
# 解决方案:用 re.search()
match = re.search(r"\d+", "abc123")
print(match.group()) # 123
坑2:参数顺序搞反
# C# 写法(错误):Regex.Replace(input, pattern, replacement)
# Python 正确写法:re.sub(pattern, replacement, input)
# 错误写法
# re.sub("hello123", r"\d+", "NUM") # 这会报错!
# 正确写法
result = re.sub(r"\d+", "NUM", "hello123")
print(result) # helloNUM
坑3:findall() 有捕获组时返回元组列表
# 无捕获组:返回字符串列表
matches = re.findall(r"\d+", "a1b2c3")
print(matches) # ['1', '2', '3']
# 有捕获组:返回元组列表(坑!)
matches = re.findall(r"(\w)(\d)", "a1b2c3")
print(matches) # [('a', '1'), ('b', '2'), ('c', '3')]
# 多个捕获组时容易翻车
matches = re.findall(r"(\w+)(\d+)", "hello123 world456")
print(matches) # [('hello', '123'), ('world', '456')]
# 每个元素是元组,不是字符串!
坑4:忘记用原始字符串
# 没用 r"",\d 被 Python 先解释一遍
# re.search("\d+", "abc123") # 可能不按预期工作
# 务必用原始字符串 r""
match = re.search(r"\d+", "abc123")
print(match.group()) # 123
坑5:贪婪模式 vs 非贪婪模式
// C# 贪婪模式
string html = "<div>hello</div><div>world</div>";
Match match = Regex.Match(html, @"<div>(.+)</div>");
Console.WriteLine(match.Groups[1].Value); // hello</div><div>world(贪婪!)
// C# 非贪婪模式
Match match2 = Regex.Match(html, @"<div>(.+?)</div>");
Console.WriteLine(match2.Groups[1].Value); // hello(非贪婪)
# Python 贪婪模式
html = "<div>hello</div><div>world</div>"
match = re.search(r"<div>(.+)</div>", html)
print(match.group(1)) # hello</div><div>world(贪婪!)
# Python 非贪婪模式
match2 = re.search(r"<div>(.+?)</div>", html)
print(match2.group(1)) # hello(非贪婪)
# .+? 和 .*? 是非贪婪模式
# 默认是贪婪模式(尽量多匹配)
# 加 ? 变成非贪婪模式(尽量少匹配)
记住 :贪婪模式(
.+、.*)尽量多匹配,非贪婪模式(.+?、.*?)尽量少匹配。大部分场景用非贪婪更安全。
坑6:编译缓存的陷阱
# Python 自动缓存正则,但缓存有上限(512个)
# 在循环里动态生成正则可能导致缓存溢出
# 不推荐:动态生成正则但不编译
for i in range(1000):
pattern = rf"error_{i}:.*"
re.search(pattern, log_text) # 每次都生成新正则,缓存可能不够
# 推荐:提前编译或复用正则
patterns = [re.compile(rf"error_{i}:.*") for i in range(1000)]
for pattern in patterns:
pattern.search(log_text) # 已编译,性能更好
一句话总结
Python 的 re 模块像瑞士军刀------轻便、直接、函数调用一行搞定。C# 的 Regex 类像精密仪器------编译优化、类型安全、功能全面。正则语法完全相同,迁移时重点关注参数顺序和命名分组的写法差异。
📦 示例代码:C# 转 Python 全系列配套练习代码(含 48 章示例)
💬 欢迎点赞、收藏、转发,你的支持是我持续创作的动力!