一:string 的概念与使用
1.string到底是什么?
在 C 语言中,字符串本质上是一个以 '\0' 结尾的字符序列,比如:
cpp
char str[] = "hello";
它实际上类似:
cpp
h e l l o \0
C 语言虽然提供了 strlen、strcpy、strcat 等函数,但字符串的数据和操作字符串的函数是分离的,而且底层空间经常需要程序员自己管理,因此容易出现越界、空间不足等问题
C++ 中可以直接:
cpp
#include <string>
using namespace std;
string s = "hello";
你可以把 string 暂时理解为:
C++ 标准库封装好的"可自动管理空间的字符串类"。
它不仅保存字符串本身,还提供了大量成员函数:
cpp
s.size();
s += "world";
s.find("or");
s.substr(...);
s.clear();
所以学习 string 时,实际上要掌握两个层次:
cpp
第一层:会使用 string
↓
构造、访问、遍历、增删、查找、截取、输入输出
第二层:理解 string 为什么能这样工作
↓
动态内存、构造函数、析构函数
↓
浅拷贝 / 深拷贝
↓
拷贝构造 / operator=
2.string 的构造:怎么创建一个字符串对象?
有几个常见构造方式
cpp
string s1;
创建空字符串:
cpp
s1 = ""
可以使用 C 风格字符串构造:
cpp
string s2("hello");
也可以:
cpp
string s2 = "hello";
还可以创建 n 个相同字符:
cpp
string s3(5, 'A');
结果:
cpp
AAAAA
还可以进行拷贝构造:
cpp
string s4(s2);
相当于:
cpp
s2 = "hello"
s4 = "hello"
因此最常见的四种情况就是:
cpp
string s1; // 空字符串
string s2("hello"); // C字符串构造
string s3(5, 'x'); // xxxxx
string s4(s2); // 拷贝构造
二:auto 和范围 for
1. auto
auto 的核心含义是:
让编译器根据初始化表达式自动推导变量类型。
例如:
cpp
int a = 10;
auto b = a;
auto c = 'A';
编译器实际上会推导成:
cpp
int b = a;
char c = 'A';
特别适合类型很长的情况。
这里要记住一个很重要的区别:
cpp
int x = 10;
auto a = x;
auto& b = x;
a 是一个新的变量:
cpp
a ──> 10
x ──> 10
修改 a 不影响 x。
而:
cpp
auto& b = x;
b 是 x 的引用:
修改 b 就是在修改 x
2.范围 for
传统遍历:
cpp
string s = "hello";
for (int i = 0; i < s.size(); i++)
{
cout << s[i];
}
C++11 可以写成:
cpp
for (auto ch : s)
{
cout << ch;
}
范围 for 可以用于数组和容器对象,它会自动完成迭代、取数据和结束判断;对于容器,底层遍历思想可以理解为利用迭代器完成
这里尤其重要的是:
cpp
for (auto ch : s)
和:
cpp
for (auto& ch : s)
是不一样的。
前者:
cpp
for (auto ch : s)
{
ch = 'A';
}
ch 是字符的副本,所以不会真正修改字符串。
而:
cpp
for (auto& ch : s)
{
ch = 'A';
}
ch 是原字符的引用,可以直接修改 string。
比如可以对字符进行乘法操作
cpp
for (auto& e : array)
{
e *= 2;
}
三:string****类的常用接口
(一)string类对象的容量操作
1.size(),length(),capacity()
size()和length()都是返回字符串有效字符长度
capacity()是返回空间总大小
他们的返回值都是size_t类型的
理解它们之前,一定要区分:
cpp
size = 当前有多少个有效字符
capacity = 当前已经准备了多少存储空间
例如:
cpp
string s = "hello";
逻辑上:
cpp
有效字符:
h e l l o
↑ ↑
共 5 个
size() = 5
而 capacity() 可能比 5 大,因为 string 为了以后追加字符,可能提前准备额外空间。
可以类比成:
cpp
宿舍当前住 5 人 → size = 5
宿舍最多能住 10 人 → capacity = 10
例如:

两者底层作用相同,通常更常使用 size(),因为它和其他容器的接口保持一致
2.empty()
empty()是检测字符串释放为空串,是返回true,否则返回****false
它的返回值类型是bool类型
cpp
string s2 = "";
if (s2.empty())
{
cout << "字符是空" << endl;
}
else
cout << "字符不为空" << endl;

相当于判断:
cpp
s.size() == 0
但:
cpp
s.empty()
可读性更好。
3.clear()
clear()只清字符,不一定释放空间
例如:
cpp
string s = "hello";
s.clear();
此时:
cpp
s.size() == 0
但是特别强调:
clear()只是清除有效字符,并不会因此改变底层已经拥有的空间大小
可以理解为:
cpp
原来:
size = 5
capacity = 15
clear以后:
size = 0
capacity = 15

宿舍的人走了,但房间没有拆
4.reserve()
reserve():提前准备空间
例如你准备往字符串里放很多字符:
cpp
string s;
s.reserve(100);
意思不是:字符串现在有100个字符
而是:提前准备能够容纳大约100个字符的空间
所以:
cpp
s.size()
仍然是:
0
注意:
reserve改变的是预留空间,不改变有效元素个数
它的用途主要是提高效率。
假如你不断:
s += 'a';
s += 'b';
s += 'c';
...
空间不够时可能需要:
申请新空间
↓
复制旧数据
↓
释放旧空间
如果提前知道大概需要 1000 个字符:
cpp
s.reserve(1000);
就可以减少重新申请空间的次数。
所以如果能够预估字符串大概会存放多少字符,可以提前使用 reserve。
5.resize()
resize():将有效字符的个数该成n个,多出的空间用字符c****填充
注意:resize() 真正改变 size(),它与reserve()不同
假设:
cpp
string s = "hello";
执行:
cpp
s.resize(3);
得到:
hel
因为有效字符数量变成了 3。
如果:
cpp
s.resize(8, 'x');
得到:
helloxxx
当增加字符个数时:
cpp
resize(n)
会用默认值填补新增位置,而:
cpp
resize(n, c)
会用字符 c 填补;如果缩小字符串,则底层空间总大小通常不会因此缩小
所以你可以牢牢记住:
(二)string****类对象的访问及遍历操作
1.operator:返回pos位置的字符
最常见的方式:
cpp
operator[]
例如:
cpp
string s = "hello";
cout << s[0];
输出:
cpp
h
也可以修改:
cpp
s[0] = 'H';
得到:
cpp
Hello
可以把:
cpp
s[i]
理解成:访问字符串中的第 i 个字符
下标从0开始
所以:
hello
01234
2.begin+end
首先我们要制动什么是迭代器呢?
你可以先把它理解成:
一种"类似指针"的对象,用来访问容器中的元素。
它的主要作用是:按照一定顺序遍历 vector、string、list、map 等容器中的元素,而不需要关心这些容器内部到底是怎么存储数据的
begin():指向第一个字符
例如:
cpp
string s = "hello";
auto it = s.begin();
cout << *it;
输出:h
因为:
cpp
hello
↑
begin()
也可以修改字符:
cpp
string s = "hello";
auto it = s.begin();
*it = 'H';
cout << s;
输出:Hello
因为:
cpp
*it
代表当前迭代器指向的字符
end():指向最后一个字符后面
这是最容易搞错的地方。
对于:
cpp
string s = "hello";
不是:
cpp
h e l l o
↑
end()
而是:
cpp
h e l l o [ ]
↑
end()
所以不能直接:
cpp
cout << *s.end();
这是错误的,因为 end() 不指向有效字符。
如果想通过 end() 得到最后一个字符,可以先往前移动一次:
cpp
string s = "hello";
auto it = s.end();
--it;
cout << *it;
输出:o
因为:
cpp
初始:
h e l l o [ ]
↑
it
执行 --it 后:
h e l l o [ ]
↑
it
3.rbegin+rend
rbegin():反向遍历的起点
rbegin() 可以理解成:
reverse begin,也就是反向遍历时的第一个位置。
对于:
cpp
string s = "hello";
rbegin() 指向最后一个字符:
cpp
h e l l o
↑
rbegin()
例如:
cpp
string s = "hello";
auto it = s.rbegin();
cout << *it;
输出:o
rend():反向遍历的结束位置
rend() 是:
reverse end,反向遍历结束的位置。
它位于第一个字符之前:
cpp
[ ] h e l l o
↑
rend()
所以反向遍历:
cpp
#include <iostream>
#include <string>
using namespace std;
int main()
{
string s = "hello";
for (auto it = s.rbegin(); it != s.rend(); ++it)
{
cout << *it << " ";
}
return 0;
}
输出:
cpp
o l l e h
注意这里一个比较有意思的地方:
cpp
++it;
虽然写的是 ++,但是由于这是反向迭代器,所以实际方向是从右往左:
cpp
o → l → l → e → h
也就是说:普通迭代器 ++是向右走。
而:反向迭代器 ++是向左走。
4.三种非常重要的遍历方式:
第一种是下标:
cpp
for (size_t i = 0; i < s.size(); ++i)
{
cout << s[i];
}
第二种是迭代器:
cpp
for (auto it = s.begin(); it != s.end(); ++it)
{
cout << *it;
}
cpp
begin()
end()
用于获得遍历区间,另外还有:
cpp
rbegin()
rend()
进行反向遍历
第三种就是最方便的范围 for:
cpp
for (auto ch : s)
{
cout << ch;
}
需要修改时:
cpp
for (auto& ch : s)
{
ch += 1;
}
(三) string****类对象的修改操作
1.push_back、append、+=
push_back:在字符串后尾插字符****c
append:在字符串后追加一个字符串
+=:在字符串后追加字符串****str
例如:
cpp
string s = "hello";
s.push_back('!');
得到:
cpp
hello!
push_back 主要添加一个字符。
而:
cpp
s.append(" world");
得到:
cpp
hello world
最常见的还是:
cpp
s += '!';
s += " world";
+= 既可以连接字符,也可以连接字符串,因此实际使用中很方便
2.c_str():
c_str() 是 std::string 提供的一个成员函数,用来把 C++ 的 string 转成 C 风格字符串 ,也就是 const char*
比如:
cpp
#include <iostream>
#include <string>
using namespace std;
int main()
{
string s = "hello";
const char* p = s.c_str();
cout << p << endl;
return 0;
}
输出:
cpp
hello
你可以把它理解为:
cpp
string s = "hello";
内部保存的是字符串内容,而:
cpp
s.c_str()
会返回一个指向字符序列的指针:
cpp
h e l l o \0
↑
p
最后的 \0 是 C 风格字符串的结束标志
所以 c_str() 最常见的用途,就是在一些只接受 const char* 的 C 函数或旧式接口里使用 std::string。
例如 printf:
cpp
#include <cstdio>
#include <string>
using namespace std;
int main()
{
string s = "hello";
printf("%s\n", s.c_str());
return 0;
}
因为 %s 需要的是:
cpp
const char*
而不是:
cpp
std::string
所以不能直接这样写:
cpp
printf("%s", s); // 错误
而要写:
cpp
printf("%s", s.c_str()); // 正确
3.find和npos
find() 用来查找字符或子串的位置;如果没找到,就返回 string::npos
假设有:
cpp
string s = "hello world";
查找字符:
cpp
size_t pos = s.find('o');
cout << pos;
输出:4
因为字符串下标是从 0 开始的:
h e l l o w o r l d
0 1 2 3 4 5 6 7 8 9 10
↑
o
所以第一个 'o' 的位置是 4。
如果查找一个不存在的字符:
cpp
size_t pos = s.find('x');
这时候不会返回 -1 来表示失败,而是返回:
cpp
string::npos
因此通常这样写:
cpp
string s = "hello world";
size_t pos = s.find('x');
if (pos == string::npos)
{
cout << "没有找到";
}
else
{
cout << "找到了,位置是:" << pos;
}
输出:没有找到
find() 也可以查找子字符串:
cpp
string s = "hello world";
size_t pos = s.find("world");
cout << pos;
输出:6
因为:
cpp
hello world
↑
6
也就是 "world" 从下标 6 开始。
完整写法通常是:
cpp
string s = "hello world";
size_t pos = s.find("world");
if (pos != string::npos)
{
cout << "找到了" << endl;
cout << "起始位置:" << pos << endl;
}
else
{
cout << "没有找到" << endl;
}
find() 默认找第一个匹配位置
例如:
cpp
string s = "banana";
size_t pos = s.find('a');
cout << pos;
输出:1
虽然 banana 中有很多个 a:
cpp
b a n a n a
0 1 2 3 4 5
↑ ↑ ↑
但是:
cpp
s.find('a')
只返回第一个:1
find可以指定从哪里开始找
例如:
cpp
string s = "banana";
size_t pos = s.find('a', 2);
cout << pos;
这里表示:
从下标
2开始寻找'a'。
字符串:
cpp
b a n a n a
0 1 2 3 4 5
↑
从这里开始
所以找到的是:3,而不是 1。
如何找到所有相同字符?
例如:
cpp
string s = "banana";
我们想找到所有 'a':
cpp
size_t pos = s.find('a');
while (pos != string::npos)
{
cout << pos << " ";
pos = s.find('a', pos + 1);
}
输出:
cpp
1 3 5
4.rfind
rfind() 是 std::string 中用来从后往前查找字符或子字符串 的函数
例如:
cpp
string s = "banana";
字符串下标是:
cpp
b a n a n a
0 1 2 3 4 5
如果写:
cpp
cout << s.find('a');
输出:1
因为 find() 找的是第一个 'a'。
而:
cpp
cout << s.rfind('a');
输出:5
因为 rfind() 从后往前找,所以找到的是最后一个 'a'
虽然 rfind() 是从右往左查找,但是它返回的仍然是字符串正常的下标
查找子字符串
rfind() 不仅能查一个字符,也可以查字符串。
例如:
cpp
string s = "abcabcabc";
size_t pos = s.rfind("abc");
cout << pos;
输出:6
因为:
cpp
a b c a b c a b c
0 1 2 3 4 5 6 7 8
↑
abc
"abc" 一共出现了三次:
cpp
abc abc abc
↑ ↑ ↑
0 3 6
rfind() 找最后一次出现的位置:
cpp
6
而:
cpp
s.find("abc")
返回:
cpp
0
所以区别非常直观:
cpp
s.find("abc"); // 第一次出现的位置
s.rfind("abc"); // 最后一次出现的位置
找不到时仍然返回 string::npos
这和 find() 完全一样。
例如:
cpp
string s = "hello";
size_t pos = s.rfind('x');
if (pos == string::npos)
{
cout << "没有找到";
}
else
{
cout << "找到了:" << pos;
}
输出:没有找到
所以标准写法依然是:
cpp
if (s.rfind("abc") != string::npos)
{
cout << "找到了";
}
可以指定从哪个位置往前找
它还有一种形式:
s.rfind(要找的内容, 起始位置);
不过这里要特别注意:
rfind()的第二个参数表示:从这个下标开始,向前查找。
例如:
cpp
string s = "banana";
size_t pos = s.rfind('a', 4);
cout << pos;
字符串:
b a n a n a
0 1 2 3 4 5
↑
从这里开始
从下标 4 开始向左:
4 → 3 → 2 → 1 → 0
首先遇到 'a' 的位置是:3
所以输出:3
5.substr
语法可以理解为:
cpp
s.substr(pos, n);
即:
从
pos位置开始,取n个字符。
定义为:从 pos 开始截取 n 个字符并返回
例如:
cpp
string s = "hello world";
string t = s.substr(6, 5);
得到:
cpp
world
所以:
cpp
hello world
↑
pos = 6
取 5 个:
cpp
world
find + substr 经常一起出现:
cpp
size_t pos = s.find(':');
string left = s.substr(0, pos);
string right = s.substr(pos + 1);