前面介绍过 set 和 map。这篇来看它们的无序版本:unordered_set 和 unordered_map。两者都用哈希表管理 Key;区别是前者只存 Key,后者存 Key 和 Value。遍历时不按 Key 排序,适不适合用它们,通常先看程序是否需要有序访问。
一、unordered_set
Key、哈希和相等判断
unordered_set 的模板参数可以简化为:
cpp
template <
class Key,
class Hash = std::hash<Key>,
class KeyEqual = std::equal_to<Key>,
class Allocator = std::allocator<Key>
>
class unordered_set;
Hash 根据 Key 算出哈希值,KeyEqual 判断两个 Key 是否相等。Key 不需要能转换成整数;如果默认的 std::hash 不支持某个自定义类型,就给它提供哈希函数和相等判断:
cpp
struct UserId {
int value;
};
struct UserIdHash {
std::size_t operator()(const UserId& id) const noexcept {
return std::hash<int>{}(id.value);
}
};
struct UserIdEqual {
bool operator()(const UserId& lhs, const UserId& rhs) const noexcept {
return lhs.value == rhs.value;
}
};
std::unordered_set<UserId, UserIdHash, UserIdEqual> user_ids;
user_ids.insert(UserId{42});
auto it = user_ids.find(UserId{42});
这里有个必须满足的约定:如果相等判断认为两个 Key 相等,它们的哈希值也必须相同。哈希值相同并不代表 Key 一定相等,容器还会用 KeyEqual 区分它们。
和 set 怎么选
set 会按 Key 排序,unordered_set 不会。set 通常基于平衡树,增删查为 O(logN);无序容器的这些操作平均为 O(1),但最坏情况下可能退化到 O(N)。因此,平均复杂度不代表每次操作都更快。
下面是插入、查找和删除的写法:
cpp
std::unordered_set<int> values{5, 1, 3};
auto [it, inserted] = values.insert(3);
std::cout << inserted << '\n'; // 3 已存在,插入失败
if (values.find(5) != values.end()) {
std::cout << "found 5\n";
}
values.erase(1);
insert 返回迭代器和插入结果;find 找不到时返回 end()。如果需要按顺序遍历,选 set;如果不在意顺序,可以考虑 unordered_set。无序容器的迭代器是前向迭代器,不能像 set 的双向迭代器那样向后移动。
二、unordered_map
unordered_map 保存 Key 和 Value 的对应关系。常用的 insert、erase、find 与 map 类似,遍历顺序则不保证按 Key 排列:
cpp
std::unordered_map<std::string, int> scores{
{"pear", 4}, {"apple", 3}, {"banana", 2}
};
scores["orange"] = 5;
++scores["apple"];
auto it = scores.find("pear");
if (it != scores.end()) {
std::cout << it->first << ": " << it->second << '\n';
}
operator[] 可以新增或修改 Value。要留意的是,访问一个不存在的 Key 时,它会先插入该 Key,并将 Value 初始化为默认值;如果只是想检查 Key 在不在,用 find 更合适。
本次程序运行里,map 遍历出的 Key 是 apple banana pear;unordered_map 则观察到 orange banana apple pear。后者只是这次运行看到的顺序,不能当作排序结果或程序保证。
三、多重无序容器
如果需要允许重复 Key,可以用 unordered_multiset 或 unordered_multimap:
cpp
std::unordered_multiset<int> values{2, 2, 5};
std::cout << values.count(2) << '\n'; // 2
std::unordered_multimap<std::string, int> scores;
scores.emplace("Ada", 90);
scores.emplace("Ada", 95);
std::cout << scores.count("Ada") << '\n'; // 2
四、哈希桶与负载因子
常见的桶接口有 bucket_count()、bucket_size() 和 bucket(key);负载因子相关接口包括 load_factor()、max_load_factor()、reserve() 和 rehash()。多数日常代码不需要直接操作桶。如果大致知道元素数量,提前调用 reserve() 可以减少插入过程中反复扩容的机会。
小结
需要按 Key 排序或顺序遍历时用 set、map;不需要排序时,可以考虑 unordered_set、unordered_map。自定义 Key 要提供一致的哈希和相等判断。无序容器的增删查平均效率为 O(1),最坏可到 O(N),最终是否更合适还需结合具体数据分布和业务需求判断。