【C++】命名空间、缺省参数、函数重载与引用

【C++】命名空间、缺省参数、函数重载与引用


🔥 H.莓飛: 个人主页

❄️个人专栏:「C++」「数据结构」「Linux系统编程」

✨ 心不畏死,踏屍前行。

📦 代码仓库:cocatrice_csdn_code


📖 博主简介:


概览

  • 第一个 C++ 程序
  • 命名空间、定义、嵌套与多文件合并
  • 三种使用方式与 std 的写法代价
  • cin、cout 与缓冲区
  • 缺省参数与函数重载
  • 引用,inline 与 nullptr

核心知识

一、第一个 C++ 程序

C 的入门程序依赖一个头文件和一个函数,换成 C++ 之后,头文件换成了流库,输出换成了往一个对象里写数据,编译命令从 gcc 换成 g++,两边各写一份,先看代码。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp03/hello.cpp
#include <iostream>
using namespace std;

int main()
{
    cout << "hello C++" << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp03
g++ -std=c++11 hello.cpp -o hello
./hello
hello C++

与 C 版相比,变化集中在三处,第一处是头文件,stdio.h 换成了 iostream,后者用来看起来像类名的那些流类型;printf 在 C++ 里仍然可以用,只是包含它的头文件要自己写清楚。第二处是输出形式,printf 是一个函数调用,参数个数与类型由调用者负责对齐,cout 则是 std 域里的一个 ostream 对象,往它里面写什么由运算符重载决定,类型不需要在调用处说明。第三处是后缀,C++ 源文件用 .cpp 结尾,编译器按后缀挑选语言。

编译这一步看起来只有一条命令,实际上背后串了四步。用 -E、-S、-c 三个开关把中间产物拦下来,就能看到每一段的位置与体量。

图 1 一个源文件走到可执行文件

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp03
g++ --version | head -1
g++ (GCC) 4.8.5 20150623 (Red Hat 4.8.5-44)
g++ -std=c++11 -E hello.cpp -o hello.i && wc -l hello.i
23705 hello.i
g++ -std=c++11 -S hello.cpp -o hello.s && wc -l hello.s
87 hello.s
g++ -std=c++11 -c hello.cpp -o hello.o && nm hello.o | grep -E " main|_Z" | head -4
0000000000000000 T main
0000000000000027 t _Z41__static_initialization_and_destruction_0ii
                 U _ZNSolsEPFRSoS_E
                 U _ZNSt8ios_base4InitC1Ev
file hello
hello: ELF 64-bit LSB executable, x86-64, version 1 (SYSV), dynamically linked (uses shared libs), for GNU/Linux 2.6.32, BuildID[sha1]=556bbf2f55428d594804a09a0cf2df9e5b563208, not stripped

一个 #include 展开之后有两万三千多行,编译完只剩 87 行汇编。hello.o 的符号表里,main 是大写 T,表示这是一个对其他目标文件可见的定义;_Z41__static_initialization_and_destruction_0ii 是小写 t,表示它只在本目标文件内可见,名字里的 41 与 0ii 是编译器给参数表编码的结果,前面已经出现了名字修饰的影子,这一段留到第十节展开。U _ZNSolsEPFRSoS_E 与 U _ZNSt8ios_base4InitC1Ev 是大写 U,表示这两个符号在本文件里被引用但还没有定义,链接阶段才会从标准库里补上。

把同一条命令换成 gcc,编译器照样按 .cpp 后缀把文件当成 C++ 来编译,语法检查全部通过,报错发生在链接阶段。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp03
gcc hello.cpp -o hello_gcc
/tmp/cc5GMb1X.o: In function `main':
hello.cpp:(.text+0xa): undefined reference to `std::cout'
hello.cpp:(.text+0xf): undefined reference to `std::basic_ostream<char, std::char_traits<char> >& std::operator<< <std::char_traits<char> >(std::basic_ostream<char, std::char_traits<char> >&, char const*)'
hello.cpp:(.text+0x14): undefined reference to `std::basic_ostream<char, std::char_traits<char> >& std::endl<char, std::char_traits<char> >(std::basic_ostream<char, std::char_traits<char> >&)'

三条报错都在说同一个意思:std::cout 这个符号在目标文件里被引用,链接时找不到定义。gcc 与 g++ 在编译阶段用的是同一个 C++ 前端,差别出现在它们默认带上哪些库。链接 C++ 程序需要 libstdc++,gcc 不会自动加上,所以要么换成 g++,要么在 gcc 后面补上 -lstdc++。

二、命名空间要解决的问题

C 语言里所有全局名字共享一个空间,两个库各自定义了一个同名函数,把它们的头文件同时包含进来就会出现重定义。标准库自己就占用了大量短名字,写一份自己的代码,撞名的机会比想象中多。下面这段代码给一个全局变量起名叫 rand。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/rand_conflict.cpp
#include <stdio.h>
#include <stdlib.h>

int rand = 10;

int main()
{
    printf("%d\n", rand);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ rand_conflict.cpp -o rand_conflict
rand_conflict.cpp:3:5: error: 'int rand' redeclared as different kind of symbol
 int rand = 10;
     ^
In file included from rand_conflict.cpp:2:0:
/usr/include/stdlib.h:374:12: error: previous declaration of 'int rand()'
 extern int rand (void) __THROW;
            ^

stdlib.h 里已经有一个 int rand(void),再定义一个同名的整型变量,编译器认为这是同一个名字的两种声明,直接拒绝。C 语言下的处理办法有三种:给名字加上模块前缀,写成 bit_rand,代价是每个名字都要手工维护前缀,库一升级就可能撞上新的名字;把变量声明成 static,让它只在当前文件里可见,代价是别的文件想用就用不到了;干脆换个名字,代价是这个名字在别人的代码里仍然可能撞车。

C++ 的解法是把名字放进一个域里,域是一层可以叠加的命名层,两个域里的 rand 互不干扰,调用处写上域的名字就能指定用哪个。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/ns_rand.cpp
#include <stdio.h>
#include <stdlib.h>

namespace bit
{
    int rand = 10;
}

int main()
{
    printf("库函数 rand 的地址 %p,第一次返回 %d\n", (void*)&rand, rand());
    printf("域里的 rand = %d\n", bit::rand);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ ns_rand.cpp -o ns_rand && ./ns_rand
库函数 rand 的地址 0x4004b0,第一次返回 1804289383
域里的 rand = 10

加了这一层域之后,同一个名字可以在全局作用域、在 bit 域、在函数内部同时存在,访问时按就近原则查找。查找顺序是编译器在编译期定下来的,编译产物里每一个名字都已经绑定到了唯一的位置。

图 2 编译器怎么找到一个名字

查找过程走完之后,如果哪一个候选都没找到,编译器的提示里往往会带上它认为最接近的那个名字。下面这段代码把变量放在域里,访问时忘了写域的名字。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/ns_use_err.cpp
#include <stdio.h>

namespace bit
{
    int a = 0;
    int b = 1;
}

int main()
{
    printf("%d\n", a);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ ns_use_err.cpp -o ns_use_err
ns_use_err.cpp: In function 'int main()':
ns_use_err.cpp:9:20: error: 'a' was not declared in this scope
     printf("%d\n", a);
                    ^
ns_use_err.cpp:9:20: note: suggested alternative:
ns_use_err.cpp:4:9: note:   'bit::a'
     int a = 0;
         ^

note 这一行给出的候选 bit::a 就是编译器查找失败之后额外做的一次尝试。C++ 的报错经常带上这类提示,遇到作用域相关的错误,先读 note 里给出的候选名字,多数时候能直接看出少写了哪一层域。

三、命名空间的定义与合并

命名空间用 namespace 关键字加一对花括号定义,花括号后面不写分号。域里可以放变量、函数、结构体、类,也可以再放一个命名空间。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp04/ns_nest.cpp
#include <stdio.h>

int a = 10;

namespace bit
{
    int a = 0;
    namespace pg
    {
        int a = 1;
        int b = 2;
    }
}

namespace bit
{
    namespace pg
    {
        int c = 3;
    }
}

int main()
{
    printf("全局 a = %d\n", ::a);
    printf("bit::a = %d\n", bit::a);
    printf("bit::pg::a = %d\n", bit::pg::a);
    printf("bit::pg::b = %d\n", bit::pg::b);
    printf("bit::pg::c = %d\n", bit::pg::c);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp04
g++ ns_nest.cpp -o ns_nest && ./ns_nest
全局 a = 10
bit::a = 0
bit::pg::a = 1
bit::pg::b = 2
bit::pg::c = 3

这段代码里有两个可以留意的细节,第一个是同一个 bit 域被打开了两次,一次放 a,一次放 pg,编译器把两次定义合成一个域,所以 bit::pg::c 能访问到。第二个是 :: 单独写在名字前面表示全局作用域,::a 拿到的就是文件开头那个 a。嵌套域的名字用一层一层的 :: 连起来读,写法上从左到右依次收窄。

C++11 的标准里没有地方定义域用的简写语法,C++17 才允许把 namespace bit::pg { } 写成一行,这一篇的实验都在 g++ 4.8.5 上做,只能用嵌套写法。

命名空间在这一点上的用处体现在两个文件之间。同一个域可以在不同的源文件里分别定义,最后合到一起,工程里常用它把同一个模块的代码拆到多个文件,下面把 Add 与 Sub 分到两个文件里编译。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp02/ns_merge.h
#pragma once
namespace bit
{
    int Add(int left, int right);
    int Sub(int left, int right);
}
cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp02/ns_merge_a.cpp
#include "ns_merge.h"
namespace bit
{
    int Add(int left, int right) { return left + right; }
}
cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp02/ns_merge_b.cpp
#include <stdio.h>
#include "ns_merge.h"
namespace bit
{
    int Sub(int left, int right) { return left - right; }
}

int main()
{
    printf("%d\n", bit::Add(3, 4));
    printf("%d\n", bit::Sub(3, 4));
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
g++ ns_merge_a.cpp ns_merge_b.cpp -o ns_merge && ./ns_merge
7
-1
g++ -c ns_merge_a.cpp -o ns_merge_a.o && nm ns_merge_a.o | grep "_ZN3bit"
0000000000000000 T _ZN3bit3AddEii
g++ -c ns_merge_b.cpp -o ns_merge_b.o && nm ns_merge_b.o | grep "_ZN3bit"
0000000000000000 T _ZN3bit3SubEii
                 U _ZN3bit3AddEii

两个目标文件里各有一个属于 bit 域的函数定义,ns_merge_b.o 里的 U _ZN3bit3AddEii 表示它在等一个定义,链接器把 a.o 里那个同名符号接上,程序才能跑出 7 与 -1。接不上时编译器与链接器给出的提示不一样:头文件里漏写声明,编译期就报名字不属于这个域;声明写了但实现文件没编进来,链接期才报未定义引用。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ ns_part_b.cpp -o ns_part_b
ns_part_b.cpp: In function 'int main()':
ns_part_b.cpp:8:20: error: 'Add' is not a member of 'bit'
     printf("%d\n", bit::Add(3, 4));
                    ^

域里的变量也一样带着域的名字进符号表,用 nm 看一眼区分静态与外部链接。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp03
nm -C ns_sym.o | grep -E "count|Add"
0000000000000000 T bit::Add(int, int)
0000000000000000 D bit::g_count
0000000000000004 d bit::s_count
nm ns_sym.o | grep "_ZN3bit"
0000000000000000 T _ZN3bit3AddEii
0000000000000000 D _ZN3bit7g_countE
0000000000000004 d _ZN3bitL7s_countE

bit::g_count 是大写 D,表示它在数据段里并且对其他目标文件可见;bit::s_count 是小写 d,编译后的符号里多出一个小写的 L,表示它只有内部链接,对应源文件里写的 static。域这一层包装改变了名字,没有改变链接属性。

图 3 两个源文件里的同名命名空间怎么合到一起

四、使用方式

指定访问、using 声明、using namespace 是把域里的名字拿出来用的三种写法,代价从低到高,方便程度也从低到高。

指定访问要写全 bit::a 这样的限定名,读代码时一眼能看出这个名字来自哪个域,同名冲突时也只需要把限定名写清楚。using 声明把一个名字引进当前作用域,using bit::a; 之后直接写 a 就等价于写 bit::a,它的影响范围到当前作用域为止。using namespace 把整个域展开,域里所有名字都变成可以直接写的名字,写起来最短,撞名的概率也最大。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp04/ns_nest.cpp
int a = 7;
printf("局部 a = %d\n", a);
printf("加 :: 的 a = %d\n", ::a);
printf("加 bit:: 的 a = %d\n", bit::a);
{
    using bit::a;
    printf("using 声明之后 a = %d\n", a);
}
printf("出了作用域 a = %d\n", a);
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp04
局部 a = 7
加 :: 的 a = 10
加 bit:: 的 a = 0
using 声明之后 a = 0
出了作用域 a = 7

这一段回显把三个作用域摆在一起看得很清楚:局部作用域里 a 是 7,全局作用域里 a 是 10,bit 域里 a 是 0,::a 这个写法绕开了就近查找,直接落到最外层的全局作用域。using 声明写在花括号里,只在这一对花括号内生效,出了花括号之后 a 又回到 7。把 using 声明和同名变量放在同一层作用域,编译器会直接拒绝。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
ns_using_clash.cpp:7:5: error: redeclaration of 'int a'
 int a = 5;
     ^
ns_using_clash.cpp:4:9: note: previous declaration 'int bit::a'
     int a = 0;

using namespace 的代价在多个域同时展开时出现。下面这段代码把 bit 与 pg 两个域都展开,两个域里各有一个 Add。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
g++ ns_amb.cpp -o ns_amb
ns_amb.cpp: In function 'int main()':
ns_amb.cpp:14:28: error: call of overloaded 'Add(int, int)' is ambiguous
     printf("%d\n", Add(1, 2));
ns_amb.cpp:14:28: note: candidates are:
ns_amb.cpp:8:9: note: int pg::Add(int, int)
ns_amb.cpp:4:9: note: int bit::Add(int, int)

两个候选的参数表完全一样,编译器无法在它们之间排出优劣,只能报二义。this 一类的错误提示里会列出全部候选,把限定名写全就能立刻定下来用哪一个。

图 4 三种使用方式把名字放到哪一层

using namespace 写在函数内部也是允许的,它的影响范围同样到最近的这对花括号为止,函数之外的地方看不到这个展开。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp04
g++ ns_func.cpp -o ns_func && ./ns_func
3

五、std

标准库的每一个名字都在 std 域里,cout、cin、string、vector 都属于它,所以每一篇 C++ 代码开头都会看到 std 这个名字。它本身没有任何特殊之处,用的就是这一节讲的这套机制。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp03/hello.cpp
#include <iostream>
using namespace std;

int main()
{
    cout << "hello C++" << endl;
    return 0;
}

这份写法把一个域里上百个名字全部展开到当前文件。写小段代码时看不出问题,规模上来之后就会撞上标准库的名字。下面这份代码同时包含流库和一个自定义的 string 结构体。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
g++ std_clash.cpp -o std_clash
std_clash.cpp: In function 'int main()':
std_clash.cpp:10:5: error: reference to 'string' is ambiguous
     string s;
std_clash.cpp:4:8: note: candidates are: struct string
In file included from /usr/include/c++/4.8.2/iosfwd:39:0,
                 from /usr/include/c++/4.8.2/ios:38,
                 from /usr/include/c++/4.8.2/ostream:38,

报错里那个候选来自 iosfwd,也就是 iostream 间接包含进来的头文件,标准库的 string 就在那里。把 using namespace std 放进头文件,代价会顺着 #include 传给每一个包含它的源文件,所以工程里的通行做法是头文件里一律不展开,要写就用 std:: 前缀,或者只在 .cpp 文件里写单名引入。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp02/ok_style.cpp
#include <iostream>
using std::cout;
using std::endl;

int main()
{
    cout << "只引入用到的两个名字" << endl;
    return 0;
}

有一种情况看起来像是撞名却没有报错,值得单独看一眼。自己写一个 swap,同时展开 std,调用处编译器会优先选那个普通的函数,而不是标准库里的模板。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
g++ std_clash2.cpp -o std_clash2 && ./std_clash2
2 1

非模板函数在参数完全匹配时优先于模板实例化出来的版本,这是重载决议里的一条规则,第九节会接着讲。程序编过并且跑出 2 1,说明调用处落在自定义的那个 swap 上。

域只管名字的可见性,不管对象的生死,域里的变量和全局变量一样具有静态存储期,程序启动时构造,main 结束之后析构,多个文件各自定义的同名域里也不会出现两份实体。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp05
g++ ns_life.cpp -o ns_life && ./ns_life
A 构造
main 开始
A 构造
A 析构
main 结束
A 析构

六、输入输出

C 语言的输入输出是一对函数,printf 用格式串决定怎么把后面的参数解释成字符,scanf 用格式串决定怎么把读到的字符装进变量。C++ 把这件事换成两个对象,cout 负责输出,cin 负责输入,<< 与 >> 是重载出来的运算符,具体打什么类型由对象自己判断。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/io_demo.cpp
#include <iostream>
#include <cstdio>
using namespace std;

int main()
{
    int a = 0;
    double b = 0.1;
    char c = 'x';

    cout << a << " " << b << " " << c << endl;

    scanf("%d%lf", &a, &b);
    printf("%d %lf\n", a, b);

    cin >> a >> b >> c;
    cout << a << " " << b << " " << c << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ io_demo.cpp -o io_demo && printf '7 3.5\n9 2.5 z\n' | ./io_demo
0 0.1 x
7 3.500000
9 2.5 z

第一行的 0、0.1、x 是 cout 按变量本身的类型打印出来的,double 默认只保留六位有效数字,所以 0.1 原样输出。接着 scanf 按 %d 与 %lf 读走 7 和 3.5,printf 用 %lf 打回来是 3.500000,六位小数一位不少。第三行由 cin 读入 9、2.5、z 之后再由 cout 打出,整型、浮点、字符三种类型在写法上没有任何区别。

cout 与 printf 最实质的差别在类型是否参与判断。printf 的格式串与实参之间没有类型检查,写错时编译器只给一条警告,程序照样编得出来,运行时按格式串去解释那块内存。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/printf_trap.cpp
#include <cstdio>
#include <iostream>
using namespace std;

int main()
{
    double d = 3.75;
    printf("printf 用 %%d 打 double: %d\n", d);
    cout << "cout 打 double: " << d << endl;

    long long n = 1099511627776LL;
    printf("printf 用 %%d 打 long long: %d\n", n);
    cout << "cout 打 long long: " << n << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ printf_trap.cpp -o printf_trap && ./printf_trap
printf 用 %d 打 double: -1773466536
cout 打 double: 3.75
printf 用 %d 打 long long: 0
cout 打 long long: 1099511627776

同一个 double 3.75,printf 用 %d 打出来是 -1773466536,cout 打出来是 3.75。原因是 %d 只从参数位置上取四个字节按 int 解释,3.75 的八个字节被切成了两半,低四个字节恰好是一个负整数。长整型那条更直接,1099511627776 的低四位是 0,所以打出 0。cout 这边没有格式串,编译器在编译期就把 operator<< 的重载选成了处理 double 与 long long 的那两个版本,类型对不上时根本编不过。

cin 出错时会把状态位留下,下面这段代码往 int 里读一个 abc。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp02/cin_fail.cpp
#include <iostream>
#include <string>
using namespace std;

int main()
{
    int i = -1;
    cin >> i;
    cout << "读到的 i = " << i << endl;
    cout << "cin 状态 good=" << cin.good() << " fail=" << cin.fail() << endl;

    int j = -1;
    cin >> j;
    cout << "第二次读到的 j = " << j << endl;

    cin.clear();
    string s;
    cin >> s;
    cout << "清掉状态后读到的字符串 = " << s << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
g++ cin_fail.cpp -o cin_fail && printf 'abc\n12 34\n' | ./cin_fail
读到的 i = 0
cin 状态 good=0 fail=1
第二次读到的 j = -1
清掉状态后读到的字符串 = abc

读失败之后 i 变成 0,good 变成 0,fail 变成 1。流一旦进入失败状态,后面的读取操作会直接返回,j 保持初始值 -1 没有被改动,直到调用 clear() 把状态位清掉,才轮到字符串把 abc 读走。工程里读用户输入时,这个状态位是判断输入是否合法的依据。

下面这一组的差别在于刷新时机,cout 默认与 stdio 共用一套缓冲区,printf 的消息会先落进这块缓冲区,等 cout 遇到 endl 触发刷新时一起写出去。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/buf_mix.cpp
#include <iostream>
#include <cstdio>
using namespace std;

int main()
{
    printf("printf 先写");
    cout << "cout 后写" << endl;

    ios_base::sync_with_stdio(false);
    printf("关同步后 printf 先写");
    cout << "关同步后 cout 后写" << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ buf_mix.cpp -o buf_mix && ./buf_mix
printf 先写cout 后写
关同步后 cout 后写
关同步后 printf 先写

第一行的两段文字合成了一行,printf 的内容先出,cout 的内容在 endl 处跟上来。关掉同步之后顺序反了过来,cout 的内容先落地,printf 的内容留在自己的缓冲区里,等到程序退出才刷出去。两套缓冲机制各管各的,混用时输出的先后就不再由代码顺序决定。sync_with_stdio(false) 换来的东西在 write 调用次数上。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
--- 十万行,换行符 ---
100.00    0.000160           1       144           write
耗时 0.007 秒
--- 十万行,endl ---
100.00    0.724636           7    100000           write
耗时 0.029 秒
--- 十万行,关同步,换行符 ---
  0.00    0.000000           0         1           write
耗时 0.006 秒
--- 十万行,关同步,endl ---
100.00    0.716553           7    100000           write
耗时 0.030 秒
--- 一万行,换行符 ---
100.00    0.000039           3        12           write
耗时 0.007 秒

数字摆在这里就很清楚了,默认同步下十万行用 \n 收尾,整段输出只触发 144 次 write,耗时 0.007 秒;同一批数据换成 endl,write 次数变成 100000 次,耗时 0.029 秒,是前者的四倍多。关掉同步之后同样的十万行只剩 1 次 write。endl 做的事比 \n 多一步,它把一个换行符放进缓冲区之后还要立刻刷新,每一行都对应一次系统调用。缓冲区满一次写一次的机制决定了 \n 的写次数与行数无关,一万行那次是 12 次 write,十万行是 144 次 write。

用一千行做同样的对照,数字的方向一致。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
--- 一千行,换行符 ---
  0.00    0.000000           0         1           write
--- 一千行,endl ---
100.00    0.006933           6      1000           write
--- 一千行,关同步,endl ---
100.00    0.007008           7      1000           write

一千行 \n 的输出只有 1 次 write,endl 的版本是 1000 次,关掉同步也救不了 endl,因为它要的是一次立即刷新。把四种组合并排放在一起,次数与耗时如下。

写法 一千行 十万行
用 \n 收尾 1 次 write 144 次 write,0.007 秒
用 endl 收尾 1000 次 write 100000 次 write,0.029 秒
关同步后用 \n 收尾 1 次 write 1 次 write,0.006 秒
关同步后用 endl 收尾 1000 次 write 100000 次 write,0.030 秒

三个流对象里还有一个默认不缓冲的,cerr 每写一次就刷新一次,下面这份代码往 cerr 里写下标从 0 到 999 的一千行错误信息。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp03
strace -c -e trace=write ./cerr_buf
    100.00    0.024461           8      3000           write

一行错误信息由字符串、数字、换行三部分组成,一共三次 <<,每次 << 都对应一次 write,所以一千行是 3000 次。cerr 与 cout 的区别只在这里,它保证写出去的内容不会因为程序崩溃而留在缓冲区里。clog 和 cout 一样带缓冲,同样是往标准错误上写,用在不需要立即刷新的大批量日志上。

顺便说一句与刷新无关的另一处不同,cout 的默认精度是六位有效数字,不是六位小数。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp03
g++ precision.cpp -o precision && ./precision
3.14159
3.141593
3.141592654
3.1415926535

同一个 double 3.1415926535,cout 打出来是 3.14159,printf 用 %f 打出来是 3.141593,一个按有效数字截,一个按小数位截。加上 setprecision(10) 之后 cout 打 3.141592654,仍是十位有效数字;printf 用 %.10f 打 3.1415926535,是十位小数。需要固定的输出格式时用 printf 的格式串更省心,需要跟着类型走时用 cout。

图 5 流库里的类与两个全局对象

图 6 一次 cout 输出经过的缓冲区

关于加速还差一句,竞赛代码里常见的三行是关同步、解开 cin 与 cout 的绑定、再关掉 endl,本节用的 g++ 4.8.5 默认按 C++98 编译,写 `cin.tie(nullptr)` 会直接报错。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ buf_sync_off_newline.cpp -o buf_sync_off_newline
buf_sync_off_newline.cpp: In function 'int main()':
buf_sync_off_newline.cpp:6:13: error: 'nullptr' was not declared in this scope
     cin.tie(nullptr);
             ^

nullptr 是 C++11 引入的关键字,默认标准里没有这个名字,这一篇的第二十四节会专门讲它。要在这台机器上用,编译命令得带 -std=c++11。

七、缺省参数

缺省参数让一个函数在调用处少写几个实参。写函数的人给某个参数一个默认值,调用的人不传这个参数时,默认值自动补上。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/default_arg.cpp
#include <iostream>
using namespace std;

void Func(int a = 0)
{
    cout << "Func a = " << a << endl;
}

void Func1(int a = 10, int b = 20, int c = 30)
{
    cout << "a = " << a << " b = " << b << " c = " << c << endl;
}

void Func2(int a, int b = 10, int c = 20)
{
    cout << "a = " << a << " b = " << b << " c = " << c << endl;
}

int main()
{
    Func();
    Func(10);
    Func1();
    Func1(1);
    Func1(1, 2);
    Func1(1, 2, 3);
    Func2(100);
    Func2(100, 200);
    Func2(100, 200, 300);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ default_arg.cpp -o default_arg && ./default_arg
Func a = 0
Func a = 10
a = 10 b = 20 c = 30
a = 1 b = 20 c = 30
a = 1 b = 2 c = 30
a = 1 b = 2 c = 3
a = 100 b = 10 c = 20
a = 100 b = 200 c = 20
a = 100 b = 200 c = 300

Func1 三个参数全有默认值,调用处可以一个都不传,也可以传满。Func2 只有后两个有默认值,调用处至少要给一个。这两种写法在说法上分别叫全缺省与半缺省。

半缺省有一条硬规则,默认值必须从右往左连续给到最右边,写成跳跃的样式编译器直接拒绝。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ def_skip.cpp -o def_skip
def_skip.cpp: In function 'void Func(int, int, int)':
def_skip.cpp:1:6: error: default argument missing for parameter 2 of 'void Func(int, int, int)'
 void Func(int a = 1, int b, int c = 3) {}
      ^

原因是调用点只能从左往右依次传参,第二个参数没有默认值而第一个有,写 Func(2) 时编译器无法知道 2 是想给第一个参数还是第二个参数,让右侧连续,这条歧义就不存在。

第二条规则管默认值写在哪里,声明与定义只能有一处写默认值,通常写在声明里,也就是头文件里。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/def_dup.h
#pragma once
void STInit(int* ps, int n = 4);
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -c def_dup.cpp -o def_dup.o
def_dup.cpp: In function 'void STInit(int*, int)':
def_dup.cpp:2:31: error: default argument given for parameter 2 of 'void STInit(int*, int)' [-fpermissive]
 void STInit(int* ps, int n = 4)
                               ^
In file included from def_dup.cpp:1:0:
def_dup.h:1:6: error: after previous specification in 'void STInit(int*, int)' [-fpermissive]
 void STInit(int* ps, int n = 4);
      ^

报错里那句 after previous specification 的意思是同一个参数的默认值已经在前面的声明里给过一次。把默认值从实现文件里删掉,只留声明里的那一份就通过了。

反过来也不行,默认值只写在实现文件里,调用方看不到它。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
g++ def_user.cpp def_impl.cpp -o def_user
def_user.cpp: In function 'int main()':
def_user.cpp:4:14: error: too few arguments to function 'void Func(int, int, int)'
     Func(1, 2);
              ^
In file included from def_user.cpp:1:0:
def_hdr.h:2:6: note: declared here
 void Func(int a, int b, int c);
      ^

编译器在处理调用点时只看得到声明,声明里没有默认值,传两个实参就是太少。把默认值挪回头文件,同样的两句调用就编得过。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
g++ def_user.cpp def_impl.cpp -o def_user && ./def_user
1 2 10
1 2 3

一句话概括这条规则的由来:默认值不是运行时去查的东西,它是编译期往调用点填的一个常量,下面这份汇编是上面两次调用的现场。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
	movl	$30, %edx
	movl	$20, %esi
	movl	$10, %edi
	call	_Z5Func1iii
	movl	$30, %edx
	movl	$20, %esi
	movl	$1, %edi
	call	_Z5Func1iii
	movl	$20, %edx
	movl	$10, %esi
	movl	$100, %edi
	call	_Z5Func2iii

Func1() 那一行在汇编里是先压 30、20、10 三个立即数再 call,跟写 Func1(10, 20, 30) 生成的指令完全一样。Func2(100) 那一行压进去的是 100、10、20,两个没传的实参被默认值填满。调用点自带了全部实参,函数体内部不需要知道哪个参数是被省略的。

由此可以推出一个使用时要注意的地方,默认值改了之后,只重新编译实现文件是不够的,调用点里压进去的还是旧值,必须把调用方一起重新编译。库的头文件里改了默认值,所有依赖它的程序都要重新构建。

实际的接口设计里,缺省参数用来给「常见取值」留一个省略的机会。栈的初始化需要一个容量,多数场合四个已足够,那就把四写成默认值,只有特殊场合才显式传一个大一点的数。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/stack_default.h
#pragma once

namespace bit
{
    typedef int STDataType;

    typedef struct Stack
    {
        STDataType* a;
        int top;
        int capacity;
    } ST;

    void STInit(ST* ps, int n = 4);
    void STDestroy(ST* ps);
    void STPush(ST* ps, STDataType x);
    void STPop(ST* ps);
    STDataType STTop(ST* ps);
    int STSize(ST* ps);
    bool STEmpty(ST* ps);
}

调用处写 STInit(&st) 就是四,写 STInit(&st, 100) 就是一百,两种写法的调用点汇编不同,函数体只有一份。

图 7 默认参数在调用点被填进去

八、缺省参数与重载冲突

缺省参数与函数重载各自都很直观,放在一起时会出现编译器无法选择的情况。两个函数名相同、参数表不同,一个靠默认值也能接收零个实参,调用处就会同时匹配两个候选。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ amb_overload.cpp -o amb_overload
amb_overload.cpp: In function 'int main()':
amb_overload.cpp:5:17: error: call of overloaded 'f1()' is ambiguous
 int main() { f1(); return 0; }
                 ^
amb_overload.cpp:5:17: note: candidates are:
amb_overload.cpp:3:6: note: void f1()
 void f1() { cout << "f1()" << endl; }
      ^
amb_overload.cpp:4:6: note: void f1(int)
 void f1(int a = 10) { cout << "f1(int a)" << endl; }
      ^

一个是不带参数的 f1,一个是带一个默认值的 f1(int),调用 f1() 时两个候选都能接住零个实参,编译器不替人做这个选择,直接报二义。把带默认值的那个函数的默认值去掉,或者把零参数的版本改个名字,冲突就消失了。

同一类冲突也会出现在参数个数不同的两个重载里。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ amb_default.cpp -o amb_default
amb_default.cpp: In function 'int main()':
amb_default.cpp:5:17: error: call of overloaded 'f(int)' is ambiguous
 int main() { f(1); return 0; }
                 ^
amb_default.cpp:5:17: note: candidates are:
amb_default.cpp:3:6: note: void f(int)
 void f(int a) { cout << "f(int)" << endl; }
      ^
amb_default.cpp:4:6: note: void f(int, int)
 void f(int a, int b = 0) { cout << "f(int,int)" << endl; }
      ^

f(int) 与 f(int, int = 0) 同时存在,写 f(1) 时前者正好匹配,后者靠默认值也能匹配,同样是二义。有意思的地方在于,实参传满之后两个候选的能力不再相同,编译器能分出高下。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
./amb_default2
f(int) 1
f(int,int) 1 2

f(1, 2) 只有第二个函数接得住,选中过程没有悬念;f(1) 这一行在传满实参的版本里落到了单参数的那个函数上。同一份代码里,调用点的实参个数决定了候选范围,少传一个实参反而让代码编不过。

这两节放在一起可以提炼出一条写法约定。给函数留默认值时,别让默认值的组合覆盖到另一个同名重载的能力范围,尤其是零参数与单参数这种最常见的调用形式。工程里遇到二义报错,把候选列表读一遍,多数情况下改一个默认值就够了。

图 8 默认参数与重载同时存在时的选择过程

九、函数重载的判定

同一作用域里可以有几个同名的函数,条件是参数表不同。参数表不同指参数的个数、类型或者顺序至少有一处不同,返回值不参与判定。下面这一组重载覆盖了整型提升、浮点提升与字面量类型三种情况。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp05/over_f.cpp
#include <iostream>
using namespace std;

void f(int v) { cout << "f(int)" << endl; }
void f(double v) { cout << "f(double)" << endl; }

int main()
{
    int i = 1;
    double d = 2;
    float fl = 3;
    char ch = 'x';
    short sh = 4;

    f(i);
    f(d);
    f(fl);
    f(ch);
    f(sh);
    f(3.5);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp05
g++ over_f.cpp -o over_f && ./over_f
f(int)
f(double)
f(double)
f(int)
f(int)
f(double)

六行输出对应六个实参,int 与 double 两个实参各自精确匹配到同名参数的那个版本。float 实参选中了 f(double),因为 float 到 double 属于浮点提升,比转成 int 的代价小。char 与 short 实参选中了 f(int),整型提升的优先级高于转成浮点。最后那个 3.5 的字面量类型是 double,而不是 float。

顺序不同也算不同的参数表,下面这两个函数的名字与参数类型集合完全一样,只是把两个参数的位置换了。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp05/over_h.cpp
void h(int a, double b) { cout << "h(int, double)" << endl; }
void h(double a, int b) { cout << "h(double, int)" << endl; }

int main()
{
    h(1, 2.5);
    h(1.5, 2);
    h(1, 2);
    return 0;
}

这份代码整体编不过,报错停在第三行的那次调用上,后面几行也就没有走到编译这一步。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp05
g++ over_h.cpp -o over_h
over_h.cpp: In function 'int main()':
over_h.cpp:9:11: error: call of overloaded 'h(int, int)' is ambiguous
     h(1, 2);
           ^
over_h.cpp:9:11: note: candidates are:
over_h.cpp:3:6: note: void h(int, double)
 void h(int a, double b) { cout << "h(int, double)" << endl; }
      ^
over_h.cpp:4:6: note: void h(double, int)
 void h(double a, int b) { cout << "h(double, int)" << endl; }
      ^

h(1, 2) 两个实参都是 int,第一个候选要把第二个参数从 int 转成 double,第二个候选要把第一个参数从 int 转成 double。两边各有一个参数需要转换,另一个参数正好匹配,编译器排不出谁更优,于是报二义。前两行调用本身是有答案的,h(1, 2.5) 落在 h(int, double) 上,h(1.5, 2) 落在 h(double, int) 上,只是同一份代码里存在一处无法选择的调用,整个文件就编不过了。

按值传递与按引用传递也会构成一对候选。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp04
g++ over_pick.cpp -o over_pick
over_pick.cpp: In function 'int main()':
over_pick.cpp:25:8: error: call of overloaded 'p(int&)' is ambiguous
     p(i);
        ^
over_pick.cpp:25:8: note: candidates are:
over_pick.cpp:7:6: note: void p(int)
 void p(int v) { cout << "p(int)" << endl; }
      ^
over_pick.cpp:8:6: note: void p(int&)
 void p(int& v) { cout << "p(int&)" << endl; }
      ^

对一个左值实参 i 来说,值传递与引用传递都是精确匹配,两个候选打平。这一条留到第十四节讲引用做参数时再看,那里会给出只写引用版本时的调用结果。

返回值不同不能让两个函数区分开,这一点编译期就会拦住。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/ov_ret.cpp
void fxx() {}
int fxx() { return 0; }
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -c ov_ret.cpp -o ov_ret.o
ov_ret.cpp: In function 'int fxx()':
ov_ret.cpp:2:9: error: new declaration 'int fxx()'
 int fxx() { return 0; }
         ^
ov_ret.cpp:1:6: error: ambiguates old declaration 'void fxx()'
 void fxx() {}
      ^

函数名加参数表组成一个函数的签名,返回值不在签名里。调用点决定用哪个函数时看的是实参,如果返回值也参与判定,编译器就得先知道调用处想要什么类型,这个方向在语法上走不通。

当两个候选各有优劣时,报错信息会把两边都列出来。下面这份代码里两个 Add 的参数表完全不同,调用处给的是 int 与 double 各一个。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -c overload.cpp -o overload.o
overload.cpp: In function 'int main()':
overload.cpp:21:17: error: call of overloaded 'Add(int, double)' is ambiguous
     Add(10, 20.5);
                 ^
overload.cpp:21:17: note: candidates are:
overload.cpp:3:5: note: int Add(int, int)
 int Add(int left, int right)
     ^
overload.cpp:8:8: note: double Add(double, double)
 double Add(double left, double right)
        ^

Add(int, int) 需要把 20.5 截断成 int,Add(double, double) 需要把 10 提升成 double,两个候选各让一步,平手。

把上面几条合起来就是重载决议的判定顺序:先找出所有能接住这组实参的同名函数,再逐个参数比较转换的代价,只有当某个候选在每一项上都不比别人差、并且至少有一项严格更优时它才被选中,出现并列就报二义。

十、重载的底层是名字修饰

C 语言编译出来的符号名与源文件里的函数名一致,一个名字只能对应一个函数。C++ 允许同名函数,编译器必须把参数信息编进符号名,这套编码叫名字修饰(name mangling)。下面这份文件里有四组函数,其中两个函数同名不同参数,另有一个同名函数放在域里。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/mangling.cpp
void f(int a, double b) {}
void f(char* p) {}
int Add(int left, int right) { return left + right; }

namespace bit
{
    int Add(int left, int right) { return left + right; }
}

int main() { return 0; }
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -c mangling.cpp -o mangling.o
nm mangling.o | grep -E " T | t "
0000000000000040 T main
0000000000000000 T _Z1fid
000000000000000e T _Z1fPc
0000000000000018 T _Z3Addii
000000000000002c T _ZN3bit3AddEii

符号的读法是:_Z 是 C++ 符号的前缀,跟着的数字是函数名的长度,再后面是函数名本身,最后是参数的编码。_Z1fid 读作 f(int, double),其中 i 表示 int,d 表示 double;_Z1fPc 里的 P 表示指针,c 表示 char;_Z3Addii 是 Add(int, int);带域限定的名字被包在 N 与 E 之间,_ZN3bit3AddEii 是 bit 域里的 Add(int, int)。nm 加一个 -C 参数就能把这些符号还原成源文件里的写法。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
nm -C mangling.o | grep -E " T | t "
0000000000000040 T main
0000000000000000 T f(int, double)
000000000000000e T f(char*)
0000000000000018 T Add(int, int)
000000000000002c T bit::Add(int, int)

同样两个函数用 C 的规则编,符号名只剩函数名本身。

c 复制代码
// 文件:/home/cocatrice/cpp01/exp01/mangling_pure.c
void f(int a, double b) {}
int Add(int left, int right) { return left + right; }
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
gcc -c mangling_pure.c -o mangling_pure.o && nm mangling_pure.o
000000000000000e T Add
0000000000000000 T f

C 语言里没有重载,两个同名函数本来就编不过,符号名按原名生成不会撞车。C++ 里如果没有这层编码,f(int, double) 与 f(char*) 会撞成同一个符号 f,链接器会报重复定义。重载能编过、能跨文件找到正确的定义,靠的都是修饰后的符号名。

参数是引用还是指针也会写进签名,下面这两个函数的参数类型在 C 语言里无法共存,C++ 靠修饰把它们分开。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/mangling_ref.cpp
void Swap(int& rx, int& ry) {}
void Swap(int* px, int* py) {}
int main() { return 0; }
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -c mangling_ref.cpp -o mangling_ref.o && nm -C mangling_ref.o
000000000000001c T main
000000000000000e T Swap(int*, int*)
0000000000000000 T Swap(int&, int&)
nm mangling_ref.o | grep -i swap
000000000000000e T _Z4SwapPiS_
0000000000000000 T _Z4SwapRiS_

Ri 表示 int&,Pi 表示 int*,末尾的 S_ 表示这个参数与上一个参数类型相同,所以 _Z4SwapRiS_ 就是 Swap(int&, int&)。这一对函数在第十四节讲引用做参数时会用上。

重载拆到多个源文件里之后,链接器认的就是修饰后的名字。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
--- 定义文件里的符号 ---
0000000000000000 T _Z1fid
0000000000000000 T _Z1fPc
--- 调用文件里的符号 ---
                 U _Z1fid
                 U _Z1fPc

两个定义分别落在两个目标文件里,主文件对这两个名字只有未定义的引用,链接器按名字把定义接上去,程序输出 f(int, double) 与 f(char*)。

图 9 函数名与参数类型怎么变成一个符号

修饰规则由编译器决定,不同编译器编出来的 C++ 符号互不通用,C++ 的目标文件因此不能随意跨编译器混链。按 C 规则生成的符号没有这个问题,下一节的 extern "C" 用的就是这个性质。

十一、extern "C" 的用途

C 与 C++ 混编时,两边对符号名的处理方式不同,C++ 里直接声明一个 C 函数再调用,链接期就会找不到定义。下面这个 C 文件里的加法函数用 gcc 编成目标文件。

c 复制代码
// 文件:/home/cocatrice/cpp01/exp01/c_add.c
int c_add(int a, int b)
{
    return a + b;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
gcc -c c_add.c -o c_add.o && nm c_add.o
0000000000000000 T c_add

符号名是 c_add,没有修饰,C++ 这边按普通函数的写法声明并且调用它。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/cpp_call_err.cpp
#include <iostream>
using namespace std;

int c_add(int a, int b);

int main()
{
    cout << c_add(3, 4) << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ cpp_call_err.cpp c_add.o -o cpp_call_err
/tmp/ccjdSCZD.o: In function `main':
cpp_call_err.cpp:(.text+0xf): undefined reference to `c_add(int, int)'
collect2: error: ld returned 1 exit status

链接器要的是 c_add(int, int) 这个修饰过的名字,c_add.o 里只有 c_add,两边对不上。把声明改到 extern "C" 里,编译器对这个名字不再做修饰,链接就能通过。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/cpp_call_ok.cpp
#include <iostream>
using namespace std;

extern "C" int c_add(int a, int b);

int main()
{
    cout << c_add(3, 4) << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ cpp_call_ok.cpp c_add.o -o cpp_call_ok && ./cpp_call_ok
7

反过来的场景是把 C++ 写的函数交给 C 程序调用。函数定义上加 extern "C",符号名退回原名,同一个文件里没加的那个函数仍然带着修饰。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/cpp_export.cpp
extern "C" int cpp_add(int a, int b) { return a + b; }
int cpp_sub(int a, int b) { return a - b; }
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -c cpp_export.cpp -o cpp_export.o && nm cpp_export.o | grep -E " T "
0000000000000000 T cpp_add
0000000000000014 T _Z7cpp_subii

C 那边按普通函数声明 cpp_add 就能调用,用 g++ 完成最后的链接。

c 复制代码
// 文件:/home/cocatrice/cpp01/exp01/c_call_cpp.c
#include <stdio.h>

int cpp_add(int a, int b);

int main()
{
    printf("%d\n", cpp_add(5, 6));
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
gcc -c c_call_cpp.c -o c_call_cpp.o && g++ c_call_cpp.o cpp_export.o -o c_call_cpp && ./c_call_cpp
11

extern "C" 里的名字不能重载。C 语言里一个符号只对应一个函数,写两个 extern "C" 的同名函数会当成重复定义。需要给 C 与 C++ 共用的头文件,惯常的写法是把声明包在条件编译里。

c 复制代码
// 文件:/home/cocatrice/cpp01/exp07/export_hdr.h
#ifdef __cplusplus
extern "C" {
#endif

int cpp_add(int a, int b);
void cpp_reset(void);

#ifdef __cplusplus
}
#endif

__cplusplus 这个宏只在 C++ 编译器里定义,C 编译器看不到 extern "C" 那一对花括号,两边可以包含同一个头文件。C 标准库的头文件在 C++ 里的用法就是这个模式,math.h、stdio.h 里的函数都能被 C++ 代码直接调用。

十二、本节的四个常见报错

这一节把前面提到的报错收在一起,都是初学阶段反复遇到的几类。

第一类是缺省参数重定义,声明与定义各写了一份默认值,编译器在处理实现文件时报错。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
def_dup.cpp: In function 'void STInit(int*, int)':
def_dup.cpp:2:31: error: default argument given for parameter 2 of 'void STInit(int*, int)' [-fpermissive]
 void STInit(int* ps, int n = 4)
                               ^
In file included from def_dup.cpp:1:0:
def_dup.h:1:6: error: after previous specification in 'void STInit(int*, int)' [-fpermissive]
 void STInit(int* ps, int n = 4);
      ^

第二类是调用二义,两个候选都能接住这组实参,编译器排不出先后。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
amb_default.cpp: In function 'int main()':
amb_default.cpp:5:17: error: call of overloaded 'f(int)' is ambiguous
 int main() { f(1); return 0; }
                 ^
amb_default.cpp:5:17: note: candidates are:
amb_default.cpp:3:6: note: void f(int)
 void f(int a) { cout << "f(int)" << endl; }
      ^
amb_default.cpp:4:6: note: void f(int, int)
 void f(int a, int b = 0) { cout << "f(int,int)" << endl; }
      ^

第三类是没有匹配的重载,实参个数或者类型与所有候选都对不上。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp06/no_match.cpp
#include <iostream>
using namespace std;

void Add(int left, int right) { cout << "Add(int, int)" << endl; }
void Add(double left, double right) { cout << "Add(double, double)" << endl; }

int main()
{
    Add(1, 2, 3);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp06
g++ no_match.cpp -o no_match
no_match.cpp: In function 'int main()':
no_match.cpp:7:16: error: no matching function for call to 'Add(int, int, int)'
     Add(1, 2, 3);
                ^
no_match.cpp:7:16: note: candidates are:
no_match.cpp:3:6: note: void Add(int, int)
 void Add(int left, int right) { cout << "Add(int, int)" << endl; }
      ^
no_match.cpp:3:6: note:   candidate expects 2 arguments, 3 provided
no_match.cpp:4:6: note: void Add(double, double)
 void Add(double left, double right) { cout << "Add(double, double)" << endl; }
      ^
no_match.cpp:4:6: note:   candidate expects 2 arguments, 3 provided

把实参换成一个字符串常量,报错里的类型写法会跟着变。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp06
no_match2.cpp:7:14: error: no matching function for call to 'Add(const char [4])'
     Add("abc");
              ^
no_match2.cpp:7:14: note: candidates are:
no_match2.cpp:3:6: note: void Add(int, int)
no_match2.cpp:3:6: note:   candidate expects 2 arguments, 1 provided
no_match2.cpp:4:6: note: void Add(double, double)
no_match2.cpp:4:6: note:   candidate expects 2 arguments, 1 provided

第四类是未声明的标识符,没有包含声明 cout 与 endl 的头文件,名字就找不到。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp06
no_include.cpp: In function 'int main()':
no_include.cpp:4:5: error: 'cout' was not declared in this scope
     cout << "hi" << endl;
     ^
no_include.cpp:4:21: error: 'endl' was not declared in this scope
     cout << "hi" << endl;
                     ^

用了没写过的函数名也是同一条报错,名字的查找只认已经出现过的声明,域里的名字还要求把域写对。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp06
no_decl.cpp:3:11: error: 'Print' was not declared in this scope
     Print();
           ^
no_member.cpp:8:20: error: 'Sub' is not a member of 'bit'
     printf("%d\n", bit::Sub(3, 4));
                    ^

编译器并不总是只报一句就结束,名字拼写相近时它会把候选写出来。第二节的例子里忘记写域限定,报错后面就跟了一条建议。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
ns_use_err.cpp:9:20: error: 'a' was not declared in this scope
     printf("%d\n", a);
                    ^
ns_use_err.cpp:9:20: note: suggested alternative:
ns_use_err.cpp:4:9: note:   'bit::a'
     int a = 0;
         ^

把 suggested alternative 那一行读出来,多数时候加上域限定名就能编过。

十三、引用的定义

引用(reference)是给一块已经存在的内存空间取一个别名,写法是在类型后面加一个 &。定义引用时必须指明它是谁的别名,之后对这个名字的所有操作都落在被引用的那块空间上。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/ref_basic.cpp
#include <iostream>
using namespace std;

int main()
{
    int a = 0;
    int& b = a;
    int& c = a;
    int& d = b;
    ++d;
    cout << "a = " << a << " b = " << b << " c = " << c << " d = " << d << endl;
    cout << &a << " " << &b << " " << &c << " " << &d << endl;

    int e = 20;
    cout << "改之前 b 的地址 " << &b << " 值 " << b << endl;
    b = e;
    cout << "b = e 之后 b 的地址 " << &b << " 值 " << b << " a 的值 " << a << endl;
    cout << "e 的地址 " << &e << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ ref_basic.cpp -o ref_basic && ./ref_basic
a = 1 b = 1 c = 1 d = 1
0x7fff5f0a81f4 0x7fff5f0a81f4 0x7fff5f0a81f4 0x7fff5f0a81f4
改之前 b 的地址 0x7fff5f0a81f4 值 1
b = e 之后 b 的地址 0x7fff5f0a81f4 值 20 a 的值 20
e 的地址 0x7fff5f0a81f0

第一行里 a、b、c、d 四个名字打出同一个值,第二行四个取地址运算打出同一个地址,说明它们指向的就是同一块空间。d 是 b 的引用,b 又是 a 的引用,引用套引用最后仍然落在 a 上,所以 ++d 之后 a 变成了 1。

第三行到第五行是引用的另一条性质,b = e 这句看着像让 b 改指到 e 上,实际做的是把 e 的值赋给 b 引用的那块空间。执行之后 b 的地址没有变化,a 的值变成了 20,e 自己的地址是另一个。引用一旦初始化就绑死在一个对象上,后面所有的赋值都只是在改那个对象的值。

定义引用时不写初始值,编译器直接拦下来。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ ref_uninit.cpp -o ref_uninit
ref_uninit.cpp: In function 'int main()':
ref_uninit.cpp:3:10: error: 'ra' declared as reference but not initialized
     int& ra;
           ^

三条性质合起来就是引用的全部规则:定义时必须初始化,初始化之后不能改绑,对它做的运算全部作用在被引用的对象上。

图 10 引用与被引用的对象共用同一块空间

十四、引用与指针的对照

引用与指针经常被放在一起比较,两者在语法上有三处明显的不同,在汇编层面却是同一件事。

第一处是定义,引用要用 int& b = a; 的形式绑到一个已经存在的对象上,指针要写成 int* p = &a;,取地址符号不能少。第二处是使用,对引用做的读写直接写名字,指针要先解引用。第三处是约束,引用必须在定义时绑好、之后不能改绑、也没有空引用这种状态;指针可以不初始化、可以改指向别的对象、还可以是空指针。

汇编层面的证据在第十七节,那里会把同一个函数按引用传参与按指针传参生成的指令摆在一起。这一节先看一个边界情况:空指针解引用后取引用。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/null_ref.cpp
int main()
{
    int* p = 0;
    int& r = *p;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -O0 null_ref.cpp -o null_ref && ./null_ref; echo "退码=$?"
退码=0

这段代码编得过也跑得完,退码是 0,从空指针取到一个引用在语法上没有被禁止,所谓引用不能为空,是写代码的人要守住的一条约定,编译器没有为它生成检查。真的去读写这个引用,访问的还是地址 0,程序会立刻崩掉。

指针与引用在符号名上的区别可以再对照一次。第十节里那一对 Swap 的符号是 _Z4SwapRiS_ 与 _Z4SwapPiS_,R 与 P 标出了参数是引用还是指针,链接器靠这个区分两个同名函数。这也是引用与指针在语言里被视为两种不同参数类型的直接体现。

十五、引用做参数

引用最常见的一个用途是当函数参数,同一个 Swap,用指针写要传地址、函数体里要解引用,用引用写这两步都可以省掉。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/swap_ref.cpp
#include <iostream>
using namespace std;
void Swap(int& rx, int& ry)
{
    int tmp = rx;
    rx = ry;
    ry = tmp;
}

int main()
{
    int x = 0, y = 1;
    cout << "调用前 " << x << " " << y << endl;
    Swap(x, y);
    cout << "调用后 " << x << " " << y << endl;
    return 0;
}
cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/swap_ptr.cpp
#include <iostream>
using namespace std;
void Swap(int* px, int* py)
{
    int tmp = *px;
    *px = *py;
    *py = tmp;
}

int main()
{
    int x = 0, y = 1;
    cout << "调用前 " << x << " " << y << endl;
    Swap(&x, &y);
    cout << "调用后 " << x << " " << y << endl;
    return 0;
}

两个版本分别在 -std=c++98 与 -std=c++11 下各编一次。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
--- -std=c++98 ---
调用前 0 1
调用后 1 0
调用前 0 1
调用后 1 0
--- -std=c++11 ---
调用前 0 1
调用后 1 0
调用前 0 1
调用后 1 0

四组输出完全一样,说明引用版与指针版改到实参的效果相同,两个标准下的行为也相同。差别只落在写法上:调用处 Swap(x, y) 与 Swap(&x, &y),函数体里 rx = ry 与 *px = *py。引用把取地址与解引用这两件事交给编译器去做,写代码的人只需要把参数当成普通变量用。

栈的接口换成引用传参之后,声明与调用处就变成下面这样。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp02/ret_ref.cpp
void STInit(ST& rs, int n = 4)
{
    rs.a = (STDataType*)malloc(n * sizeof(STDataType));
    rs.top = 0;
    rs.capacity = n;
}

int main()
{
    ST st1;
    STInit(st1);
    STInit(st1, 100);
    STPush(st1, 1);
    return 0;
}

STInit(st1) 把栈对象本身交给函数,函数里对 rs.a、rs.top、rs.capacity 的修改都落在 st1 上。调用处不需要写 &st1,STPush(st1, 1) 也一样。缺省参数与引用在同一份声明里同时用上,前一个省掉实参,后一个省掉取地址。

有一类参数必须写成指针的引用,就是指针本身要被函数修改的时候。链表头插要把新的头结点挂到链表上,头指针变量的值要跟着变。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp09/list_ref.cpp
typedef int LTDataType;
typedef struct LTNode
{
    LTDataType data;
    LTNode* next;
} LTNode;

LTNode* BuyNode(LTDataType x)
{
    LTNode* node = (LTNode*)malloc(sizeof(LTNode));
    node->data = x;
    node->next = NULL;
    return node;
}

void LTPushFront(LTNode*& phead, LTDataType x)
{
    LTNode* node = BuyNode(x);
    node->next = phead;
    phead = node;
}

int main()
{
    LTNode* plist = NULL;
    LTPushFront(plist, 1);
    LTPushFront(plist, 2);
    LTPushFront(plist, 3);
    LTPrint(plist);
    cout << "头指针本身的值 " << plist << ",头结点的值 " << plist->data << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp09
g++ list_ref.cpp -o list_ref && ./list_ref
3 2 1 
头指针本身的值 0x1651c60,头结点的值 3

phead = node; 改的是调用方那个指针变量。参数写成 LTNode* phead 就不行,函数里的 phead 只是一份拷贝,改完出了函数就没了,链表头还是空的。同一件事用二级指针也能做,参数改成 LTNode** pphead,函数体里每个 phead 都要多写一个星号。

两种写法在符号名上分得也很清楚,引用版本与指针版本各自有一串自己的编码。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp09
nm -C list_ref.o | grep LTPushFront
0000000000000034 T LTPushFront(LTNode*&, int)
nm -C list_ptr.o | grep LTPushFront
0000000000000000 T LTPushFront(LTNode**, int)

不带 -C 时它们分别是 _Z11LTPushFrontRP6LTNodei 与 _Z11LTPushFrontPP6LTNodei。R 后面跟的是被引用类型 P6LTNode,也就是「指向 LTNode 的指针」的引用;后一个是 P6LTNode 再套一层指针。这与第十节里 _Z4SwapRiS_ 与 _Z4SwapPiS_ 的区分规则一致。

引用参数有一个限制:只写引用版时,字面量传不进去。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp08/lit_ref.cpp
void q(int& v) { cout << "q(int&)" << endl; }

int main()
{
    q(30);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp08
g++ lit_ref.cpp -o lit_ref
lit_ref.cpp: In function 'int main()':
lit_ref.cpp:6:9: error: invalid initialization of non-const reference of type 'int&' from an rvalue of type 'int'
     q(30);
         ^
lit_ref.cpp:3:6: error: in passing argument 1 of 'void q(int&)'
 void q(int& v) { cout << "q(int&)" << endl; }
      ^

30 是一个常量,没有对应的内存空间,自然没有别名的落脚处。把参数改成 const int& 之后同一个调用就能编过,编译为它开一块临时空间,这就是下一节的内容。

十六、引用做返回值

返回值也可以写成引用,写 int STTop(ST&) 时函数把栈顶元素的值拷贝一份给调用方,写 int& STTop(ST&) 时返回的是那块空间本身的别名。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp02/ret_ref.cpp
typedef int STDataType;
typedef struct Stack
{
    STDataType* a;
    int top;
    int capacity;
} ST;

void STPush(ST& rs, STDataType x)
{
    if (rs.top == rs.capacity)
    {
        int newcapacity = rs.capacity == 0 ? 4 : rs.capacity * 2;
        STDataType* tmp = (STDataType*)realloc(rs.a, newcapacity * sizeof(STDataType));
        if (tmp == NULL)
        {
            perror("realloc fail");
            return;
        }
        rs.a = tmp;
        rs.capacity = newcapacity;
    }
    rs.a[rs.top] = x;
    rs.top++;
}

int& STTop(ST& rs)
{
    return rs.a[rs.top - 1];
}

int main()
{
    ST st1;
    STInit(st1);
    STPush(st1, 1);
    STPush(st1, 2);
    cout << STTop(st1) << endl;
    STTop(st1) += 10;
    cout << STTop(st1) << endl;
    STTop(st1) = 4;
    cout << STTop(st1) << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
g++ ret_ref.cpp -o ret_ref && ./ret_ref
2
12
4

第一次打印栈顶元素 2,第二次的 STTop(st1) += 10 能编过,正因为 STTop(st1) 是一个和 rs.a[rs.top - 1] 等价的左值,加 10 直接改在栈里,第三次打印变成 12。第三次的 STTop(st1) = 4 把新值写进同一块空间,第四次打印变成 4。如果返回值那行写成 int STTop(ST& rs),这两句都编不过,因为函数调用的结果是一个临时量,不能赋值。

返回引用与返回值在汇编里的区别很小但很关键。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp02/ret_asm.cpp
int g_a = 1;
int& ret_ref(int& r) { return r; }
int ret_val(int v) { return v; }
int& ret_global() { return g_a; }
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
--- ret_ref ---
_Z7ret_refRi:
.LFB0:
	pushq	%rbp
	movq	%rsp, %rbp
	movq	%rdi, -8(%rbp)
	movq	-8(%rbp), %rax
	popq	%rbp
	ret
--- ret_val ---
_Z7ret_vali:
.LFB1:
	pushq	%rbp
	movq	%rsp, %rbp
	movl	%edi, -4(%rbp)
	movl	-4(%rbp), %eax
	popq	%rbp
	ret
--- ret_global ---
_Z10ret_globalv:
.LFB2:
	pushq	%rbp
	movq	%rsp, %rbp
	movl	$g_a, %eax
	popq	%rbp
	ret

三个函数都通过 eax 把结果交给调用方。ret_ref 交给调用方的是参数里那个对象的地址,ret_global 交给调用方的是全局变量 g_a 的地址,ret_val 交给调用方的是值本身。返回引用时返回的是一个地址,调用方拿到它之后可以继续在这个地址上读写,这就是 STTop(st1) = 4 能编译出来的原因。

返回值写成引用时有一条硬性要求:返回的对象必须活得比这个函数调用久。全局变量、静态变量、堆上申请的空间、参数传进来的对象都满足,函数里的局部变量不满足。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/ret_dangling.cpp
int& func()
{
    int a = 0;
    return a;
}

int main()
{
    int& ra = func();
    cout << ra << endl;
    cout << ra << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -Wall ret_dangling.cpp -o ret_dangling
ret_dangling.cpp: In function 'int& func()':
ret_dangling.cpp:5:9: warning: reference to local variable 'a' returned [-Wreturn-local-addr]
     int a = 0;
         ^
./ret_dangling; echo "运行结束退码=$?"
0
32564
运行结束退码=0

编译器只给了一条警告,程序照样编出来也照样跑完,退码是 0。func 里的 a 在函数返回时已经销毁,它占的那块栈空间紧接着被 cout 的调用覆盖,所以第二次打印打出来的是 32564。这种结果每次运行都可能不同,也可能当场崩掉。把函数那行改成 int func(),同一段逻辑打出来的就是两个 0。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -Wall ret_value.cpp -o ret_value && ./ret_value
0

-Wall 是必须加的一个开关,上面那条警告默认不开,不加 -Wall 时编译器一个字都不说,代码里留着一个悬空引用也看不出来。

十七、值与引用的效率

这一节把引用参数与指针参数生成的机器码放在一起对照。两个函数一个按引用收参,一个按指针收参,函数体做的事情相同。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/asm_ref_ptr.cpp
void by_ref(int& x) { x++; }
void by_ptr(int* p) { (*p)++; }
void call_ref()
{
    int a = 1;
    by_ref(a);
}
void call_ptr()
{
    int a = 1;
    by_ptr(&a);
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
--- by_ref ---
_Z6by_refRi:
.LFB0:
	pushq	%rbp
	movq	%rsp, %rbp
	movq	%rdi, -8(%rbp)
	movq	-8(%rbp), %rax
	movl	(%rax), %eax
	leal	1(%rax), %edx
	movq	-8(%rbp), %rax
	movl	%edx, (%rax)
	popq	%rbp
	ret
--- by_ptr ---
_Z6by_ptrPi:
.LFB1:
	pushq	%rbp
	movq	%rsp, %rbp
	movq	%rdi, -8(%rbp)
	movq	-8(%rbp), %rax
	movl	(%rax), %eax
	leal	1(%rax), %edx
	movq	-8(%rbp), %rax
	movl	%edx, (%rax)
	popq	%rbp
	ret

两个函数各 13 行,逐行一致,连寄存器都用同一套。参数都从 %rdi 取,都先存到 -8(%rbp) 这个栈位,再取出来读写指向的整数,调用点的两段汇编同样只有符号名不同。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
--- call_ref ---
_Z8call_refv:
.LFB2:
	pushq	%rbp
	movq	%rsp, %rbp
	subq	$16, %rsp
	movl	$1, -4(%rbp)
	leaq	-4(%rbp), %rax
	movq	%rax, %rdi
	call	_Z6by_refRi
	leave
	ret
--- call_ptr ---
_Z8call_ptrv:
.LFB3:
	pushq	%rbp
	movq	%rsp, %rbp
	subq	$16, %rsp
	movl	$1, -4(%rbp)
	leaq	-4(%rbp), %rax
	movq	%rax, %rdi
	call	_Z6by_ptrPi
	leave
	ret

call_ref 里没有取地址运算,call_ptr 里写的是 by_ptr(&a),两段汇编都先用 leaq -4(%rbp), %rax 算出 a 的地址再放进 %rdi。也就是说,取地址这件事在引用版里同样发生了,只是编译器替你写。相对的,指针版里的解引用在引用版里也照样发生,movq -8(%rbp), %rax 取到的就是那个地址。

sizeof 打出来的几个数字印证的是同一件事,写 sizeof 的时候编译器取的是被引用对象的类型大小。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ sizeof_ref.cpp -o sizeof_ref && ./sizeof_ref
sizeof(a) = 4
sizeof(ra) = 4
sizeof(int&) = 4
sizeof(p) = 8
sizeof(int*) = 8

sizeof(ra) 与 sizeof(a) 都是 4,引用本身不占额外空间,对它做 sizeof 拿到的是被引用类型的大小。指针变量本身是 8 个字节,那是一台 64 位机器上地址的宽度。所以引用不是指针的变体,它在语法层没有独立身份,编译器只在需要的时候把地址放进寄存器。

图 11 引用参数与指针参数生成的汇编逐行对照

由此可以得出这一节要说的结论:引用参数相对指针参数在运行期没有额外开销,也没有额外收益,机器码是一样的。引用真正省下的是书写上的两步,以及它把「这个参数一定有效」写进了类型里。要谈效率,比较的对象应该是传值与传引用,一次传值意味着实参被完整拷贝一份,对象大的时候这个拷贝才是开销所在。

把这个结论落到一个结构体上,对比就更直观。Big 里有一个 64 字节的字符数组和一个 int,一共 68 字节。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp10/big_arg.cpp
struct Big
{
    char buf[64];
    int n;
};
int sum1(Big b)
{
    return b.n;
}
int sum2(const Big& b)
{
    return b.n;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp10
# 所属目录:/home/cocatrice/cpp01/exp10
7 7
--- sum1 按值传参的函数体指令数 ---
7
--- sum2 按引用传参的函数体指令数 ---
9
--- 整体汇编里 mov 类指令的总条数 ---
47
--- 调用点:传值 ---
67-	movq	-80(%rbp), %rax
68-	movq	%rax, 16(%rsp)
69-	movq	-72(%rbp), %rax
70-	movq	%rax, 24(%rsp)
71-	movq	-64(%rbp), %rax
72-	movq	%rax, 32(%rsp)
73-	movq	-56(%rbp), %rax
74-	movq	%rax, 40(%rsp)
75-	movq	-48(%rbp), %rax
76-	movq	%rax, 48(%rsp)
77-	movq	-40(%rbp), %rax
78-	movq	%rax, 56(%rsp)
79-	movl	-32(%rbp), %eax
80-	movl	%eax, 64(%rsp)
81:	call	_Z4sum13Big
--- 调用点:传引用 ---
53-	movq	%rsp, %rbp
54-	.cfi_def_cfa_register 6
55-	pushq	%rbx
56-	subq	$168, %rsp
57-	.cfi_offset 3, -24
58-	movl	$7, -32(%rbp)
59-	leaq	-96(%rbp), %rax
60-	movq	%rax, %rdi
61:	call	_Z4sum2RK3Big
ax, %rdi
61:	call	_Z4sum2RK3Big

两个函数打出来的都是 7,传值那一版在调用点用了十四条 mov 类指令把整个结构体搬到栈上,传引用那一版只用了两条 mov 与一条取地址指令。函数体本身的长度差不多,一处是七行,一处是九行,差的那些行是在按地址取字段。真正被省掉的是调用点的整份拷贝,这个拷贝的字节数等于实参的大小,对象越大省下的越多。

十八、const 引用

在引用前面加一个 const,得到的就是 const 引用。它限制的是通过这个名字修改对象的能力,被引用的对象本身不受影响。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/const_ref.cpp
#include <iostream>
using namespace std;

int main()
{
    const int a = 10;
    const int& ra = a;
    int b = 20;
    const int& rb = b;
    b++;
    cout << "a = " << a << " rb = " << rb << " b = " << b << endl;

    const int& rc = 30;
    const int& rd = a * 3;
    double d = 12.34;
    const int& re = d;
    cout << "rc = " << rc << " rd = " << rd << " re = " << re << " d = " << d << endl;
    cout << &rc << " " << &rd << " " << &re << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ const_ref.cpp -o const_ref && ./const_ref
a = 10 rb = 21 b = 21
rc = 30 rd = 30 re = 12 d = 12.34
0x7ffc9c410a24 0x7ffc9c410a28 0x7ffc9c410a2c

第一行里 b 自增之后 rb 跟着变成 21,说明 rb 引用的仍然是 b 那块空间,const 只是不允许通过 rb 去改它。第二行三个引用各自打出了值,第三行的三个地址互不相同,说明编译器为每个临时量单独开了一块空间:30 是一个 int 临时量,a * 3 的结果是另一个 int 临时量,double 类型的 d 被转成 int 之后又是一个临时量。re 打出 12 而不是 12.34,这个截断就发生在临时量上。

引用的权限可以缩小,不能放大,三种放大写法各对应一条报错。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/const_ref_err.cpp
int main()
{
    const int a = 10;
    int& ra = a;
    return 0;
}
cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/const_ref_err2.cpp
int main()
{
    int a = 10;
    int& rb = a * 3;
    return 0;
}
cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/const_ref_err3.cpp
int main()
{
    double d = 12.34;
    int& rd = d;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
const_ref_err.cpp: In function 'int main()':
const_ref_err.cpp:4:15: error: invalid initialization of reference of type 'int&' from expression of type 'const int'
     int& ra = a;
               ^
const_ref_err2.cpp: In function 'int main()':
const_ref_err2.cpp:4:17: error: invalid initialization of non-const reference of type 'int&' from an rvalue of type 'int'
     int& rb = a * 3;
                 ^
const_ref_err3.cpp: In function 'int main()':
const_ref_err3.cpp:4:15: error: invalid initialization of reference of type 'int&' from expression of type 'double'
     int& rd = d;
               ^

三种情况被拒绝的原因相同:普通引用是一块可读可写的别名,绑定之后调用方随时可能通过它写数据。绑定一个 const 对象等于把只读权限放大成可写,绑定一个临时量等于要给一个马上就要消失的对象留位置,绑定一个类型不同的值等于允许通过引用改坏原始对象的字节。const 引用只读,这三种风险都不存在,所以编译器愿意为它开口子,其中前两种直接拒绝、后一种由编译器开一块临时空间再转成目标类型。

临时量的生命期也被这条规则延长了,下面这个函数按 const 引用收参,三次调用分别传了字面量、表达式结果与另一个类型的变量。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/const_ref_life.cpp
#include <iostream>
using namespace std;

int g_count = 0;
void func(const int& v)
{
    cout << "第 " << ++g_count << " 次调用 v = " << v << " 地址 " << &v << endl;
}

int main()
{
    int a = 1, b = 2;
    func(30);
    func(a + b);
    double d = 7.5;
    func(d);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ const_ref_life.cpp -o const_ref_life && ./const_ref_life
第 1 次调用 v = 30 地址 0x7ffe949b3264
第 2 次调用 v = 3 地址 0x7ffe949b3268
第 3 次调用 v = 7 地址 0x7ffe949b326c

三次调用拿到的地址各不同,每个临时量都活到了函数调用结束,函数体里读到的值是完整的。如果参数写成 int&,第三次调用会被拒绝,第二次也会被拒绝,只有第一次传进去的那个 int 左值能通过。const 引用让这三类实参都能传进来,同时守住只读这条线。

这两条性质合起来,让 const 引用成为工程里最常用的参数写法:既能接左值也能接右值,既不负责拷贝也不允许修改实参。第十五节里那个 STInit(ST& rs, int n = 4) 之所以不写成 const 引用,是因为它要在函数里修改栈对象;只读的接口,例如打印、计算、比较,参数写成 const 引用更合适。

十九、引用与指针的关系

把前几节的结论放在一起,引用与指针的关系可以概括成一句话:语法上用两个名字表达两件事,机器码里是同一件事。

同一件事的证据在第十七节:by_ref 与 by_ptr 两段汇编逐行一致,调用点都是一条 leaq 加一条 movq 把地址送进 %rdi,sizeof 的结果也印证引用没有自己独立的空间。

两件事指的是它们在语言层面完全不同的约束。指针是一个对象,它自己有地址、能改指向、能为空、能与整数做运算;引用不是对象,没有自己的地址,不能改绑,不能为空。这些差别在编译器需要生成代码的地方显现出来。

引用不能组成数组,这一点在声明时就会被拒绝。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp08/ref_array.cpp
int main()
{
    int a = 1, b = 2, c = 3;
    int& arr[3] = { a, b, c };
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp08
g++ ref_array.cpp -o ref_array
ref_array.cpp: In function 'int main()':
ref_array.cpp:4:15: error: declaration of 'arr' as array of references
     int& arr[3] = { a, b, c };
               ^

数组的每个元素都要有一块自己的空间,引用没有自己的空间,这条要求无法满足。指针数组是允许的,int* arr[3] 里放的是三个指针对象,每个占八个字节。对引用取地址拿到的是被引用对象的地址,这也是引用没有独立身份的一个侧面。

选用的场合可以按三条来判断,要修改实参并且这个参数一定存在,用引用;参数可能不存在、需要显式表示空的状态,用指针;要和 C 的接口对接,用指针;要修改调用方的一个指针变量,用指针的引用。剩下的一种情况是函数返回值,返回的对象由调用方负责继续使用、并且一定活得比调用久时,返回引用可以省掉一次拷贝,第十六节的 STTop 就属于这一类。

初学阶段容易写出的两种错法各有一处特征。把引用当成指针用,写成 int& r = &a;,报错会说类型不匹配;把引用当成值用,在函数里返回局部变量的引用,编译器只给一条警告。前一类错误编译期就能拦住,后一类要靠 -Wall 才能看见。

二十、inline

让一个小函数不产生调用开销,C 语言里的办法是写宏。宏由预处理器做文本替换,编译器在语法分析之前就已经看不到函数名,也就没有函数调用的开销。代价是替换发生在语法检查之前,参数没有类型,写出来的东西容易和上下文串味。

第一类问题是括号,下面两个宏做同一件事,一个在参数外加了括号,一个没有。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp02/macro_trap.cpp
#include <iostream>
using namespace std;

#define MUL_NO_PAREN(a, b) a * b
#define MUL_PAREN(a, b) ((a) * (b))

int main()
{
    cout << "MUL_NO_PAREN(1+2, 3+4) = " << MUL_NO_PAREN(1 + 2, 3 + 4) << endl;
    cout << "MUL_PAREN(1+2, 3+4) = " << MUL_PAREN(1 + 2, 3 + 4) << endl;
    int x = 1;
    cout << "MUL_NO_PAREN(x++, 2) = " << MUL_NO_PAREN(x++, 2) << ",x = " << x << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
g++ macro_trap.cpp -o macro_trap && ./macro_trap
MUL_NO_PAREN(1+2, 3+4) = 11
MUL_PAREN(1+2, 3+4) = 21
MUL_NO_PAREN(x++, 2) = 2,x = 2

同一个调用 MUL_NO_PAREN(1 + 2, 3 + 4) 打出 11,因为展开之后是 1 + 2 * 3 + 4,乘法先算。加了括号的版本展开成 ((1 + 2) * (3 + 4)),得到 21。第三个调用里的参数带了自增,展开之后 x++ 只出现一次,所以 x 从 1 变成 2,这一点在两个乘法宏上恰好没有暴露出来,换成加法形式的宏就会连续自增两次。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp03/macro_twice.cpp
#include <iostream>
using namespace std;

#define ADD(a, b) a + b

int main()
{
    int x = 1;
    int y = ADD(x++, x++);
    cout << "y = " << y << ",x = " << x << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp03
g++ macro_twice.cpp -o macro_twice && ./macro_twice
y = 3,x = 3

x 从 1 变成了 3,说明 x++ 被求值了两次。函数调用不会有这个问题,实参在传进去之前只算一次。

第二类问题是宏体里带语句时对上下文的依赖,下面这个宏交换两个变量。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp03/macro_stmt.cpp
#include <iostream>
using namespace std;

#define SWAP(a, b) int t = a; a = b; b = t;

int main()
{
    int m = 1, n = 2;
    if (m < n)
        SWAP(m, n)
    cout << m << " " << n << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp03
g++ macro_stmt.cpp -o macro_stmt
macro_stmt.cpp: In function 'int main()':
macro_stmt.cpp:3:42: error: 't' was not declared in this scope
 #define SWAP(a, b) int t = a; a = b; b = t;
                                          ^
macro_stmt.cpp:8:9: note: in expansion of macro 'SWAP'
         SWAP(m, n)
         ^

if 后面没有花括号,只有紧跟着的第一条语句属于 if,展开之后 b = t; 掉到了 if 之外,t 在那一层已经出了作用域。给宏体加上 do { } while (0) 能解决这一处,但参数被求值两次的问题仍然留在原地。

宏还有两点在检查上帮不上忙,宏没有参数类型,传进去的东西是什么类型由展开后的表达式决定,编译器报错时报的是展开之后那一行,位置指向宏定义而不是调用处,上面这条报错的 note 就是为此补的。宏也没有作用域,#define 从定义那一行开始对后面所有代码生效,直到被 #undef 取消或者文件结束,中间包含进来的头文件里如果有一个同名标识符,同样会被替换掉。

inline 函数补的正是这两处,它首先是一个函数,有参数表和返回类型,实参在调用前按类型检查并只求值一次;其次它的名字遵循普通的名字查找规则,可以放在命名空间里、可以重载、能在调试时按名字下断点。除此之外它还带来一条语法上的便利:一个 inline 函数可以在多个源文件里出现同一份定义,链接器负责把它们合成一份。这一点与它的来历有关,宏本来就是每个源文件各自展开一份的。

二十一、inline 仅为建议

声明一个函数为 inline,写给编译器的是一个建议,编译器可以接受,也可以不理会。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/inline_add.cpp
#include <iostream>
using namespace std;

inline int Add(int x, int y)
{
    int ret = x + y;
    ret += 1;
    ret += 1;
    return ret;
}

int main()
{
    cout << Add(1, 2) * 5 << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -S -O0 inline_add.cpp -o inline_add_O0.s && grep -c "call" inline_add_O0.s
6
grep "call" inline_add_O0.s | head -5
	call	_Z3Addii
	call	_ZNSolsEi
	call	_ZNSolsEPFRSoS_E
	call	_ZNSt8ios_base4InitC1Ev
	call	__cxa_atexit
g++ -S -O2 inline_add.cpp -o inline_add_O2.s && grep -c "call" inline_add_O2.s
3

-O0 的汇编里有 6 条 call,第一条就是 call _Z3Addii,函数照常被调用。同一条命令加上 -O2 之后 call 只剩 3 条,_Z3Addii 那一条消失了,函数体被复制到了调用点。同一份源码、同一个 inline 关键字,展开与否由优化档决定。

把两份汇编换成目标文件,看符号表也能看出这一层区别。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/inline_vs_plain.cpp
inline int AddInline(int x, int y) { return x + y; }
int AddPlain(int x, int y) { return x + y; }

int main()
{
    int a = AddInline(1, 2);
    int b = AddPlain(3, 4);
    return a + b;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
--- 符号表 O0 ---
0000000000000000 T _Z8AddPlainii
0000000000000000 W _Z9AddInlineii
--- 符号表 O2 ---
0000000000000000 T _Z8AddPlainii

-O0 下两个函数都在符号表里,普通函数是大写 T,inline 函数是大写 W。W 表示弱符号,同一份定义在多个目标文件里出现时链接器只保留一份,不会报重复定义。-O2 下 inline 的那个符号没有单独生成,调用点已经换成了展开后的指令;普通函数仍然留在符号表里,因为它的定义对其他目标文件可见。

反过来也成立,没有写 inline 的小函数,在优化打开之后同样可能被展开到调用点,编译器有自己的成本模型,函数体只有一两条指令、被调用的次数少或是参数是常量时,展开比自己调用更合适。所以 inline 关键字与是否展开之间没有必然联系,把它理解成一个提示更贴近事实。

展开本身也不是没有代价,函数体被复制到每一个调用点,代码段会变大。一个函数体有上百条指令,在代码里被调用一万次,展开之后指令数量会增加上百万条,指令缓存装不下,运行速度反而可能下降。

inline 真正能确定下来的作用只有一条:允许同一个函数在多个源文件里各有一份定义而不违反单一定义规则。这正好解决了头文件里定义函数的问题,下一节看它的写法与报错。

二十二、inline 的写法

inline 函数要写在头文件里,如果只写声明,把定义留在另一个源文件里,编译器在调用点看不到函数体,既没法展开,也生不出这个符号。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/if_inline.h
#pragma once
#include <iostream>

inline void f(int i);
cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/if_inline.cpp
#include "if_inline.h"

void f(int i)
{
    std::cout << i << std::endl;
}
cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/if_inline_main.cpp
#include "if_inline.h"

int main()
{
    f(10);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ if_inline.cpp if_inline_main.cpp -o if_inline_app
In file included from if_inline_main.cpp:1:0:
if_inline.h:3:13: warning: inline function 'void f(int)' used but never defined [enabled by default]
 inline void f(int i);
             ^
/tmp/ccrkDQWV.o: In function `main':
if_inline_main.cpp:(.text+0xa): undefined reference to `f(int)'
collect2: error: ld returned 1 exit status

警告与报错各一句,头文件里只有声明,编译器按 inline 的约定认为定义会在别处以某种方式合并进来,于是既不生成符号也不报缺定义;到了链接阶段,定义在 if_inline.cpp 里生成的是普通符号,两边对不上,报未定义引用。

把函数的定义直接放进头文件之后,同样的两个源文件就能正常编译并链接。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/ok_inline.h
#pragma once
#include <iostream>

inline void f(int i)
{
    std::cout << i << std::endl;
}
cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/ok_inline_a.cpp
#include "ok_inline.h"

void a() { f(1); }
cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp01/ok_inline_main.cpp
#include "ok_inline.h"

void a();

int main()
{
    f(10);
    a();
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ ok_inline_a.cpp ok_inline_main.cpp -o ok_inline_app && ./ok_inline_app
10
1

两个源文件都包含了同一个头文件,各自都有一份 f 的定义,链接时按弱符号合并成一份,程序正常跑出 10 与 1。这就是上一节说的那条确定作用落在实际代码里的样子。

写的时候有三条要记住,定义放在头文件里,声明与定义都在前面写 inline;不要把一个 inline 函数的定义拆到头文件与 .cpp 两个地方;需要内部链接时用 static inline,这样每个源文件各有一份自己的副本,符号表里看到的是小写 t,与 inline 的弱符号合并是两种不同的做法。

二十三、nullptr 前的 NULL

指针要表示空值,在 C 里靠一个宏,这个宏在不同语言下的定义不一样,把它展开出来看最清楚。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
echo | gcc -E -dM -x c -include stddef.h - | grep -w NULL
#define NULL ((void *)0)
echo | g++ -E -dM -x c++ -include cstddef - | grep -w NULL
#define NULL __null

C 下面它是 ((void *)0),一个指到地址 0 的 void 指针。C++ 下面变成了 __null,一个编译器内置的常量,在 g++ 4.8.5 里它的类型是 long,长度与指针一样是 8 字节。同一个名字,一个是指针,一个是数值常量。

数值常量这四个字带来的第一处后果,是它可以被赋给整型与浮点型。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp11/null_to_int.cpp
#include <iostream>
using namespace std;

int main()
{
    int i = NULL;
    double d = NULL;
    int* p = NULL;
    cout << "i = " << i << ",d = " << d << ",p == 0 时 " << (p == 0) << endl;
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp11
g++ -Wall null_to_int.cpp -o null_to_int
null_to_int.cpp: In function 'int main()':
null_to_int.cpp:5:13: warning: converting to non-pointer type 'int' from NULL [-Wconversion-null]
     int i = NULL;
             ^
null_to_int.cpp:6:16: warning: converting to non-pointer type 'double' from NULL [-Wconversion-null]
     double d = NULL;
                ^
./null_to_int
i = 0,d = 0,p == 0 时 1

编译器只给警告,程序照样跑出 i = 0 与 d = 0。同样的位置换成 nullptr,得到的是错误而不是警告。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -std=c++11 null_int.cpp -o null_int
null_int.cpp: In function 'int main()':
null_int.cpp:3:13: error: cannot convert 'std::nullptr_t' to 'int' in initialization
     int i = nullptr;
             ^

第二处后果落在重载上,下面两个函数只差参数类型。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp11/null_pick1.cpp
#include <iostream>
using namespace std;

void f(int x) { cout << "f(int) " << x << endl; }
void f(double x) { cout << "f(double) " << x << endl; }

int main()
{
    f(NULL);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp11
g++ -Wall null_pick1.cpp -o null_pick1
null_pick1.cpp: In function 'int main()':
null_pick1.cpp:7:11: error: call of overloaded 'f(NULL)' is ambiguous
     f(NULL);
           ^
null_pick1.cpp:7:11: note: candidates are:
null_pick1.cpp:3:6: note: void f(int)

long 转 int 与 long 转 double 都是数值转换,两个候选谁也说不上更合适,编译器拒绝在两者之间选择,把参数换成指针与整型,结论一样。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp11
g++ -Wall null_pick5.cpp -o null_pick5
null_pick5.cpp: In function 'int main()':
null_pick5.cpp:8:11: error: call of overloaded 's(NULL)' is ambiguous
     s(NULL);
           ^
null_pick5.cpp:8:11: note: candidates are:
null_pick5.cpp:3:6: note: void s(int)
null_pick5.cpp:4:6: note: void s(int*)

s 的两个版本都在候选里,一个收整数,一个收指针。想让编译器选中指针那一版,只能自己写出转换。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp11/null_pick3.cpp
#include <iostream>
using namespace std;

void q(int x) { cout << "q(int) " << x << endl; }
void q(int* p) { cout << "q(int*) " << (p == 0) << endl; }

int main()
{
    q(nullptr);
    q((int*)NULL);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp11
g++ -std=c++11 -Wall null_pick3.cpp -o null_pick3 && ./null_pick3
q(int*) 1
q(int*) 1

两种写法都进了指针那一版,区别在 nullptr 不需要写强制转换。这里还有一种写法是把 NULL 定义成 ((void *)0),在 C++ 里同样选不出结果。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ null_void.cpp -o null_void
null_void.cpp: In function 'int main()':
null_void.cpp:5:15: error: call of overloaded 'f(void*)' is ambiguous
     f((void*)0);
           ^
null_void.cpp:5:15: note: candidates are:
null_void.cpp:1:6: note: void f(int) <near match>

void* 到 int 之间没有隐式转换,编译器在 note 里标出 f(int) 只是接近匹配,但 f(int*) 那一版同样要自己做一次指针转换,两边都不成立,结论仍然是二义。

整型字面量 0 是唯一一种能直接进指针重载的写法,它被当作空指针常量,也仍然可以被整型版本接走。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp11/null_pick4.cpp
#include <iostream>
using namespace std;

void r(int x) { cout << "r(int) " << x << endl; }
void r(char* p) { cout << "r(char*) " << (p == 0) << endl; }

int main()
{
    r(0);
    r('\0');
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp11
g++ -Wall null_pick4.cpp -o null_pick4 && ./null_pick4
r(int) 0
r(int) 0

0 是 int 类型,到 r(int) 是精确匹配,到 r(char*) 需要一次指针转换,精确匹配优先,所以选中了整型那一版。写 r(0) 的时候以为会调指针版本,是这类代码里常见的一处误判。

二十四、nullptr 的类型

C++11 加了一个关键字 nullptr,用来表示空指针。它是一个独立的类型 std::nullptr_t,只有一种取值。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/exp02/null_sizeof2.cpp
#include <iostream>
#include <cstddef>
#include <typeinfo>
#include <type_traits>
using namespace std;

void g(int x) { cout << "g(int)" << endl; }
void g(char* p) { cout << "g(char*)" << endl; }

int main()
{
    cout << "sizeof(NULL) = " << sizeof(NULL) << endl;
    cout << "sizeof(nullptr) = " << sizeof(nullptr) << endl;
    cout << "typeid(nullptr).name() = " << typeid(nullptr).name() << endl;
    cout << "is_same<nullptr_t, decltype(nullptr)> = "
         << is_same<nullptr_t, decltype(nullptr)>::value << endl;
    g(nullptr);
    g((char*)0);
    return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp02
g++ -std=c++11 null_sizeof2.cpp -o null_sizeof2 && ./null_sizeof2
sizeof(NULL) = 8
sizeof(nullptr) = 8
typeid(nullptr).name() = Dn
is_same<nullptr_t, decltype(nullptr)> = 1
g(char*)
g(char*)

typeid 打出的是编译器内部的名字,Dn 是 g++ 给 std::nullptr_t 的修饰名;下一行的 1 说明它的类型确实是 nullptr_t。两次调用都进了 g(char*),一次传 nullptr,一次传强制转换成 char* 的 0,两条路指向同一件事。

这个类型与整型分得很开,同一个 f(int)/f(double) 的重载,参数换成 nullptr 之后不再是二义,而是没有匹配。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp11
g++ -std=c++11 -Wall null_pick2.cpp -o null_pick2
null_pick2.cpp: In function 'int main()':
null_pick2.cpp:7:14: error: no matching function for call to 'f(std::nullptr_t)'
     f(nullptr);
              ^
null_pick2.cpp:7:14: note: candidates are:
null_pick2.cpp:3:6: note: void f(int)
 void f(int x) { cout << "f(int) " << x << endl; }
      ^

报错信息里直接写出了实参的类型 std::nullptr_t。整型与浮点型的候选都列在下面,nullptr 到它们之间没有可用的转换,与 NULL 那种两边都能凑上去的情形正好相反。

用之前还要确认标准,nullptr 是 C++11 起才有的关键字,把编译命令退回 C++98,它连名字都不认识。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -std=c++98 null_c98.cpp -o null_c98
null_c98.cpp: In function 'int main()':
null_c98.cpp:4:14: error: 'nullptr' was not declared in this scope
     int* p = nullptr;
              ^

至此这一节的两条结论可以直接记下来:需要空指针的地方写 nullptr,需要整数 0 的地方写 0,两者不要混用。写 nullptr 遇到 was not declared in this scope,先检查编译命令里的标准选项有没有到 C++11。

图 12 NULL 与 nullptr 在重载选择上的两条路

二十五、本节的三个报错

这一节把前面几处编译期报错与一条运行期警告收在一起,都来自引用与空指针这几节里的实测。

第一处报错来自引用声明时没有初始化的写法,编译器在编译期就把它拦住了。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ ref_uninit.cpp -o ref_uninit
ref_uninit.cpp: In function 'int main()':
ref_uninit.cpp:3:10: error: 'ra' declared as reference but not initialized
     int& ra;
          ^

引用是一块已有空间的别名,声明的时候必须说明它绑定谁。写成 int& ra; 之后再赋值,编译器不会当成延迟初始化,直接报未初始化。改成 int& ra = a; 即可。

第二处是函数把局部变量的引用当返回值交了出去,编译器在这一处只给出一条警告。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ -Wall ret_dangling.cpp -o ret_dangling
ret_dangling.cpp: In function 'int& func()':
ret_dangling.cpp:5:9: warning: reference to local variable 'a' returned [-Wreturn-local-addr]
     int a = 0;
         ^

这条只给警告,程序能编出来,函数里的局部变量在函数返回时随栈帧一起消失,调用方拿到的是一个指向已失效地址的引用,读到的值取决于那块栈空间后来被谁覆盖。改动办法是让被引用的对象活得比函数久,例如返回全局变量、静态变量,或者调用方传进来的对象。

第三处报错出现在 NULL 进重载的时候,编译器在几个候选之间无法确定该调用哪一个。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/exp01
g++ null_overload.cpp -o null_overload
null_overload.cpp: In function 'int main()':
null_overload.cpp:8:11: error: call of overloaded 'f(NULL)' is ambiguous
     f(NULL);
           ^
null_overload.cpp:8:11: note: candidates are:
null_overload.cpp:3:6: note: void f(int)

NULL 是一个数值常量,在 int 与指针两版重载之间没有优先关系,编译器列出候选就停住了。把调用处的 NULL 换成 nullptr,调用直接落到指针那一版;如果非要保留 NULL,就得自己写出 (int*)NULL 这样的转换。

三处报错指向的是同一件事的三种表现:引用与指针既有类型信息又要在编译期确定绑定关系,越界的写法都会在这一步被拦住,只是拦住的力度不同,未初始化与二义是错误,返回局部变量的引用只到警告。

实例代码

工程结构

text 复制代码
demo/
├── bit_stack.h     栈的定义与实现,全部放在命名空间 bit 里
├── bit_utils.h     三个 Add 重载与两个 Swap 重载
├── c_legacy.h      C 侧函数的声明,带条件编译
├── c_legacy.c      C 侧函数的实现,用 gcc 编译
├── main.cpp        入口,走一遍命名空间、重载、引用、缺省参数与 nullptr
├── edge.cpp        五组边界用例
├── err1.cpp        把 nullptr 赋给 int
├── err2.cpp        Add(1, 2.5) 的二义
└── err3.cpp        Add(1) 参数个数不够

编译命令是 gcc -c c_legacy.c -o c_legacy.o 加 g++ -std=c++11 -Wall main.cpp c_legacy.o -o demo,零告警。

完整代码

栈的定义与实现合在一个头文件里,所有函数都写了 inline,函数名都在 bit 这个命名空间内。接口参数从 C 版的 ST* ps 换成了 Stack& rs,调用处不用再取地址。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/demo/bit_stack.h
#pragma once
#include <cstdio>
#include <cstdlib>

namespace bit
{
	typedef int STDataType;

	struct Stack
	{
		STDataType* a;
		int top;
		int capacity;
	};

	inline void STInit(Stack& rs, int n = 4)
	{
		rs.a = (STDataType*)malloc(sizeof(STDataType) * n);
		if (rs.a == nullptr)
		{
			perror("malloc fail");
			exit(-1);
		}
		rs.top = 0;
		rs.capacity = n;
	}

	inline void STPush(Stack& rs, STDataType x)
	{
		if (rs.top == rs.capacity)
		{
			int newcapacity = rs.capacity == 0 ? 4 : rs.capacity * 2;
			STDataType* tmp = (STDataType*)realloc(rs.a, sizeof(STDataType) * newcapacity);
			if (tmp == nullptr)
			{
				perror("realloc fail");
				exit(-1);
			}
			rs.a = tmp;
			rs.capacity = newcapacity;
		}
		rs.a[rs.top] = x;
		rs.top++;
	}

	inline STDataType& STTop(Stack& rs)
	{
		return rs.a[rs.top - 1];
	}

	inline int STSize(const Stack& rs)
	{
		return rs.top;
	}

	inline void STDestroy(Stack& rs)
	{
		free(rs.a);
		rs.a = nullptr;
		rs.top = 0;
		rs.capacity = 0;
	}
}

三个 Add 靠参数类型与个数区分,两个 Swap 一个收引用一个收指针,调用处传变量名与传地址各走一版。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/demo/bit_utils.h
#pragma once

namespace bit
{
	inline int Add(int x, int y)
	{
		return x + y;
	}

	inline double Add(double x, double y)
	{
		return x + y;
	}

	inline int Add(int x, int y, int z)
	{
		return x + y + z;
	}

	inline void Swap(int& rx, int& ry)
	{
		int tmp = rx;
		rx = ry;
		ry = tmp;
	}

	inline void Swap(int* px, int* py)
	{
		int tmp = *px;
		*px = *py;
		*py = tmp;
	}
}

C 侧的函数按 C 的规则写,声明外面套一层 extern "C",用条件编译保证同一个头文件在 C 与 C++ 下都能包含。

c 复制代码
/* 文件:/home/cocatrice/cpp01/demo/c_legacy.h */
#ifndef C_LEGACY_H
#define C_LEGACY_H

#ifdef __cplusplus
extern "C" {
#endif

int c_sum(int* a, int n);

#ifdef __cplusplus
}
#endif

#endif
c 复制代码
/* 文件:/home/cocatrice/cpp01/demo/c_legacy.c */
#include "c_legacy.h"

int c_sum(int* a, int n)
{
	int s = 0;
	int i = 0;
	for (i = 0; i < n; i++)
	{
		s += a[i];
	}
	return s;
}

入口里把这一篇的语法各用一次,读输入用的是 cin,空指针比较用的是 nullptr。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/demo/main.cpp
#include <iostream>
#include "bit_stack.h"
#include "bit_utils.h"
#include "c_legacy.h"

using std::cin;
using std::cout;
using std::endl;

int main()
{
	bit::Stack st;
	bit::STInit(st);

	int i = 0;
	for (i = 1; i <= 6; i++)
	{
		bit::STPush(st, i * 10);
	}

	cout << "元素个数 " << bit::STSize(st) << ",容量 " << st.capacity << endl;
	bit::STTop(st) += 5;
	cout << "栈顶 " << bit::STTop(st) << endl;

	int a = 1, b = 2;
	bit::Swap(a, b);
	bit::Swap(&a, &b);
	cout << "交换两轮之后 a = " << a << ",b = " << b << endl;

	cout << "Add(1, 2) = " << bit::Add(1, 2) << endl;
	cout << "Add(1.5, 2.5) = " << bit::Add(1.5, 2.5) << endl;
	cout << "Add(1, 2, 3) = " << bit::Add(1, 2, 3) << endl;

	int arr[5] = { 1, 2, 3, 4, 5 };
	cout << "C 侧求和 " << c_sum(arr, 5) << endl;

	int n = 0;
	cout << "请输入一个整数:";
	cin >> n;
	cout << "输入的是 " << n << endl;

	bit::STDestroy(st);
	return 0;
}

正常情况

text 复制代码
# 所属目录:/home/cocatrice/cpp01/demo
gcc -c c_legacy.c -o c_legacy.o
g++ -std=c++11 -Wall main.cpp c_legacy.o -o demo && echo '编译零告警'
编译零告警
echo '4 5 6' | ./demo
元素个数 6,容量 8
栈顶 65
交换两轮之后 a = 1,b = 2
Add(1, 2) = 3
Add(1.5, 2.5) = 4
Add(1, 2, 3) = 6
C 侧求和 15
请输入一个整数:输入的是 4
退码=0

六次 push 之后容量从默认的 4 扩到 8,栈顶那个 60 被 STTop(st) += 5 改成 65,说明返回引用的时候调用方拿到的是栈里那块空间本身。交换两轮之后 a 与 b 回到 1 与 2,两轮各自走了引用版与指针版。cin 读到 4 就停了,后面两个数留在输入缓冲区里,程序没有再一次读,所以直接结束。

常见报错

报错一:把 nullptr 赋给整型。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/demo/err1.cpp
#include <iostream>
using namespace std;

int main()
{
	int i = nullptr;
	cout << i << endl;
	return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/demo
g++ -std=c++11 -Wall err1.cpp -o err1
err1.cpp: In function 'int main()':
err1.cpp:6:10: error: cannot convert 'std::nullptr_t' to 'int' in initialization
  int i = nullptr;
          ^

报错二:一个 int 实参配一个 double 实参。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/demo/err2.cpp
#include <iostream>
#include "bit_utils.h"
using namespace std;

int main()
{
	cout << bit::Add(1, 2.5) << endl;
	return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/demo
g++ -std=c++11 -Wall err2.cpp -o err2
err2.cpp: In function 'int main()':
err2.cpp:7:25: error: call of overloaded 'Add(int, double)' is ambiguous
  cout << bit::Add(1, 2.5) << endl;
                         ^
err2.cpp:7:25: note: candidates are:
In file included from err2.cpp:2:0:
bit_utils.h:5:13: note: int bit::Add(int, int)
  inline int Add(int x, int y)
             ^
bit_utils.h:10:16: note: double bit::Add(double, double)
  inline double Add(double x, double y)
                ^

两个候选各有一处需要转换,一处是第二个参数 int 转 double,一处是第二个参数 double 转 int,成本相同,编译器不选。把实参写成 Add(1, 2) 或者 Add(1.0, 2.5) 都能编过。

报错三:实参个数不够。

cpp 复制代码
// 文件:/home/cocatrice/cpp01/demo/err3.cpp
#include <iostream>
#include "bit_utils.h"
using namespace std;

int main()
{
	cout << bit::Add(1) << endl;
	return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/demo
g++ -std=c++11 -Wall err3.cpp -o err3
err3.cpp: In function 'int main()':
err3.cpp:7:20: error: no matching function for call to 'Add(int)'
  cout << bit::Add(1) << endl;
                    ^
err3.cpp:7:20: note: candidates are:
In file included from err3.cpp:2:0:
bit_utils.h:5:13: note: int bit::Add(int, int)
  inline int Add(int x, int y)
             ^
bit_utils.h:5:13: note:   candidate expects 2 arguments, 1 provided

这里没有缺省参数可以兜底,Add 的三个版本都要求两个或三个实参,编译器把候选挨个列出来,并在每个候选后面标出期望的实参个数。

边界情况

cpp 复制代码
// 文件:/home/cocatrice/cpp01/demo/edge.cpp
#include <iostream>
#include "bit_stack.h"
#include "bit_utils.h"
#include "c_legacy.h"

using std::cout;
using std::endl;

int main()
{
	bit::Stack s0;
	bit::STInit(s0, 0);
	cout << "容量给 0:容量 " << s0.capacity << ",个数 " << bit::STSize(s0) << endl;
	bit::STPush(s0, 7);
	cout << "第一次 push 之后容量 " << s0.capacity << ",栈顶 " << bit::STTop(s0) << endl;

	bit::Stack s1;
	bit::STInit(s1, 1);
	bit::STPush(s1, 1);
	bit::STPush(s1, 2);
	cout << "只给一个槽位再 push 两次:容量 " << s1.capacity << ",个数 " << bit::STSize(s1) << endl;

	bit::STTop(s1) = 100;
	cout << "给栈顶赋值之后 " << bit::STTop(s1) << endl;

	cout << "c_sum(nullptr, 0) = " << c_sum(nullptr, 0) << endl;

	int x = 5;
	bit::Swap(x, x);
	cout << "自己和自己交换之后 " << x << endl;

	bit::STDestroy(s0);
	bit::STDestroy(s1);
	return 0;
}
text 复制代码
# 所属目录:/home/cocatrice/cpp01/demo
g++ -std=c++11 -Wall edge.cpp c_legacy.o -o edge && ./edge
容量给 0:容量 0,个数 0
第一次 push 之后容量 4,栈顶 7
只给一个槽位再 push 两次:容量 2,个数 2
给栈顶赋值之后 100
c_sum(nullptr, 0) = 0
自己和自己交换之后 5

容量给 0 的那一组走的是扩容函数里的三元表达式,第一次 push 时把容量改成 4。只给一个槽位的栈第二次 push 时扩到 2。c_sum(nullptr, 0) 返回 0,因为循环次数是 0,指针从来没有被解引用,这是空指针在参数里作为「没有数据」标记的正常用法。Swap 传同一个变量两次不会出错,中间那次赋值写回的是同一个地址。

实测数据

符号表能看出两件事,C 侧目标文件里的函数名是原样的 c_sum,C++ 侧目标文件里引用它写的是 U c_sum,两边不需要名字修饰就能对上。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/demo
nm c_legacy.o
0000000000000000 T c_sum
nm main.o | grep -E '_ZN3bit|_Z4Swap|_Z3Add|c_sum'
                 U c_sum
0000000000000000 W _ZN3bit3AddEdd
0000000000000000 W _ZN3bit3AddEii
0000000000000000 W _ZN3bit3AddEiii
0000000000000000 W _ZN3bit4SwapEPiS0_
0000000000000000 W _ZN3bit4SwapERiS0_
0000000000000000 W _ZN3bit5STTopERNS_5StackE
0000000000000000 W _ZN3bit6STInitERNS_5StackEi
0000000000000000 W _ZN3bit6STPushERNS_5StackEi
0000000000000000 W _ZN3bit6STSizeERKNS_5StackE
0000000000000000 W _ZN3bit9STDestroyERNS_5StackE
nm -C main.o | grep -E 'bit::|Swap|c_sum'
                 U c_sum
0000000000000000 W bit::Add(double, double)
0000000000000000 W bit::Add(int, int)
0000000000000000 W bit::Add(int, int, int)
0000000000000000 W bit::Swap(int*, int*)
0000000000000000 W bit::Swap(int&, int&)
0000000000000000 W bit::STTop(bit::Stack&)
0000000000000000 W bit::STInit(bit::Stack&, int)
0000000000000000 W bit::STPush(bit::Stack&, int)
0000000000000000 W bit::STSize(bit::Stack const&)
0000000000000000 W bit::STDestroy(bit::Stack&)

自己写的函数全是 W,表示弱符号,它们在头文件里定义、被多个源文件包含时由链接器合并成一份。名字里能读出参数类型,_ZN3bit4SwapERiS0_ 末尾的 Ri 是引用,_ZN3bit4SwapEPiS0_ 里的 Pi 是指针,同一个 Swap 的两版在符号表里也是两个名字。

优化档对内联的影响可以直接数调用指令与查看段大小。

text 复制代码
# 所属目录:/home/cocatrice/cpp01/demo
g++ -std=c++11 -S -O0 main.cpp -o main_O0.s
g++ -std=c++11 -S -O2 main.cpp -o main_O2.s
grep -c call main_O0.s
52
grep -c call main_O2.s
39
grep -E 'call.*_ZN3bit' main_O0.s
	call	_ZN3bit6STInitERNS_5StackEi
	call	_ZN3bit6STPushERNS_5StackEi
	call	_ZN3bit6STSizeERKNS_5StackE
	call	_ZN3bit5STTopERNS_5StackE
	call	_ZN3bit5STTopERNS_5StackE
	call	_ZN3bit4SwapERiS0_
	call	_ZN3bit4SwapEPiS0_
	call	_ZN3bit3AddEii
	call	_ZN3bit3AddEdd
	call	_ZN3bit3AddEiii
	call	_ZN3bit9STDestroyERNS_5StackE
size demo_O0 demo_O2
   text	   data	    bss	    dec	    hex	filename
   4749	    708	    568	   6025	   1789	demo_O0
   3542	    700	    568	   4810	   12ca	demo_O2

-O0 下 11 个自己写的函数都留了调用指令,-O2 之后 call 总数从 52 条降到 39 条,text 段从 4749 字节降到 3542 字节。两个优化档跑出来的输出完全一样,元素个数、容量、栈顶三行逐字相同。两个优化档的 call 条数与 text 段字节数并排如下。

优化档 call 条数 text 段字节数 运行输出
-O0 52 4749 与 -O2 逐字相同
-O2 39 3542 与 -O0 逐字相同

实例:把 C 写的模块接到 C++ 上

上面这个工程里 c_legacy.c 就是一份按 C 规则写的代码,它不知道什么是命名空间,也不会做名字修饰。接入的方式是给它配一个头文件,声明外面套条件编译的 extern "C",C 源文件用 gcc 编译,C++ 源文件用 g++ 编译,最后一起交给链接器。链接的时候 C++ 那边发出的符号名是 c_sum,C 那边定义的名字也是 c_sum,两边直接对上。

命名空间在这件事上提供的帮助是把新写的代码隔开。bit 里的 Add、Swap、STPush 即使与某个库里的名字重名,也不会互相干扰,因为符号表里的全名带着 _ZN3bit 前缀。真正需要跨语言的那几个函数,就在头文件里单独用 extern "C" 放开名字修饰。

引用、缺省参数与函数重载把接口的调用形式变多了:同一个 STInit 既能带容量调用也能省略容量,同一组 Swap 既能传值也能传地址,栈顶接口返回引用之后可以直接当左值用。这些语法在后面的类与对象里还会接着用,那时栈会被封装成一个类,这些自由函数会变成成员函数,接口的名字不变。

踩坑点

  • int rand = 10; 在包含 之后报重复定义,因为库里的 rand 是全局函数,改放进命名空间或者换名字。
  • 头文件里只想用某个名字时写 using std::cout;,写 using namespace std; 会把整个标准库的名字都拉进包含它的每个源文件。
  • using bit::a; 之后再写 int a = 5; 报重复定义,同一作用域里同名的 using 声明与变量不能共存。
  • 缺省参数只写在声明里,定义里再写一次会报 default argument given for parameter。
  • 半缺省必须从右往左连续给,void Func2(int a, int b = 1, int c); 这种写法报 default argument missing for parameter 2。
  • 只有返回值不同的两个同名函数不构成重载,报 new declaration 与 ambiguates old declaration。
  • f(int) 与 f(int, int = 0) 同时存在时,f(1) 报调用二义。
  • 名字修饰只在 C++ 里做,C 编译出的目标文件里函数名是原样的,跨语言调用必须写 extern "C"。
  • extern "C" 里不能放两个同名的重载函数,报 declaration of C function conflicts with previous declaration。
  • 引用声明时必须初始化,int& ra; 报 declared as reference but not initialized。
  • 引用一旦绑定不能改绑,b = e; 改的是被引用对象的值,不是让 b 指向 e。
  • 返回局部变量的引用只给 -Wreturn-local-addr 警告,运行结果不确定。
  • 空指针解引用取引用这一句本身不崩,int& r = *p; 在处理到 r 之前不会访问内存。
  • const int& 能绑右值与类型不同的值,代价是编译器开一块临时空间,改临时量的地址不是原变量的地址。
  • 引用占用的空间与指针一样,sizeof(int&) 是 4 是因为 sizeof 取的是被引用对象的类型,不是引用本身的大小。
  • 引用不能组成数组,int& arr[3]; 报 declaration of arr as array of references。
  • 宏的参数不带类型,MUL_NO_PAREN(1+2, 3+4) 展开成 1 + 2 * 3 + 4 得到 11,写宏要给参数与整体各加一层括号。
  • 宏的参数出现两次就求值两次,ADD(x++, x++) 之后 x 从 1 变成 3。
  • 宏体是几条语句时放在 if 后面会漏出作用域,加 do while 包起来也只是补上一半。
  • inline 只是建议,-O0 下照样生成 call 指令。
  • 没写 inline 的小函数在 -O2 下也会被展开,是否展开由优化档与成本模型决定。
  • inline 函数的定义要放在头文件里,只放声明会报 inline function used but never defined 与未定义引用。
  • 返回局部变量的引用在 -O2 下有可能把乱码优化成一个固定值,掩盖掉问题,别靠优化档掩盖告警。
  • NULL 在 C++ 里是数值常量,赋给 int 只给 -Wconversion-null 警告,换成 nullptr 直接报错。
  • f(NULL) 在 int 与指针两版重载之间二义,写 f((void*)0) 同样二义。
  • nullptr 要 C++11 起才能用,-std=c++98 下报 was not declared in this scope。

本篇总结(模拟面试问题)

问:命名空间解决了什么问题,为什么 C 里做不到同样的事?

核心要点:C 语言在不同源文件里出现同名函数时,链接器会报符号重复,只能靠改名字或者加前缀躲开。命名空间把名字放进一个作用域里,符号表里带上域的全名,两个库里都叫 Add 的函数可以同时存在。C 语言没有作用域这套机制,函数名在编译之后就只剩一个扁平的名字。

问:using 声明与 using 指令的区别是什么?

核心要点:using 声明只把某一个名字引入当前作用域,using std::cout; 之后就只认识 cout。using 指令把整个命名空间展开,using namespace std; 之后标准库里的所有名字都可见,代价是可能与自定义的名字冲突,并且一旦写在头文件里,所有包含它的源文件都会受影响。

问:缺省参数为什么只能从右往左连续给?

核心要点:实参按位置对应形参,从左往右填。如果中间某个参数有默认值而它右边的参数没有,那么少传实参时编译器无法判断缺的是哪一个,因此默认值必须连续地贴在参数表最右边。

问:缺省参数为什么只写在声明里?

核心要点:默认值是在调用点填入的,编译器需要在看到调用处时就知道有哪些默认值,所以默认值要出现在声明里并且被调用方包含。声明与定义各写一次会报重定义,只有定义里有默认值而声明里没有,调用方看不到,仍然报实参不足。

问:函数重载的判定依据是什么,为什么不看返回值?

核心要点:同名函数只要参数的类型、个数、顺序有一处不同就构成重载,编译器按实参与形参的匹配程度选最合适的一版。返回值不参与重载判定,因为调用处可以只写函数名而不用返回值,写 f(); 的时候编译器无法从上下文判断该选哪一版。

问:重载在底层是怎么实现的?

核心要点:编译器把函数名与参数类型一起编码成新的符号名,g++ 把 f(int, double) 编成 _Z1fid,把命名空间也编进去变成 _ZN3bit3AddEii。链接器看到的是这些新名字,因此同名函数不会撞车。C 语言不做这件事,跨语言调用要用 extern "C" 关掉名字修饰。

问:extern "C" 的作用是什么?

核心要点:它告诉 C++ 编译器,这一段声明与定义按 C 的规则生成符号名,也就是保持原名不做修饰。这样 C 编译出的目标文件能被 C++ 调用,C++ 里包装过的函数也能被 C 调用。它按声明生效而不是按文件生效,同一个函数在 C++ 里可以只把对外的那几个放开。

问:引用与指针的区别是什么?

核心要点:引用是给一块已经存在的空间取别名,声明时必须初始化,绑定之后不能改绑,使用时不需要取地址与解引用。指针本身是一个变量,可以改指向、可以为空指针、可以做算术运算。语法上是两套写法,机器码里引用与指针生成的指令可以逐行相同。

问:引用做返回值要注意什么?

核心要点:返回的引用必须指向一个活得比函数久的对象,例如全局变量、静态变量或调用方传进来的对象。返回局部变量的引用,函数结束后那块栈空间已经失效,读到的值取决于栈被谁覆盖,编译器只给一条 -Wreturn-local-addr 警告。

问:const 引用为什么能绑定临时量?

核心要点:非 const 引用要求绑定一块可修改的现有空间,临时量没有名字也不可修改,所以绑不上。const 引用允许编译器为临时量开辟一块空间并把引用绑上去,同时把这块空间的生命周期延长到引用的作用域结束,这也是 func(const string&) 能直接接收字符串字面量的原因。

问:inline 的作用是什么?

核心要点:它首先是一个建议,编译器可以展开也可以照常生成调用指令,实际是否展开由优化档与成本模型决定。它能确定下来的作用是允许同一份函数定义出现在多个源文件里而不违反单一定义规则,因此可以把函数定义直接放进头文件。

问:宏和内联函数差别在哪?

核心要点:宏是预处理阶段的文本替换,参数没有类型、可能被求值多次、可能因为缺少括号改变运算顺序,也没有作用域与调试信息。inline 函数有参数表与返回类型,实参只求值一次,按普通的名字查找规则参与重载,调试时也按函数名出现。

问:NULL 和 nullptr 有什么不一样?

核心要点:NULL 是一个宏,在 C 里展开成 ((void *)0),在 g++ 里展开成内置常量 __null,它的类型是数值类型,能赋给 int 与 double,只在开了 -Wconversion-null 时给警告。nullptr 是 C++11 的关键字,类型是 std::nullptr_t,能隐式转成任意指针类型,但不能转成整型。

问:为什么 f(NULL) 会报二义而 f(nullptr) 不会?

核心要点:NULL 是数值常量,在 int 与指针两版重载之间,两边都需要一次转换,没有哪一边更合适,编译器停住不选。nullptr 有自己的类型,见到 f(int) 之类的整型候选时根本没有可用的转换,直接报没有匹配的函数,见到指针候选时走标准转换,选择是明确的。

参考

相关推荐
91刘仁德1 小时前
IP协议详解:从IP协议头到网段划分、路由与NAT
linux·服务器·网络·网络协议·tcp/ip
xxwl5851 小时前
数据结构知识点和代码实现总结(C语言实现)
c语言·开发语言·数据结构
硅基手札2 小时前
【Linux内核专栏 14】网络协议栈
linux·运维·网络协议
汉克老师2 小时前
GESP2026年9月认证C++八级( 第一部分选择题(1~7题)精讲
c++·gesp·小学生·学c++编程
郝学胜_神的一滴4 小时前
C++ Templates 06:搞懂模板代码的三种组织方式
c++·visual studio
C++ 老炮儿的技术栈4 小时前
sizeof操作符
c语言·c++·人工智能·mfc·c
沫璃染墨4 小时前
《Linux工程实践篇(一):认识设计模式——从日志系统看策略模式》
linux·c++·安全·设计模式·策略模式
工作10年+,存储芯片行业4 小时前
Linux NVMe 中断排查与性能优化:CPU 亲和性
linux·运维·服务器·windows·性能优化·ssd·pcie
wuminyu4 小时前
ForkJoinPool内部WorkQueue的Lock-Free数组操作以及并发任务窃取原理剖析
java·linux·c语言·jvm·c++