cJSON 源码解析:parse_string 函数深入剖析

1. 引言

cJSON 是一个轻量级的 C 语言 JSON 解析库,因其代码简洁、无外部依赖而广受欢迎。在 cJSON 的解析流程中,parse_string 函数负责将 JSON 字符串字面量(如 "hello"、"a\\nb")解析为 C 字符串,是整个解析器中最核心、也最容易出错的环节之一。

本文将逐行剖析 cJSON 源码中的 parse_string 实现,讲解其转义字符处理、Unicode 解码、内存分配等关键细节,帮助你彻底理解这个函数的工作原理。

2. parse_string 函数概览

parse_string 函数定义在 cJSON.c 文件中,其函数签名如下:

c 复制代码
static cJSON_bool parse_string(cJSON *item, const unsigned char * const input, const unsigned char ** ep);

三个参数的含义:

  • item:指向待填充的 cJSON 节点,解析成功后字符串值会存入 item->valuestring。
  • input:指向 JSON 文本中字符串的起始位置(即开头的双引号 ")。
  • ep:指向指针的指针,解析失败时用于记录错误位置。

函数返回 cJSON_bool(即 int),成功返回 true,失败返回 false。

3. 整体流程

parse_string 的整体流程可以概括为以下几个步骤:
#mermaid-svg-tznls9BOulICLZga{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;fill:#333;}@keyframes edge-animation-frame{from{stroke-dashoffset:0;}}@keyframes dash{to{stroke-dashoffset:0;}}#mermaid-svg-tznls9BOulICLZga .edge-animation-slow{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 50s linear infinite;stroke-linecap:round;}#mermaid-svg-tznls9BOulICLZga .edge-animation-fast{stroke-dasharray:9,5!important;stroke-dashoffset:900;animation:dash 20s linear infinite;stroke-linecap:round;}#mermaid-svg-tznls9BOulICLZga .error-icon{fill:#552222;}#mermaid-svg-tznls9BOulICLZga .error-text{fill:#552222;stroke:#552222;}#mermaid-svg-tznls9BOulICLZga .edge-thickness-normal{stroke-width:1px;}#mermaid-svg-tznls9BOulICLZga .edge-thickness-thick{stroke-width:3.5px;}#mermaid-svg-tznls9BOulICLZga .edge-pattern-solid{stroke-dasharray:0;}#mermaid-svg-tznls9BOulICLZga .edge-thickness-invisible{stroke-width:0;fill:none;}#mermaid-svg-tznls9BOulICLZga .edge-pattern-dashed{stroke-dasharray:3;}#mermaid-svg-tznls9BOulICLZga .edge-pattern-dotted{stroke-dasharray:2;}#mermaid-svg-tznls9BOulICLZga .marker{fill:#333333;stroke:#333333;}#mermaid-svg-tznls9BOulICLZga .marker.cross{stroke:#333333;}#mermaid-svg-tznls9BOulICLZga svg{font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:16px;}#mermaid-svg-tznls9BOulICLZga p{margin:0;}#mermaid-svg-tznls9BOulICLZga .label{font-family:"trebuchet ms",verdana,arial,sans-serif;color:#333;}#mermaid-svg-tznls9BOulICLZga .cluster-label text{fill:#333;}#mermaid-svg-tznls9BOulICLZga .cluster-label span{color:#333;}#mermaid-svg-tznls9BOulICLZga .cluster-label span p{background-color:transparent;}#mermaid-svg-tznls9BOulICLZga .label text,#mermaid-svg-tznls9BOulICLZga span{fill:#333;color:#333;}#mermaid-svg-tznls9BOulICLZga .node rect,#mermaid-svg-tznls9BOulICLZga .node circle,#mermaid-svg-tznls9BOulICLZga .node ellipse,#mermaid-svg-tznls9BOulICLZga .node polygon,#mermaid-svg-tznls9BOulICLZga .node path{fill:#ECECFF;stroke:#9370DB;stroke-width:1px;}#mermaid-svg-tznls9BOulICLZga .rough-node .label text,#mermaid-svg-tznls9BOulICLZga .node .label text,#mermaid-svg-tznls9BOulICLZga .image-shape .label,#mermaid-svg-tznls9BOulICLZga .icon-shape .label{text-anchor:middle;}#mermaid-svg-tznls9BOulICLZga .node .katex path{fill:#000;stroke:#000;stroke-width:1px;}#mermaid-svg-tznls9BOulICLZga .rough-node .label,#mermaid-svg-tznls9BOulICLZga .node .label,#mermaid-svg-tznls9BOulICLZga .image-shape .label,#mermaid-svg-tznls9BOulICLZga .icon-shape .label{text-align:center;}#mermaid-svg-tznls9BOulICLZga .node.clickable{cursor:pointer;}#mermaid-svg-tznls9BOulICLZga .root .anchor path{fill:#333333!important;stroke-width:0;stroke:#333333;}#mermaid-svg-tznls9BOulICLZga .arrowheadPath{fill:#333333;}#mermaid-svg-tznls9BOulICLZga .edgePath .path{stroke:#333333;stroke-width:2.0px;}#mermaid-svg-tznls9BOulICLZga .flowchart-link{stroke:#333333;fill:none;}#mermaid-svg-tznls9BOulICLZga .edgeLabel{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-tznls9BOulICLZga .edgeLabel p{background-color:rgba(232,232,232, 0.8);}#mermaid-svg-tznls9BOulICLZga .edgeLabel rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-tznls9BOulICLZga .labelBkg{background-color:rgba(232, 232, 232, 0.5);}#mermaid-svg-tznls9BOulICLZga .cluster rect{fill:#ffffde;stroke:#aaaa33;stroke-width:1px;}#mermaid-svg-tznls9BOulICLZga .cluster text{fill:#333;}#mermaid-svg-tznls9BOulICLZga .cluster span{color:#333;}#mermaid-svg-tznls9BOulICLZga div.mermaidTooltip{position:absolute;text-align:center;max-width:200px;padding:2px;font-family:"trebuchet ms",verdana,arial,sans-serif;font-size:12px;background:hsl(80, 100%, 96.2745098039%);border:1px solid #aaaa33;border-radius:2px;pointer-events:none;z-index:100;}#mermaid-svg-tznls9BOulICLZga .flowchartTitleText{text-anchor:middle;font-size:18px;fill:#333;}#mermaid-svg-tznls9BOulICLZga rect.text{fill:none;stroke-width:0;}#mermaid-svg-tznls9BOulICLZga .icon-shape,#mermaid-svg-tznls9BOulICLZga .image-shape{background-color:rgba(232,232,232, 0.8);text-align:center;}#mermaid-svg-tznls9BOulICLZga .icon-shape p,#mermaid-svg-tznls9BOulICLZga .image-shape p{background-color:rgba(232,232,232, 0.8);padding:2px;}#mermaid-svg-tznls9BOulICLZga .icon-shape .label rect,#mermaid-svg-tznls9BOulICLZga .image-shape .label rect{opacity:0.5;background-color:rgba(232,232,232, 0.8);fill:rgba(232,232,232, 0.8);}#mermaid-svg-tznls9BOulICLZga .label-icon{display:inline-block;height:1em;overflow:visible;vertical-align:-0.125em;}#mermaid-svg-tznls9BOulICLZga .node .label-icon path{fill:currentColor;stroke:revert;stroke-width:revert;}#mermaid-svg-tznls9BOulICLZga :root{--mermaid-font-family:"trebuchet ms",verdana,arial,sans-serif;} 否
是
检查开头引号
是否含转义字符?
直接计算长度并复制
逐字符解析转义
处理 Unicode 转义
分配内存并填充
写入 valuestring
更新解析位置 ep

下面我们按照源码顺序,逐步拆解每个环节。

4. 源码逐段解析

4.1 前置检查:确认起始引号

c 复制代码
if (input == NULL || *input != '\"') {
    return false;
}

函数首先检查 input 是否为空,以及当前位置是否为双引号 "。如果不是,直接返回 false,表示解析失败。这是字符串解析的入口校验。

4.2 第一遍扫描:计算字符串长度

c 复制代码
const unsigned char *pointer = input + 1;
size_t length = 0;

while (*pointer != '\"' && *pointer != '\0') {
    if (*pointer == '\\') {
        pointer++;
        if (*pointer == '\0') {
            return false;
        }
        pointer++;
    } else {
        pointer++;
    }
    length++;
}

这一遍扫描的目的是计算字符串的实际长度(不含转义符本身)。注意:

  • 遇到 \ 时,跳过转义符及其后面的一个字符(如 \n、\" 都算一个字符)。
  • 遇到 " 或字符串结束符 \0 时停止。
  • 如果遇到 \ 后紧跟 \0,说明转义不完整,返回 false。

这里有个细节:length 统计的是转义后 的字符数,而不是原始字节数。例如 "a\\nb" 的原始字节数是 5(a、\、n、b、"),但转义后实际只有 3 个字符(a、换行、b)。

4.3 判断是否需要二次扫描

c 复制代码
if (*pointer != '\"') {
    return false;
}

第一遍扫描结束后,pointer 应指向结束引号 "。如果不是,说明字符串没有正常闭合,返回 false。

c 复制代码
if (pointer == input + 1) {
    /* 空字符串 */
    item->valuestring = (char*)cJSON_malloc(1);
    if (item->valuestring == NULL) {
        return false;
    }
    item->valuestring[0] = '\0';
    *ep = pointer + 1;
    return true;
}

如果 pointer == input + 1,说明字符串是空的(""),直接分配 1 字节内存存放 \0 即可。

c 复制代码
if (length >= (size_t)(pointer - input) - 1) {
    /* 没有转义字符,直接复制 */
    item->valuestring = (char*)cJSON_malloc(length + 1);
    if (item->valuestring == NULL) {
        return false;
    }
    memcpy(item->valuestring, input + 1, length);
    item->valuestring[length] = '\0';
    *ep = pointer + 1;
    return true;
}

这里是一个性能优化 :如果 length 不小于原始字节数减一(即没有转义字符被压缩),说明字符串中没有转义序列,可以直接用 memcpy 批量复制,无需逐字符处理。

4.4 第二遍扫描:处理转义字符

c 复制代码
/* 有转义字符,需要逐字符处理 */
pointer = input + 1;
unsigned char *output = (unsigned char*)cJSON_malloc(length + 1);
if (output == NULL) {
    return false;
}

while (*pointer != '\"' && *pointer != '\0') {
    if (*pointer != '\\') {
        *output++ = *pointer++;
    } else {
        pointer++;
        switch (*pointer) {
            case 'b': *output++ = '\b'; break;
            case 'f': *output++ = '\f'; break;
            case 'n': *output++ = '\n'; break;
            case 'r': *output++ = '\r'; break;
            case 't': *output++ = '\t'; break;
            case '\"': *output++ = '\"'; break;
            case '\\': *output++ = '\\'; break;
            case '/': *output++ = '/'; break;
            case 'u':
                /* Unicode 转义,稍后详解 */
                ...
            default:
                cJSON_free(output);
                return false;
        }
        pointer++;
    }
}

第二遍扫描逐字符处理:

  • 普通字符直接复制。
  • 遇到 \ 时,根据后面的字符进行转义映射。
  • 支持的标准转义:\b、\f、\n、\r、\t、\"、\\、\/、\uXXXX。
  • 遇到未知转义(如 \x),释放内存并返回 false。

4.5 Unicode 转义处理

c 复制代码
case 'u':
    /* 解析 4 位十六进制 Unicode 码点 */
    if (parse_hex4(pointer + 1, &uc)) {
        /* 处理代理对 */
        if ((uc >= 0xD800) && (uc <= 0xDBFF) && (*(pointer + 5) == '\\') && (*(pointer + 6) == 'u')) {
            unsigned short uc2 = 0;
            if (parse_hex4(pointer + 7, &uc2)) {
                if ((uc2 >= 0xDC00) && (uc2 <= 0xDFFF)) {
                    /* 计算最终 Unicode 码点 */
                    unsigned long cp = ((uc - 0xD800) << 10) + (uc2 - 0xDC00) + 0x10000;
                    /* 编码为 UTF-8 */
                    ...
                }
            }
        }
        /* 将码点编码为 UTF-8 字节序列 */
        ...
    }
    break;

Unicode 转义是 parse_string 中最复杂的部分,涉及:

  1. 解析 4 位十六进制 :parse_hex4 读取 \uXXXX 中的 4 个十六进制字符。
  2. 代理对处理 :如果码点在 0xD800~0xDBFF(高代理区),且后面紧跟 \u,则继续读取低代理区的 4 位十六进制,组合成完整的 Unicode 码点(范围 0x10000~0x10FFFF)。
  3. UTF-8 编码:将 Unicode 码点转换为 1~4 字节的 UTF-8 序列。

4.6 收尾:写入 valuestring

c 复制代码
*output = '\0';
item->valuestring = (char*)output;
*ep = pointer + 1;
return true;

第二遍扫描结束后,在输出末尾写入 \0,将 output 赋值给 item->valuestring,并更新 ep 指向结束引号之后的位置,返回成功。

5. 关键细节与注意事项

5.1 内存分配策略

parse_string 采用两遍扫描策略:

  • 第一遍只统计长度,不分配内存。
  • 第二遍才分配内存并填充。

这样做的目的是精确分配内存 ,避免过度分配或二次扩容。对于没有转义的字符串,直接用 memcpy 一次复制,性能更高。

5.2 错误处理

函数在以下情况返回 false:

  • 输入为空或开头不是引号。
  • 字符串未闭合(缺少结束引号)。
  • 转义序列不完整(\ 后紧跟 \0)。
  • 未知的转义字符。
  • 内存分配失败。

每次失败前,都会释放已分配的内存,避免内存泄漏。

5.3 与 parse_value 的协作

parse_string 通常由 parse_value 调用。当 parse_value 遇到 " 时,会调用 parse_string 来解析字符串值:

c 复制代码
case '\"':
    return parse_string(item, input, ep);

解析成功后,item->type 会被设置为 cJSON_String,item->valuestring 指向解析出的字符串。

6. 完整源码(精简注释版)

c 复制代码
static cJSON_bool parse_string(cJSON *item, const unsigned char * const input, const unsigned char ** ep)
{
    const unsigned char *pointer = input + 1;
    size_t length = 0;

    if (input == NULL || *input != '\"') {
        return false;
    }

    /* 第一遍:计算长度 */
    while (*pointer != '\"' && *pointer != '\0') {
        if (*pointer == '\\') {
            pointer++;
            if (*pointer == '\0') {
                return false;
            }
            pointer++;
        } else {
            pointer++;
        }
        length++;
    }

    if (*pointer != '\"') {
        return false;
    }

    if (pointer == input + 1) {
        /* 空字符串 */
        item->valuestring = (char*)cJSON_malloc(1);
        if (item->valuestring == NULL) {
            return false;
        }
        item->valuestring[0] = '\0';
        *ep = pointer + 1;
        return true;
    }

    if (length >= (size_t)(pointer - input) - 1) {
        /* 无转义,直接复制 */
        item->valuestring = (char*)cJSON_malloc(length + 1);
        if (item->valuestring == NULL) {
            return false;
        }
        memcpy(item->valuestring, input + 1, length);
        item->valuestring[length] = '\0';
        *ep = pointer + 1;
        return true;
    }

    /* 第二遍:处理转义 */
    pointer = input + 1;
    unsigned char *output = (unsigned char*)cJSON_malloc(length + 1);
    if (output == NULL) {
        return false;
    }

    while (*pointer != '\"' && *pointer != '\0') {
        if (*pointer != '\\') {
            *output++ = *pointer++;
        } else {
            pointer++;
            switch (*pointer) {
                case 'b': *output++ = '\b'; break;
                case 'f': *output++ = '\f'; break;
                case 'n': *output++ = '\n'; break;
                case 'r': *output++ = '\r'; break;
                case 't': *output++ = '\t'; break;
                case '\"': *output++ = '\"'; break;
                case '\\': *output++ = '\\'; break;
                case '/': *output++ = '/'; break;
                case 'u':
                    /* Unicode 处理(省略细节) */
                    break;
                default:
                    cJSON_free(output);
                    return false;
            }
            pointer++;
        }
    }

    *output = '\0';
    item->valuestring = (char*)output;
    *ep = pointer + 1;
    return true;
}

7. 总结

parse_string 是 cJSON 解析器中处理字符串的核心函数,其设计体现了几个重要思想:

  1. 两遍扫描:先统计长度再分配内存,兼顾效率与精确性。
  2. 无转义快路径 :对不含转义的字符串使用 memcpy 批量复制,避免逐字符开销。
  3. 完整的转义支持:覆盖 JSON 规范定义的所有标准转义,并正确处理 Unicode 代理对。
  4. 严谨的错误处理:每个失败路径都释放内存,避免泄漏。

理解 parse_string 的实现,不仅能帮助你深入掌握 cJSON 的工作原理,也能为阅读其他 JSON 解析库(如 yyjson、jansson)的字符串处理逻辑打下良好基础。

相关推荐
小小、码农1 小时前
〖Linux进程间通信〗IPC全家桶:匿名管道、FIFO、共享内存、消息队列与信号量一篇打通
linux·运维·服务器
阳光九叶草LXGZXJ1 小时前
达梦数据库-学习-67-SSL加密认证
linux·运维·数据库·sql·学习·ssl
弹简特1 小时前
【Java项目-企悦抽】13-奖品管理模块-奖品列表和创建奖品的实现
java·开发语言·网络·springboot
0+1111 小时前
Linux --应用层自定义协议与序列化
linux·运维·服务器·网络·tcp
dyxal1 小时前
Linux Crontab 防重复执行利器:flock 文件锁实战详解(脱敏版)
linux·运维·服务器
niuTaylor1 小时前
RK3568 Linux SDK 详解:SDK 是什么、板级差异与完整构建流程
linux·运维·服务器
傻啦嘿哟2 小时前
爬虫代理IP池从0到1:构建高可用代理池,彻底解决IP被封问题
网络·爬虫·tcp/ip
拾贰_C2 小时前
【Ubuntu | port】Linux进程端口号冲突问题解决
linux·运维·ubuntu
szial2 小时前
网络爬虫与 CDP 实战(一):网页数据到底在哪?从 HTTP 到分页采集
网络·爬虫·python·网络爬虫