本篇定位 :Linux "一切皆文件"的根基是 VFS(虚拟文件系统)。11 篇写的 file_operations 是 VFS 给驱动的接口;本篇讲 VFS 怎么把"ext4/proc/sysfs/设备驱动"统一成一套 open/read/write。讲清四大对象(superblock/inode/dentry/file)、挂载、路径查找、procfs/sysfs/debugfs、ext4 概念。读完能理解
cat /proc/cpuinfo和cat /etc/hosts走同一套 VFS 但到不同后端。
VFS 是 Linux "一切皆文件"的实现层裸机操作文件要懂具体文件系统(FAT/自己写)。Linux 在用户态和具体文件系统(ext4/FAT/proc)之间插了 VFS 抽象层 :用户 open() → VFS → 据路径找具体文件系统 → 调它的实现。11 篇的 file_operations 就是 VFS 给字符设备的接口。VFS 让"ext4 文件、proc 信息、设备节点、网络套接字"都用同一套 syscall。
一、VFS 是什么
1.1 VFS 的角色
用户态: open("/etc/hosts") / read / write ...
│
▼ syscall
VFS 层: 统一抽象(open/read/write 通用于所有文件系统)
│
▼ 分发到具体后端
具体 FS: ext4 / FAT / procfs / sysfs / 你的字符驱动
│
▼
底层: 块设备(磁盘)/ 内存(虚拟 FS)/ 设备寄存器
1.2 VFS 的统一接口
不管后端是什么,VFS 提供统一 syscall:
open/read/write/close/lseekmmap/munmapioctlstat/fstatmkdir/readdir/unlink/rename
后端(ext4/proc/驱动)实现这套接口(struct file_operations / inode_operations),VFS 转发。
1.3 支持的文件系统类型
| 类型 | 例子 | 后端 |
|---|---|---|
| 磁盘 FS | ext4 / xfs / btrfs / FAT / NTFS | 块设备 |
| 网络 FS | NFS / CIFS | 网络 |
| 虚拟 FS | procfs / sysfs / debugfs / tmpfs | 内存 |
| 设备 FS | devtmpfs / 字符设备 | 设备 |
| 特殊 | sockfs / pipefs / futexfs | 内核对象 |
"一切皆文件"靠 VFS 落地
/etc/hosts(ext4 磁盘)、/proc/cpuinfo(procfs 内存)、/dev/ttyS0(设备)、/sys/...(sysfs)------都用open/read操作,因为 VFS 统一。你 11 篇的字符驱动 file_operations,就是 VFS 让你接入"一切皆文件"的钩子。
二、VFS 四大对象⭐
2.1 四大对象总览
| 对象 | 代表 | 生命周期 | 存在哪 |
|---|---|---|---|
| superblock | 一个已挂载的文件系统 | 挂载到卸载 | 内存(可能回写磁盘) |
| inode | 一个文件(元数据) | 文件存在 | 内存 + 磁盘 |
| dentry | 路径的一个分量(目录项) | 访问时缓存 | 内存(dcache) |
| file | 一个打开的文件实例 | open 到 close | 内存 |
2.2 superblock(超级块)
c
struct super_block {
dev_t s_dev; // 设备号
unsigned long s_blocksize; // 块大小
unsigned long s_flags; // 挂载标志(MS_RDONLY...)
struct file_system_type *s_type; // 文件系统类型(ext4)
const struct super_operations *s_op; // 操作
struct dentry *s_root; // 根 dentry
struct list_head s_inodes; // 所有 inode 链表
...
};
struct super_operations {
struct inode *(*alloc_inode)(struct super_block *sb);
void (*destroy_inode)(struct inode *);
void (*evict_inode)(struct inode *); // inode 回收
int (*statfs)(...); // 文件系统状态
int (*sync_fs)(...); // 同步
...
};
- 一个挂载的 FS 一个 superblock
- 含 FS 全局信息(块大小/类型/操作表)
- 对应磁盘上的超级块(ext4 在磁盘有副本)
2.3 inode(索引节点)
c
struct inode {
umode_t i_mode; // 文件类型 + 权限
unsigned long i_ino; // inode 号(FS 内唯一)
uid_t i_uid; gid_t i_gid; // 所有者
loff_t i_size; // 文件大小
struct timespec i_atime/mtime/ctime; // 访问/修改/改变时间
const struct inode_operations *i_op; // inode 操作
const struct file_operations *i_fop; // 默认 file 操作
struct address_space *i_mapping; // 页缓存(04 篇)
struct super_block *i_sb; // 所属 superblock
...
};
struct inode_operations {
int (*create)(...); // 创建文件
struct dentry *(*lookup)(...); // 查找
int (*link/unlink/mkdir/rmdir/rename)(...);
int (*getattr/setattr)(...);
...
};
- inode = 一个文件的元数据(大小/权限/时间/数据块位置)
- inode 不含文件名(文件名在 dentry)
- 一个文件一个 inode,即使被多个硬链接
- 磁盘 FS 的 inode 在磁盘有对应(ext4 inode 表)
2.4 dentry(目录项)
c
struct dentry {
struct dentry *d_parent; // 父 dentry
struct qstr d_name; // 名字
struct inode *d_inode; // 关联 inode
const struct dentry_operations *d_op;
struct super_block *d_sb;
unsigned char d_flags;
struct list_head d_subdirs; // 子 dentry
...
};
- dentry = 路径的一个分量(/etc/hosts → "etc" 和 "hosts" 两个 dentry)
- dentry 把"文件名"和"inode"关联
- dcache:dentry 缓存,加速路径查找
- 纯内存对象(磁盘 FS 不存 dentry,存"目录文件"里)
2.5 file(打开的文件)
c
struct file {
const struct file_operations *f_op; // 操作(你 11 篇写的)
struct dentry *f_dentry; // 关联 dentry
struct inode *f_inode;
loff_t f_pos; // 当前偏移(每个 file 独立)
unsigned int f_flags; // open 标志(O_RDONLY...)
fmode_t f_mode; // 读/写模式
void *private_data; // 驱动私有(你 11 篇用)
...
};
- file = 一次 open 的实例
- 每次 open 创建一个 file(同文件多次 open 有多个 file)
- f_pos 是该实例的偏移(独立)
- 11 篇的
filp->private_data就是这里
2.6 四对象关系
superblock (一个 FS)
└─ 根 dentry
└─ 子 dentry (etc)
└─ 子 dentry (hosts) ──> inode (hosts 的元数据)
└─ 多个 file (多次 open 的实例)
└─ file_operations (驱动实现)
四对象是 VFS 的核心数据结构
你读内核文件系统代码,全是这四对象。记忆:superblock=FS;inode=文件元数据(无名字);dentry=名字+inode(路径分量);file=open 实例(有偏移)。你 11 篇的 file_operations 挂在 inode->i_fop,open 时拷到 file->f_op。
三、挂载(mount)
3.1 挂载做什么
- 把一个文件系统的根 dentry 挂到挂载点目录
- 访问挂载点 = 访问该 FS
mount -t ext4 /dev/sda1 /mnt
3.2 挂载流程
mount("/dev/sda1", "/mnt", "ext4", ...)
→ VFS 找 ext4 文件系统类型
→ 调 ext4_mount → 读磁盘超级块 → 建内存 superblock
→ 建根 dentry
→ 挂到 /mnt(挂载点 dentry)
3.3 文件系统类型注册
c
// 每个文件系统注册一个 file_system_type
struct file_system_type ext4_fs_type = {
.name = "ext4",
.mount = ext4_mount,
.kill_sb = ext4_kill_sb,
...
};
register_filesystem(&ext4_fs_type);
3.4 看挂载
bash
mount # 列所有挂载
cat /proc/mounts # 同上(更详细)
findmnt # 树形
3.5 根文件系统(rootfs)
- 内核启动第一个挂载的 FS(02 篇)
- 通常是 initramfs(内存)或 ext4(磁盘)
- PID 1(init)在 rootfs 上
嵌入式视角:嵌入式挂载
嵌入式常见:① rootfs = initramfs(打包进内核);② /mnt/data = ext4(eMMC 分区);③ /mnt/usb = FAT(U 盘)。你 BSP 要配挂载脚本(/etc/fstab 或 init 脚本)。
四、路径查找(path lookup)
4.1 open("/etc/hosts") 怎么找
1. 从根 dentry(/)开始
2. 取下一分量 "etc" → 在根 dentry 的子目录找 "etc" dentry
3. 取下一分量 "hosts" → 在 "etc" dentry 的子目录找 "hosts" dentry
4. "hosts" dentry 关联 inode
5. 创建 file,关联 inode,调 inode->i_fop->open
4.2 dcache 加速
- 找过的 dentry 缓存到 dcache
- 下次同路径直接命中,不查磁盘
cat /etc/hosts第二次比第一次快(dcache 命中)
4.3 符号链接
- 符号链接是特殊 dentry,内容是"目标路径"
- 查找时遇到符号链接,跳到目标继续找
五、read/write 的完整路径(结合 VFS)
5.1 read(fd, buf, n) 完整
用户态: read(fd, buf, n)
→ glibc → syscall(06 篇)
→ sys_read
→ VFS vfs_read
→ 从 fd 找 file
→ 调 file->f_op->read(你驱动的 my_read,或 ext4 的 ext4_read)
→ 字符设备:copy_to_user(11 篇)
→ ext4:走 page cache(04 篇)→ 块 IO(13 篇)→ 磁盘
→ 返回读字节数
→ sret 回用户态
5.2 不同后端,f_op 不同
| 后端 | file_operations |
|---|---|
| 字符设备(你写的) | 你的 my_fops |
| ext4 | ext4_file_operations |
| procfs | proc_file_operations |
| sysfs | sysfs_file_operations |
| 套接字 | socket_file_ops |
VFS 据 file->f_op 转发,统一接口不同实现------这就是多态。
VFS 是 C 实现的多态
11 篇写 file_operations 就是"实现 VFS 接口"。VFS 据 file->f_op 调对应实现------字符驱动调你的,ext4 调 ext4 的。C 用函数指针表(file_operations)实现多态,裸机的"函数指针表注册"经验直接迁移,Linux 是它的系统化。
六、inode 操作 vs file 操作
6.1 两种操作表
| 操作表 | 关联 | 作用 |
|---|---|---|
inode_operations (i_op) |
inode | 文件级(create/lookup/mkdir/unlink/rename) |
file_operations (i_fop / f_op) |
inode/file | 内容级(open/read/write/ioctl) |
6.2 区别
- inode_operations:操作"文件本身"(创建/删除/改名/查找)------目录操作为主
- file_operations:操作"文件内容"(读/写/控制)------打开文件后
6.3 例子
mkdir /tmp/newdir
→ VFS → 父目录 inode_operations.mkdir
echo hi > /tmp/newdir/file
→ open(create): 父目录 inode_operations.create
→ write: file_operations.write
cat /tmp/newdir/file
→ open: inode_operations.lookup(找 file)
→ read: file_operations.read
七、Page Cache 与 VFS(衔接 04 篇)
7.1 address_space
- 每个 inode 有
address_space,管该文件的页缓存 - read 命中 page cache 直接返回,不命中读盘
- write 写 page cache,标记 dirty,writeback 写盘
7.2 VFS read(磁盘文件)
vfs_read → ext4_read_iter
→ address_space 读 page cache
命中 → 拷到用户
不命中 → 发块 IO 请求读盘(13 篇)→ 填 page cache → 拷用户
- 字符设备不走 page cache(你 my_read 自己处理)
- 磁盘 FS 走 page cache(04 篇)
八、虚拟文件系统:procfs / sysfs / debugfs
8.1 procfs(/proc)
-
内核状态导出,内存 FS
-
旧接口,杂(进程信息 + 内核参数混杂)
/proc/cpuinfo CPU 信息
/proc/meminfo 内存
/proc/interrupts 中断
/proc// 进程信息(maps/status/fd/...)
/proc/sys/ 可改内核参数(sysctl)
8.2 procfs 编程接口
c
// 创建 /proc/myinfo
struct proc_dir_entry *entry = proc_create("myinfo", 0444, NULL, &my_fops);
// 用 file_operations(read 返回字符串)
remove_proc_entry("myinfo", NULL);
- 老接口
create_proc_read_entry弃用 - 新用
proc_create+ seq_file(处理大数据量)
8.3 sysfs(/sys)
- 设备模型导出(10 篇),每个 kobject 一个目录
- attribute 一个文件,show/store 回调
- 比 proc 规范(只放设备模型相关)
8.4 debugfs(/sys/kernel/debug)
- 调试用,无 API 稳定性保证
- 简单易用,适合驱动调试
c
struct dentry *d = debugfs_create_u32("myval", 0644, NULL, &my_var);
debugfs_create_file("myinfo", 0444, NULL, dev, &my_fops);
debugfs_remove_recursive(d);
8.5 三个虚拟 FS 的取舍
| FS | 用途 | 稳定性 |
|---|---|---|
| procfs | 进程信息 + 内核参数(老) | 半稳定(/proc/ 稳定,/proc/sys 稳定,其他不保证) |
| sysfs | 设备模型(10 篇) | 稳定(设备模型接口) |
| debugfs | 调试 | 不保证(随时改) |
驱动调试信息放哪
BSP 写驱动要导出调试信息:① 生产配置 用 sysfs(DEVICE_ATTR,11 篇);② 调试临时 用 debugfs(简单);③ 进程/内核状态用 proc(老接口)。别把调试信息塞 proc(污染)。这是规范。
九、ext4 概念(磁盘 FS 代表)
9.1 ext4 磁盘结构
超级块(Superblock) ← FS 全局信息(块大小/inode 数/空闲块)
块组描述符表(GDT) ← 每块组的信息
块组(Block Group):
├─ 超级块副本(部分块组)
├─ 块位图(Block Bitmap) ← 哪些块空闲
├─ inode 位图 ← 哪些 inode 空闲
├─ inode 表 ← 该组所有 inode
└─ 数据块 ← 文件数据
9.2 ext4 特性
- ** extents**(区段):文件数据用 extents 描述(连续块),比 ext2/3 的间接块高效
- journaling(日志):元数据变更先写日志,崩溃可恢复
- 大文件支持:up to 16TB(块 4K)
- 延迟分配:数据先缓存,批量分配块,减少碎片
9.3 ext4 inode
- 每个 inode 固定大小(常 256 字节)
- 含文件元数据 + 数据块指针(extents)
- inode 号是 FS 内唯一标识
9.4 写文件(ext4)
write(fd, buf, n)
→ vfs_write → ext4_file_write_iter
→ 写 page cache(标记 dirty)
→ 延迟分配(不立即分配块)
→ writeback 线程:分配块(extents)+ 写日志 + 写数据块
9.5 fsck 与日志
- 日志模式:journal(数据+元数据)/ordered(元数据,数据先写)/writeback(元数据,数据不保证)
- 崩溃后日志恢复快(不用全盘 fsck)
嵌入式视角:嵌入式常用更轻的 FS
ext4 功能强但复杂。嵌入式常:① SquashFS (只读压缩,rootfs);② JFFS2/UBIFS (NAND Flash,带磨损均衡);③ FAT (兼容性,U 盘);④ ext4(eMMC/SD 大容量)。BSP 按 storage 选。
十、dentry/inode 缓存
10.1 dcache(dentry 缓存)
- 访问过的 dentry 缓存
- 内存压力时回收
10.2 icache(inode 缓存)
- inode 缓存
- 配合 inode_hashtable(ino → inode)
10.3 看缓存
bash
cat /proc/sys/fs/dentry-state
# nr_dentry nr_unused age_limit want_pages nr_negative
十一、文件描述符(fd)
11.1 fd 是什么
- 进程级的"打开文件"索引
- 每个 fd 对应一个 file 对象
11.2 fd 表
c
struct task_struct {
struct files_struct *files; // fd 表
};
struct files_struct {
struct file __rcu **fdt; // fd 数组
...
};
fdt[fd]→ struct file- 0/1/2 默认 stdin/stdout/stderr
- open 返回最小可用 fd
11.3 看进程 fd
bash
ls -l /proc/<pid>/fd/
# 0 -> /dev/null
# 1 -> /dev/pts/0
# 3 -> /etc/hosts
# 4 -> socket:[12345]
十二、对照总表
| 概念 | 裸机/FreeRTOS | Linux VFS |
|---|---|---|
| 文件操作 | 自己写 FAT | VFS 统一 open/read/write |
| 文件元数据 | FAT 表项 | inode |
| 文件名 | FAT 目录项 | dentry |
| 打开文件 | 自己管句柄 | file + fd |
| 多 FS 支持 | 单一 | VFS 抽象 + 多后端 |
| 设备 | 直接访问 | file_operations(11 篇) |
| 内核状态 | 自己打印 | procfs/sysfs |
| 挂载 | 无 | mount |
十三、本篇小结
- VFS:用户态和具体 FS 之间的抽象层,统一 open/read/write,后端可换(ext4/proc/驱动)
- 四对象:superblock(FS)/inode(文件元数据,无名)/dentry(名字+inode,路径分量)/file(open 实例,有偏移)
- 挂载:把 FS 根 dentry 挂到挂载点;file_system_type 注册
- 路径查找:从根 dentry 逐级查,dcache 加速
- read 路径:sys_read → vfs_read → file->f_op->read(字符驱动直接/ext4 走 page cache)
- inode_operations(文件级:create/lookup/mkdir)vs file_operations(内容级:read/write)
- procfs :进程/内核信息(老);sysfs :设备模型(规范);debugfs:调试(临时)
- ext4:超级块/块组/inode 表/extents/日志,延迟分配
- dcache/icache 加速;fd 是进程级打开文件索引
- VFS 是 C 用函数指针表实现的多态,你 11 篇的 file_operations 就是接入点
速查表
| 想干啥 | 用什么 |
|---|---|
| 看挂载 | mount / cat /proc/mounts / findmnt |
| 看 fd | ls -l /proc//fd/ |
| 看进程内存映射 | cat /proc//maps |
| 看文件系统缓存 | cat /proc/sys/fs/dentry-state |
| 看 inode 信息 | stat file |
| 内核参数 | /proc/sys/ + sysctl |
| 设备信息 | /sys/(sysfs) |
| 调试信息 | /sys/kernel/debug/(debugfs) |
| 创建 proc 项 | proc_create(name, perm, parent, fops) |
| 创建 debugfs | debugfs_create_u32/file |
| 注册文件系统 | register_filesystem(&fs_type) |
| 看 FS 类型 | df -T |
💡技术之路漫漫,分享是为了更好地交流。如果本文的内容对你有启发,希望能得到你的 点赞 👍 和 收藏 ⭐。
如果你在调试过程中遇到了其他问题,欢迎在 评论区 💬 留言,我们一起探讨。也欢迎 关注 👀 我,一起交流底层开发的那些事儿。