Using Huge Pages in Linux for Big Data Processing

Using Huge Pages in Linux for Big Data Processing

Huge pages can significantly improve performance for big data processing by reducing TLB (Translation Lookaside Buffer) misses and memory management overhead. Here's how to use them in Linux with C/C++ examples.

1. Configuring Huge Pages in Linux

First, configure huge pages on your system:

bash 复制代码
# Check current huge page settings
cat /proc/meminfo | grep Huge

# Set number of huge pages (e.g., 1024 pages of 2MB each = 2GB)
sudo sysctl vm.nr_hugepages=1024

# Make it persistent by adding to /etc/sysctl.conf
echo "vm.nr_hugepages=1024" | sudo tee -a /etc/sysctl.conf
sudo sysctl -p

2. C/C++ Example Using Huge Pages

Here's a complete example demonstrating huge page allocation:

c 复制代码
#include <stdio.h>
#include <stdlib.h>
#include <sys/mman.h>
#include <fcntl.h>
#include <unistd.h>

#define HUGE_PAGE_SIZE (2 * 1024 * 1024)  // 2MB for x86_64
#define ARRAY_SIZE (1024 * 1024 * 1024)    // 1GB array

// Method 1: Using mmap with MAP_HUGETLB flag
void* allocate_huge_pages_mmap(size_t size) {
    void* ptr = mmap(NULL, size, PROT_READ | PROT_WRITE,
                    MAP_PRIVATE | MAP_ANONYMOUS | MAP_HUGETLB,
                    -1, 0);
    
    if (ptr == MAP_FAILED) {
        perror("mmap");
        return NULL;
    }
    
    printf("Allocated %zu bytes using mmap+MAP_HUGETLB at %p\n", size, ptr);
    return ptr;
}

// Method 2: Using hugetlbfs filesystem
void* allocate_huge_pages_hugetlbfs(size_t size) {
    char path[] = "/dev/hugepages/hugepagefile";
    int fd = open(path, O_CREAT | O_RDWR, 0755);
    
    if (fd < 0) {
        perror("open");
        return NULL;
    }
    
    void* ptr = mmap(NULL, size, PROT_READ | PROT_WRITE, MAP_SHARED, fd, 0);
    if (ptr == MAP_FAILED) {
        perror("mmap");
        close(fd);
        return NULL;
    }
    
    printf("Allocated %zu bytes using hugetlbfs at %p\n", size, ptr);
    close(fd);
    return ptr;
}

void process_large_data(double* data, size_t size) {
    // Simulate big data processing
    for (size_t i = 0; i < size / sizeof(double); i++) {
        data[i] = i * 0.5;
    }
    
    // Do some computation
    double sum = 0;
    for (size_t i = 0; i < size / sizeof(double); i++) {
        sum += data[i];
    }
    
    printf("Processing completed. Sum: %f\n", sum);
}

int main() {
    // Allocate memory using huge pages
    double* huge_data = (double*)allocate_huge_pages_mmap(ARRAY_SIZE);
    if (!huge_data) {
        fprintf(stderr, "Failed to allocate using mmap+MAP_HUGETLB. Trying hugetlbfs...\n");
        huge_data = (double*)allocate_huge_pages_hugetlbfs(ARRAY_SIZE);
        if (!huge_data) {
            fprintf(stderr, "Failed to allocate huge pages. Falling back to regular pages.\n");
            huge_data = (double*)malloc(ARRAY_SIZE);
            if (!huge_data) {
                perror("malloc");
                return 1;
            }
        }
    }
    
    // Process data
    process_large_data(huge_data, ARRAY_SIZE);
    
    // Free memory
    if (munmap(huge_data, ARRAY_SIZE) {
        perror("munmap");
    }
    
    return 0;
}

3. Compiling and Running

Compile the program with:

bash 复制代码
gcc -o hugepage_demo hugepage_demo.c

Run it with:

bash 复制代码
./hugepage_demo

4. Verifying Huge Page Usage

Check huge page usage after running your program:

bash 复制代码
cat /proc/meminfo | grep Huge

5. Important Notes

  1. Permissions: Your program may need appropriate permissions to use huge pages.

  2. Page Size: Default huge page size is typically 2MB. 1GB pages are also available on some systems.

  3. Allocation: Huge page allocation must be contiguous in physical memory.

  4. Transparent Huge Pages (THP) : Linux also supports THP which automatically promotes regular pages to huge pages. Enable with:

    bash 复制代码
    echo "always" | sudo tee /sys/kernel/mm/transparent_hugepage/enabled

6. When to Use Huge Pages

Huge pages are particularly beneficial for:

  • Large in-memory databases
  • Scientific computing applications
  • Big data processing frameworks
  • Any memory-intensive application processing large datasets

The performance improvement comes from reduced TLB pressure and fewer page faults when working with large datasets.


资料

Understanding Huge Pages: Optimizing Memory Usage
Linux HugePages(大内存页) 原理与使用
Performance Benefits of Using Huge Pages for Code
Optimizing Linux for AMD EPYC™ 9005 Series Processors with SUSE Linux Enterprise 15 SP6

相关推荐
爱写代码的森5 小时前
鸿蒙三方库 | harmony-utils之ImageUtil图片保存到本地详解
服务器·华为·harmonyos·鸿蒙·huawei
HLC++8 小时前
Linux的进程间通信
android·linux·服务器
华清远见IT开放实验室9 小时前
实验室建设案例 | 石家庄科技信息职业学院嵌入式实验室——从底层硬件到系统应用,一所应用型高校的嵌入式人才培养这样落地
linux·arm开发·stm32·嵌入式硬件·高校·实验室建设
白露与泡影10 小时前
Arthas 实战指南:从方法耗时定位到 JVM 变量热修改
服务器·jvm·c#
神奇霸王龙11 小时前
Claude Code屠榜:MiMo与Grok紧追Codex
服务器·网络·人工智能·gpt·ai·ai编程
Web极客码11 小时前
WordPress SEO优化:提升网站排名的13个关键步骤
服务器·seo·wordpress
groundhappy12 小时前
idalib安装和codex ida-mcp配置
linux·开发语言·python
通信小小昕13 小时前
Ubuntu 26.04 中文输入法安装
linux·运维·ubuntu
张小姐的猫14 小时前
【Linux】网络编程 —— HTTP协议(上)
linux·运维·服务器·网络·http·单例模式·策略模式
国科安芯14 小时前
星间光链路:AS32S601型抗辐射MCU在空间激光通信终端控制中的技术实现
服务器·网络·单片机·嵌入式硬件·物联网·安全·信息与通信