2.1概要说明
本文章的目标,深入分析时钟虚拟化原理比如kvm-clock的初始化流程,以及解释时间从host端传到guest端的原理等。阅读该部分内容需要了解内存qemu内存虚拟化,vm-exit vm-enter虚拟化机制,调试验证的方式有:编译kvm增加打印日志、用bpftrace追踪或者添加dump_stack()函数。
首先需要了解时钟在linux系统中是处于的层级,虚拟机中用户态执行一次clock_gettime() 或者 gettimeofday()或者time()后,用户态和内核态的调用逻辑。对于用户态接口来说获取时间有两种方式:一种是通过vDSO(virtual dynamic shared object)的方式直接读取,另一种是间接通过syscall的方式调用内核接口。在内核态获取时间首先交互的是内核态的时钟模型time_keeping,该数据结构是对系统时间的建模,在host端可以说无论是内核态还是用户态要获取时间都离不开time_keeping里的数据。而time_keeping里的数据是通过对下层时钟接口调用的返回值组装的即clock_source,clock_source有很多tsc、hpte、kvm_clock等,不同的时钟区别是获取时间或者说是计量时间的方式不同,因此时钟设备和系统时间是两个概念。时钟虚拟化本质就是对计量时间方式的虚拟化,具体来说不同的clock_source对应的read函数是不一样的,获取wall_clock的函数也不一样。
2.2 kvm-clock时钟虚拟化
Kvm-clock是qemu-kvm虚拟机安装linux系统时官方推荐的虚拟时钟,具有准确、稳定、快速等优点的半虚拟化时钟是主流的虚拟时钟。本文章从该时钟的初始化和时钟更新等方面进行介绍。由于设计到的内容是非常多的,所以我们以从整体到模块再到细节的顺序对kvm-clock进行介绍。首先整体上根据交互对象可分为三部分,host部分、guest部分和guest的用户程部分。接着分别介绍这三部分主要的工作,重点介绍交互流程。最后在guest的用户程序中介绍如何拿到时间数据的。然后介绍time_keeping模型,kvm_clock模型中涉及的的时间概念,明确交互变量的含义,最后结合工程实践说明时间虚拟化中可能会遇到的问题。
2.1.1整体数据传输框架
我们通过一个具体的实例来了解kvm_clock的数据传输。回答当虚拟机内的应用程序执行clock_gettime(CLOCK_MONOTONIC,&ts)时数据是怎么获取的,获取的又是什么内容? kvm_clock整体的框架就是guest和host通过ram内存进行传输数据,host端计算相应的数据通过kvm对内存的虚拟化,找到ram内存的hva地址,进行数据写入工作,完成对数据的传递,这是本质内容。

从上图可以看到(1)guest端从ram内存中申请了struct pvclock_vsyscall_time_info*类型的per_cpu内存和struct pvclock_wall_clock类型的wall_clock内存,guest通过rwmsr的方式与host端的kvm交互。(2)host端通过kvm对guest虚拟内存的地址映射表,找到guest中per_cpu变量hv_clock_per_cpu和wall_clock的gha地址对应的hva地址。这样host端就可以对per_cpu的hv_clock_per_cpu变量和wall_clock变量进行写入,guest端就直接获取了相应的时钟信息。(3)对于guest的用户程序,通过vDSO(virtual dynamic shared object)技术用mapping的方式把用户态程序特定的地址映射到host端per_cpu变量hv_clock_per_cpu的pha上实现无系统调用的访问,这样用户程序就拿到了时钟信息。
2.1.2 三个模块的具体工作
从前面的叙述我们已经了解到时钟信息是在host端获取的,通过内存传输到了guest端,用户程序又通过vDSO(virtual dynamic shared object)实现0拷贝的访问,需要补充的一点从host端到guest端的传递在虚拟机启动和重启的时候会出现,在此之后guest内部通过rdtsc的方式维护guest kernel中的内核时间模型,因此与host的交互就很少了。现在具体看看三个模块都做了什么工作。
2.1.2.1 guest端kvm-clock时钟的初始化
在guest内核中如果能查到kvm_clock的配置,就会调用相应的初始化以及注册函数。初始化函数调用逻辑start_kernel=>setup_arch=>init_hypervisor_platform=>x86_init.hyper.init_platform()=>x86_hyper_kvm.init.init_platform = kvm_init_platform=>kvm_init_platform->kvmclock_init
相应的初始化函数源码为:
cpp
void __init kvmclock_init(void)
{
u8 flags;
if (!kvm_para_available() || !kvmclock)
return;
------------------msr寄存器地址的设置---------------------------------------------------
if (kvm_para_has_feature(KVM_FEATURE_CLOCKSOURCE2)) {
msr_kvm_system_time = MSR_KVM_SYSTEM_TIME_NEW;
msr_kvm_wall_clock = MSR_KVM_WALL_CLOCK_NEW;
} else if (kvm_para_has_feature(KVM_FEATURE_CLOCKSOURCE)) {
msr_kvm_system_time = MSR_KVM_SYSTEM_TIME;
msr_kvm_wall_clock = MSR_KVM_WALL_CLOCK;
} else {
return;
}
/ * 初始化其他的percpu */
if (cpuhp_setup_state(CPUHP_BP_PREPARE_DYN, "kvmclock:setup_percpu",
kvmclock_setup_percpu, NULL) < 0) {
return;
}
pr_info("kvm-clock: Using msrs %x and %x",
msr_kvm_system_time, msr_kvm_wall_clock);
-------------------设置percpu变量hv_clock_per_cpu---------------------------------------
this_cpu_write(hv_clock_per_cpu, &hv_clock_boot[0]);
kvm_register_clock("primary cpu clock");
pvclock_set_pvti_cpu0_va(hv_clock_boot);
if (kvm_para_has_feature(KVM_FEATURE_CLOCKSOURCE_STABLE_BIT))
pvclock_set_flags(PVCLOCK_TSC_STABLE_BIT);
flags = pvclock_read_flags(&hv_clock_boot[0].pvti)
kvm_sched_clock_init(flags & PVCLOCK_TSC_STABLE_BIT);
-------------------设置时钟跟新回调函数---------------------------------------------------
x86_platform.calibrate_tsc = kvm_get_tsc_khz;
x86_platform.calibrate_cpu = kvm_get_tsc_khz;
x86_platform.get_wallclock = kvm_get_wallclock;
x86_platform.set_wallclock = kvm_set_wallclock;
#ifdef CONFIG_X86_LOCAL_APIC
x86_cpuinit.early_percpu_clock_init = kvm_setup_secondary_clock;
#endif
x86_platform.save_sched_clock_state = kvm_save_sched_clock_state;
x86_platform.restore_sched_clock_state = kvm_restore_sched_clock_state;
kvm_get_preset_lpj();
/*
* X86_FEATURE_NONSTOP_TSC is TSC runs at constant rate
* with P/T states and does not stop in deep C-states.
*
* Invariant TSC exposed by host means kvmclock is not necessary:
* can use TSC as clocksource.
*
*/
if (boot_cpu_has(X86_FEATURE_CONSTANT_TSC) &&
boot_cpu_has(X86_FEATURE_NONSTOP_TSC) &&
!check_tsc_unstable())
kvm_clock.rating = 299;
/* 注册kvm_clock 时钟设备 */
clocksource_register_hz(&kvm_clock, NSEC_PER_SEC);
pv_info.name = "KVM";
}
以上是完整的kvmclock_init代码,主要是设置变量地址,初始化相关回调函数。我们先着重明确变量的初始化以及相应的回调函数。
1:sysclock变量设置
1.1 percpu变量hv_clock_per_cpu设置定义hv_clock_per_cpu
cpp
------------------------定义hv_clock_boot数组的大小---------------------------------------
#define HVC_BOOT_ARRAYSIZE \
(PAGE_SIZE / sizeof(struct pvclock_vsyscall_time_info))
--------------------------定义hv_clock_boot数组在.bss.decrypted区域并4k对齐---------------
static struct pvclock_vsyscall_time_info
hv_clock_boot[HVC_BOOT_ARRAY_SIZE] __bss_decrypted __aligned(PAGE_SIZE);
-----------------------定义hv_clock_per_cpu----------------------------------------------
DEFINE_PER_CPU(struct pvclock_vsyscall_time_info *, hv_clock_per_cpu);
EXPORT_PER_CPU_SYMBOL_GPL(hv_clock_per_cpu);
static int kvmclock_setup_percpu(unsigned int cpu)
{
struct pvclock_vsyscall_time_info *p = per_cpu(hv_clock_per_cpu, cpu);
/*
* The per cpu area setup replicates CPU0 data to all cpu
* pointers. So carefully check. CPU0 has been set up in init
* already.
*/
if (!cpu || (p && p != per_cpu(hv_clock_per_cpu, 0)))
return 0;
/* Use the static page for the first CPUs, allocate otherwise */
if (cpu < HVC_BOOT_ARRAY_SIZE)
p = & hv_clock_boot [cpu];
else if (hvclock_mem)
p = hvclock_mem + cpu - HVC_BOOT_ARRAY_SIZE;
else
return -ENOMEM;
per_cpu(hv_clock_per_cpu, cpu) = p;
return p ? 0 : -ENOMEM;
}
这个是开了个page空间内存,然后创建pvclock_vsyscall_time_info类型的数组,因此这个数据的空间肯定是在虚拟机的ram里的,然后注册struct pvclock_vsyscall_time_info *类型的per_cpu变量hv_clock_per_cpu并在函数kvmclock_setup_percpu中把hv_clock_per_cpu赋值为hv_clock_boot数组中的一个元素地址,即per_cpu的hv_clock_per_cpu变量执行的是ram内存中的一段空间。
对应的pvclock_vsyscall_time_info数据结构:
cpp
struct pvclock_vsyscall_time_info {
struct pvclock_vcpu_time_info pvti;
} __attribute__((__aligned__(SMP_CACHE_BYTES)));
//一切不言尽在注释中
/*
* These structs MUST NOT be changed.
* They are the ABI between hypervisor and guest OS.
* Both Xen and KVM are using this.
*
* pvclock_vcpu_time_info holds the system time and the tsc timestamp
* of the last update. So the guest can use the tsc delta to get a
* more precise system time. There is one per virtual cpu.
*
* pvclock_wall_clock references the point in time when the system
* time was zero (usually boot time), thus the guest calculates the
* current wall clock by adding the system time.
*
* Protocol for the "version" fields is: hypervisor raises it (making
* it uneven) before it starts updating the fields and raises it again
* (making it even) when it is done. Thus the guest can make sure the
* time values it got are consistent by checking the version before
* and after reading them.
*/
struct pvclock_vcpu_time_info {
u32 version;
u32 pad0;
u64 tsc_timestamp;
u64 system_time;
u32 tsc_to_system_mul;
s8 tsc_shift;
u8 flags;
u8 pad[2];
} __attribute__((__packed__)); /* 32 bytes */
1.2 wall_clock数据结构的定义
cpp
static struct pvclock_wall_clock wall_clock __bss_decrypted;
static struct pvclock_vsyscall_time_info *hvclock_mem;
/* It is OK to have a 12 bytes struct with no padding because it is packed */
struct pvclock_wall_clock {
u32 version;
u32 sec;
u32 nsec;
u32 sec_hi;
} __attribute__((__packed__));
明确了这些变量我们在梳理一下kvmcloc_init里的逻辑关系。首先是寄存器地址的赋值:
cpp
msr_kvm_system_time = MSR_KVM_SYSTEM_TIME_NEW;
msr_kvm_wall_clock = MSR_KVM_WALL_CLOCK_NEW;
/*对应的地址是*/
#define MSR_KVM_WALL_CLOCK_NEW 0x4b564d00
#define MSR_KVM_SYSTEM_TIME_NEW 0x4b564d01
接着注册kvm_clock时钟源
cpp
static void kvm_register_clock(char *txt)
{
/* 获取percpu中的struct pvclock_vsyscall_time_info *地址 */
struct pvclock_vsyscall_time_info *src = this_cpu_hvclock();
u64 pa;
if (!src)
return;
/* 转化为pha 虚拟机的物理地址 */
pa = slow_virt_to_phys(&src->pvti) | 0x01ULL;
/* 写入到对应percpu的msr寄存器中 */
wrmsrl(msr_kvm_system_time, pa);
pr_debug("kvm-clock: cpu %d, msr %llx, %s", smp_processor_id(), pa, txt);
}
void pvclock_set_pvti_cpu0_va(struct pvclock_vsyscall_time_info *pvti)
{
WARN_ON(vclock_was_used(VDSO_CLOCKMODE_PVCLOCK));
pvti_cpu0_va = pvti;
}
3.设置回调函数:
/*
* The wallclock is the time of day when we booted. Since then, some time may
* have elapsed since the hypervisor wrote the data. So we try to account for
* that with system time
*/
static void kvm_get_wallclock(struct timespec64 *now)
{
-------------------------写入MSR寄存器触发vm exit-----------------------------------------
//回忆kvmclock_init中msr_kvm_wall_clock = MSR_KVM_WALL_CLOCK_NEW
// wall_clock为定义的全局变量
wrmsrl(msr_kvm_wall_clock, slow_virt_to_phys(&wall_clock));
preempt_disable();
------------------------读取从host端返回来的信息------------------------------------------
pvclock_read_wallclock(&wall_clock, this_cpu_pvti(), now);
preempt_enable();
}
static int kvm_set_wallclock(const struct timespec64 *now)
{
return -ENODEV;
}
void pvclock_read_wallclock(struct pvclock_wall_clock *wall_clock,
struct pvclock_vcpu_time_info *vcpu_time,
struct timespec64 *ts)
{
u32 version;
u64 delta;
struct timespec64 now;
/* get wallclock at system boot */
do {
version = wall_clock->version;
rmb(); /* fetch version before time */
/*
* Note: wall_clock->sec is a u32 value, so it can
* only store dates between 1970 and 2106. To allow
* times beyond that, we need to create a new hypercall
* interface with an extended pvclock_wall_clock structure
* like ARM has.
*/
now.tv_sec = wall_clock->sec;
now.tv_nsec = wall_clock->nsec;
rmb(); /* fetch time before checking version */
} while ((wall_clock->version & 1) || (version != wall_clock->version));
/* wall_clock和 sys_clock结合最后得到now的时间 */
delta = pvclock_clocksource_read(vcpu_time); /* time since system boot */
delta += now.tv_sec * NSEC_PER_SEC + now.tv_nsec;
now.tv_nsec = do_div(delta, NSEC_PER_SEC);
now.tv_sec = delta;
set_normalized_timespec64(ts, now.tv_sec, now.tv_nsec);
}
阅读代码可分析出kvm_get_wallclock就是通过wrmsr退出虚拟机进入kvm读取host里的信息拼接wall_clock然后再复制到guest系统的变量中。
cpp
/*
* If we don't do that, there is the possibility that the guest
* will calibrate under heavy load - thus, getting a lower lpj -
* and execute the delays themselves without load. This is wrong,
* because no delay loop can finish beforehand.
* Any heuristics is subject to fail, because ultimately, a large
* poll of guests can be running and trouble each other. So we preset
* lpj here
*/
static unsigned long kvm_get_tsc_khz(void)
{
setup_force_cpu_cap(X86_FEATURE_TSC_KNOWN_FREQ);
/* 获取percpu里的hv_clock_per_cpu变量并返回pvti成员 */
return pvclock_tsc_khz(this_cpu_pvti());
}
unsigned long pvclock_tsc_khz(struct pvclock_vcpu_time_info *src)
{
u64 pv_tsc_khz = 1000000ULL << 32;
do_div(pv_tsc_khz, src->tsc_to_system_mul);
if (src->tsc_shift < 0)
pv_tsc_khz <<= -src->tsc_shift;
else
pv_tsc_khz >>= src->tsc_shift;
return pv_tsc_khz;
}
static __always_inline struct pvclock_vcpu_time_info *this_cpu_pvti(void)
{
return &this_cpu_read(hv_clock_per_cpu)->pvti;
}
根据这些回调函数可以明确的感知到都跟percpu的pvti变量有关。x86_platform里面的回调函数是time_keep相关的很重要的回调函数,都是为了时钟初始化服务的。
最后是注册时钟设备频率:
cpp
static struct clocksource kvm_clock = {
.name = "kvm-clock",
.read = kvm_clock_get_cycles,
.rating = 400,
.mask = CLOCKSOURCE_MASK(64),
.flags = CLOCK_SOURCE_IS_CONTINUOUS,
.id = CSID_X86_KVM_CLK,
.enable = kvm_cs_enable,
};
我们看一下这个read函数:
static u64 kvm_clock_get_cycles(struct clocksource *cs)
{
return kvm_clock_read();
}
static u64 kvm_clock_read(void)
{
u64 ret;
preempt_disable_notrace();
ret = pvclock_clocksource_read_nowd(this_cpu_pvti());
preempt_enable_notrace();
return ret;
}
noinstr u64 pvclock_clocksource_read_nowd(struct pvclock_vcpu_time_info *src)
{
return __pvclock_clocksource_read(src, false);
}
static __always_inline
u64 __pvclock_clocksource_read(struct pvclock_vcpu_time_info *src, bool dowd)
{
unsigned version;
u64 ret;
u64 last;
u8 flags;
do {
version = pvclock_read_begin(src);
ret = __pvclock_read_cycles(src, rdtsc_ordered());
flags = src->flags;
} while (pvclock_read_retry(src, version));
if (dowd && unlikely((flags & PVCLOCK_GUEST_STOPPED) != 0)) {
src->flags &= ~PVCLOCK_GUEST_STOPPED;
pvclock_touch_watchdogs();
}
if ((valid_flags & PVCLOCK_TSC_STABLE_BIT) &&
(flags & PVCLOCK_TSC_STABLE_BIT))
return ret;
/*
* Assumption here is that last_value, a global accumulator, always goes
* forward. If we are less than that, we should not be much smaller.
* We assume there is an error margin we're inside, and then the correction
* does not sacrifice accuracy.
*
* For reads: global may have changed between test and return,
* but this means someone else updated poked the clock at a later time.
* We just need to make sure we are not seeing a backwards event.
*
* For updates: last_value = ret is not enough, since two vcpus could be
* updating at the same time, and one of them could be slightly behind,
* making the assumption that last_value always go forward fail to hold.
*/
last = raw_atomic64_read(&last_value);
do {
if (ret <= last)
return last;
} while (!raw_atomic64_try_cmpxchg(&last_value, &last, ret));
return ret;
}
static __always_inline
u64 __pvclock_read_cycles(const struct pvclock_vcpu_time_info *src, u64 tsc)
{
u64 delta = tsc - src->tsc_timestamp;
u64 offset = pvclock_scale_delta(delta, src->tsc_to_system_mul,
src->tsc_shift);
return src->system_time + offset;
}
可以看到读取的就是注册的pvclock_vcpu_time_info信息与tsc值的组合
以上内容需要了解到,kvmclock_init干了两件事(1)申请了空间注册了变量,(2)赋值函数。这些都是为了给time_keep使用的,最终要作用到time_keep的变量或者函数里。这时候time_keeping(内核时钟模型)的材料已经准备好。
4.回调函数的调用时机
内核启动时:start_kernel->timekeeping_init->read_persistent_wall_and_boot_offset->read_persistent_clock64-> x86_platform.get_wallclock(ts)->x86_platform.get_wallclock = kvm_get_wallclock
我们就着重在timekeeping_init中进行梳理:
cpp
void __init timekeeping_init(void)
{
struct timespec64 wall_time, boot_offset, wall_to_mono;
struct timekeeper *tk = &tk_core.timekeeper;
struct clocksource *clock;
unsigned long flags;
read_persistent_wall_and_boot_offset(&wall_time, &boot_offset);
if (timespec64_valid_settod(&wall_time) &&
timespec64_to_ns(&wall_time) > 0) {
persistent_clock_exists = true;
} else if (timespec64_to_ns(&wall_time) != 0) {
pr_warn("Persistent clock returned invalid value");
wall_time = (struct timespec64){0};
}
if (timespec64_compare(&wall_time, &boot_offset) < 0)
boot_offset = (struct timespec64){0};
/*
* We want set wall_to_mono, so the following is true:
* wall time + wall_to_mono = boot time
*/
wall_to_mono = timespec64_sub(boot_offset, wall_time);
raw_spin_lock_irqsave(&timekeeper_lock, flags);
write_seqcount_begin(&tk_core.seq);
ntp_init();
clock = clocksource_default_clock();
if (clock->enable)
clock->enable(clock);
tk_setup_internals(tk, clock);
tk_set_xtime(tk, &wall_time);
tk->raw_sec = 0;
tk_set_wall_to_mono(tk, wall_to_mono);
/* 把tk值更新到pvclock_gtod_data中 */
timekeeping_update(tk, TK_MIRROR | TK_CLOCK_WAS_SET);
write_seqcount_end(&tk_core.seq);
raw_spin_unlock_irqrestore(&timekeeper_lock, flags);
}
void __init read_persistent_wall_and_boot_offset(struct timespec64 *wall_time,
struct timespec64 *boot_offset)
{
struct timespec64 boot_time;
union tod_clock clk;
u64 delta;
delta = initial_leap_seconds + TOD_UNIX_EPOCH;
clk = tod_clock_base;
clk.eitod -= delta;
ext_to_timespec64(&clk, &boot_time);
read_persistent_clock64(wall_time);
*boot_offset = timespec64_sub(*wall_time, boot_time);
}
/* not static: needed by APM */
void read_persistent_clock64(struct timespec64 *ts)
{
/* 在host端就是mach_get_cmos_time*/
x86_platform.get_wallclock(ts);
}
到这里我们已经明确在kvmclock_init中初始化的get_wallclock是如何被使用的,下面就是明确其功能是要干什么的。
如下是其对应的源码:
cpp
static void kvm_get_wallclock(struct timespec64 *now)
{
---------------------------------写入MSR寄存器触发vm exit-----------------------------------
//回忆kvmclock_init中msr_kvm_wall_clock = MSR_KVM_WALL_CLOCK_NEW
// wall_clock为定义的全局变量
wrmsrl(msr_kvm_wall_clock, slow_virt_to_phys(&wall_clock));
preempt_disable();
---------------------------------读取从host端返回来的信息-----------------------------------
pvclock_read_wallclock(&wall_clock, this_cpu_pvti(), now);
preempt_enable();
}
我们先看wall_clock是如何获取的。在guest环境写入smr操作会触发vm-exit,并在host kvm端找到对应的回调函数,可以在host端执行bpftrace脚本进行验证;
bash
#!/usr/bin/bpftrace
#include <linux/fs.h>
#include <linux/virtio.h>
#define MSR_IA32_TSC 0x00000010
#define MSR_P6_PERFCTR1 0x000000C2
#define MSR_KVM_SYSTEM_TIME_NEW 0x4b564d01
#define MSR_KVM_WALL_CLOCK_NEW 0x4b564d00
struct msr_data {
bool host_initiated;
u32 index;
u64 data;
};
kprobe:kvm_set_msr_common
{
@msr_data = arg1;
$index = ((struct msr_data*)@msr_data)->index;
$data = ((struct msr_data*)@msr_data)->data;
if($index == MSR_KVM_WALL_CLOCK_NEW)
{
printf("[%s]kvm_set_msr_common msr:0x%08x data:%lu\n",strftime("%H:%M:%S",nsecs),$index,$data);
}
}
END
{
clear(@msr_data);
}
在虚拟机启动和重启的时候会有打印日志;因为data是从guest端传过来的gpa(客户端物理地址)因此无法直接获取相应数据结构的值。
2.1.2.2 host端kvm_clock的处理
在host端kvm代码里接受guest因为写入msr地址MSR_KVM_WALL_CLOCK_NEW而产生的vm exit并执行写入msr的回调函数,对应的调用逻辑是handle_wrmsr->kvm_set_msr->vmx_set_msr->kvm_set_msr_common (case MSR_KVM_WALL_CLOCK_NEW)
具体涉及到的代码内容为:(该代码在host端的kvm模块中)
cpp
int kvm_set_msr_common(struct kvm_vcpu *vcpu, struct msr_data *msr_info)
{
u64 data = msr_info->data;
...
case MSR_KVM_WALL_CLOCK_NEW:
if (!guest_pv_has(vcpu, KVM_FEATURE_CLOCKSOURCE2))
return 1;
/* 把guest端的gpa复制到arch.wall_clock成员中 */
vcpu->kvm->arch.wall_clock = data;
kvm_write_wall_clock(vcpu->kvm, data, 0);
break;
...
}
/* 具体的组装wall_clock的逻辑 */
static void kvm_write_wall_clock(struct kvm *kvm, gpa_t wall_clock, int sec_hi_ofs)
{
int version;
int r;
struct pvclock_wall_clock wc;
u32 wc_sec_hi;
u64 wall_nsec;
if (!wall_clock)
return;
r = kvm_read_guest(kvm, wall_clock, &version, sizeof(version));
if (r)
return;
if (version & 1)
++version; /* first time write, random junk */
++version;
if (kvm_write_guest(kvm, wall_clock, &version, sizeof(version)))
return;
/*
* ktime_get_real_ns()返回值是现在host时间拒1970-01-01.00:00:00的ns数,比如host现在时
* 间为2026-8-27左右其返回值为1787796713940544512ns也就是56年7个月26天2小时11分跟
* 在时间差不多。get_kvmclock_ns返回值为8190378453ns约8秒,因为qemu中设置的rtc base=20* 26-08-27T 02:11:44 系统启动时间为2026-08-27 10:11:52因为qemu有8小时的时间差所以基本
* 能对上。因此wall_nsec就是当时host的时间表示与guest的时间表示的间隔,不是误差。本质就
* 是qemu虚拟机系统启动时对应的host系统时间。
*/
wall_nsec = ktime_get_real_ns() - get_kvmclock_ns(kvm);
wc.nsec = do_div(wall_nsec, 1000000000);
wc.sec = (u32)wall_nsec; /* overflow in 2106 guest time */
wc.version = version;
kvm_write_guest(kvm, wall_clock, &wc, sizeof(wc));
if (sec_hi_ofs) {
wc_sec_hi = wall_nsec >> 32;
kvm_write_guest(kvm, wall_clock + sec_hi_ofs,
&wc_sec_hi, sizeof(wc_sec_hi));
}
version++;
kvm_write_guest(kvm, wall_clock, &version, sizeof(version));
}
/**
* ktime_get_real_ns - Get the current real/wall time in nanoseconds
*
* Returns: current real time converted to nanoseconds
*/
static inline u64 ktime_get_real_ns(void)
{
return ktime_to_ns(ktime_get_real());
}
//根据gpa地址和偏移在memslot中查找对应的hva地址
int kvm_write_guest(struct kvm *kvm, gpa_t gpa, const void *data,
unsigned long len)
{
gfn_t gfn = gpa >> PAGE_SHIFT;
int seg;
int offset = offset_in_page(gpa);
int ret;
while ((seg = next_segment(len, offset)) != 0) {
//查找memslot,并根据地址映射找到对应的hva然后进行写入
ret = kvm_write_guest_page(kvm, gfn, data, offset, seg);
if (ret < 0)
return ret;
offset = 0;
len -= seg;
data += seg;
++gfn;
}
return 0;
}
EXPORT_SYMBOL_GPL(kvm_write_guest);
我们要抓住思路主线,现在是guest端通过写入smr寄存器获取到wall_clock,根据代码这个值是kvm_write_wall_clock中的wc值的内容,写入的方式是kvm_write_guest,重点函数是ktime_get_real_ns() - get_kvmclock_ns(kvm)和kvm_write_guest,一个是用于理解wall_clock内容的含义,即往wall_clock写入的是什么;一个是表示如何把时间的值写入到guest的内存中。kvm_write_wall_clock函数实现了对wall_clock的赋值,并通过kvm_write_guest函数把值写入到gpa对应的hva中。获取的wall_clock值的关键变量是wall_nsec,其赋值函数是ktime_get_real_ns()和get_kvmclock_ns(kvm),ktime_get_real_ns和get_kvmclock_ns是需要重点分析的函数,通过调试打印这两个函数的返回值可以基本推测ktime_get_real_ns可粗略的认为是host端此时的时间拒1970-01-01的纳秒数,get_kvmclock_ns是可粗略的认为是kvm_clock初始化时的时间到rtc base的差,调试的记录在注释中已经标注。wall_nsec的含就是这两个差值的差值,把这个差值写入到wall_clock中并传递给guest使用。而且这个写入操作根据现在的测试只有虚拟机启动的时候执行一次。第二点就是guest wall_clock的gpa保存在kvm->arch.wall_clock。
还有一个寄存器是MSR_KVM_SYSTEM_TIME_NEW这个smr寄存器也会触发虚拟机退出,回忆一下guest内部的代码:
cpp
static void kvm_register_clock(char *txt)
{
struct pvclock_vsyscall_time_info *src = this_cpu_hvclock();
u64 pa;
if (!src)
return;
pa = slow_virt_to_phys(&src->pvti) | 0x01ULL;
wrmsrl(msr_kvm_system_time, pa);
pr_debug("kvm-clock: cpu %d, msr %llx, %s", smp_processor_id(), pa, txt);
}
host端因为写入寄存器发生vm-exit行为,并调用相应的回调函数。
cpp
int kvm_set_msr_common(struct kvm_vcpu *vcpu, struct msr_data *msr_info)
{
...
case MSR_KVM_SYSTEM_TIME_NEW:
if (!guest_pv_has(vcpu, KVM_FEATURE_CLOCKSOURCE2))
return 1;
kvm_write_system_time(vcpu, data, false, msr_info->host_initiated);
break;
...
}
/** system_time是客户端写smr的src->pvti的gpa 这点很重要,就是说guest写入smr的都是gpa
* 然后其次就是谁的gpa这样才能好理解逻辑
*/
static void kvm_write_system_time(struct kvm_vcpu *vcpu, gpa_t system_time,
bool old_msr, bool host_initiated)
{
struct kvm_arch *ka = &vcpu->kvm->arch;
if (vcpu->vcpu_id == 0 && !host_initiated) {
if (ka->boot_vcpu_runs_old_kvmclock != old_msr)
kvm_make_request(KVM_REQ_MASTERCLOCK_UPDATE, vcpu);
ka->boot_vcpu_runs_old_kvmclock = old_msr;
}
vcpu->arch.time = system_time;
kvm_make_request(KVM_REQ_GLOBAL_CLOCK_UPDATE, vcpu);
/* we verify if the enable bit is set... */
if (system_time & 1)
/* 把system-time即传过来的src->pvti gpa跟arch.pv_time相关联 */
kvm_gpc_activate (&vcpu->arch.pv_time, system_time & ~1ULL,
sizeof(struct pvclock_vcpu_time_info));
else
kvm_gpc_deactivate(&vcpu->arch.pv_time);
return;
}
这里面做的三件事:1 把pvti的gpa保存在vcpu->arch.time中,区分与wall_clock保存在kvm->arch.wall_clock中,因为pvti是每个cpu都有一份,而wall_clock只有一份。2,提交了vcpu的request这个我理解是vcpu的任务机制,当发生vm-exit时根据时间和任务要计划在vm-enter时要做的时,用request请求任务实现。3转换地址把gpa转化为hva并保存在vcpu->arch.pv_time,实现函数是kvm_gpc_activate
cpp
int kvm_gpc_activate(struct gfn_to_pfn_cache *gpc, gpa_t gpa, unsigned long len)
{
/*
* Explicitly disallow INVALID_GPA so that the magic value can be used
* by KVM to differentiate between GPA-based and HVA-based caches.
*/
if (WARN_ON_ONCE(kvm_is_error_gpa(gpa)))
return -EINVAL;
return __kvm_gpc_activate(gpc, gpa, KVM_HVA_ERR_BAD, len);
}
static int __kvm_gpc_activate(struct gfn_to_pfn_cache *gpc, gpa_t gpa, unsigned long uhva,
unsigned long len)
{
struct kvm *kvm = gpc->kvm;
if (!kvm_gpc_is_valid_len(gpa, uhva, len))
return -EINVAL;
guard(mutex)(&gpc->refresh_lock);
if (!gpc->active) {
if (KVM_BUG_ON(gpc->valid, kvm))
return -EIO;
spin_lock(&kvm->gpc_lock);
list_add(&gpc->list, &kvm->gpc_list);
spin_unlock(&kvm->gpc_lock);
/*
* Activate the cache after adding it to the list, a concurrent
* refresh must not establish a mapping until the cache is
* reachable by mmu_notifier events.
*/
write_lock_irq(&gpc->lock);
gpc->active = true;
write_unlock_irq(&gpc->lock);
}
return __kvm_gpc_refresh(gpc, gpa, uhva);
}
static int __kvm_gpc_refresh(struct gfn_to_pfn_cache *gpc, gpa_t gpa, unsigned long uhva)
{
unsigned long page_offset;
bool unmap_old = false;
unsigned long old_uhva;
kvm_pfn_t old_pfn;
bool hva_change = false;
void *old_khva;
int ret;
/* Either gpa or uhva must be valid, but not both */
if (WARN_ON_ONCE(kvm_is_error_gpa(gpa) == kvm_is_error_hva(uhva)))
return -EINVAL;
lockdep_assert_held(&gpc->refresh_lock);
write_lock_irq(&gpc->lock);
if (!gpc->active) {
ret = -EINVAL;
goto out_unlock;
}
old_pfn = gpc->pfn;
old_khva = (void *)PAGE_ALIGN_DOWN((uintptr_t)gpc->khva);
old_uhva = PAGE_ALIGN_DOWN(gpc->uhva);
if (kvm_is_error_gpa(gpa)) {
page_offset = offset_in_page(uhva);
gpc->gpa = INVALID_GPA;
gpc->memslot = NULL;
gpc->uhva = PAGE_ALIGN_DOWN(uhva);
if (gpc->uhva != old_uhva)
hva_change = true;
} else {
struct kvm_memslots *slots = kvm_memslots(gpc->kvm);
page_offset = offset_in_page(gpa);
if (gpc->gpa != gpa || gpc->generation != slots->generation ||
kvm_is_error_hva(gpc->uhva)) {
gfn_t gfn = gpa_to_gfn(gpa);
/*
* 实际就是想知道guest穿过来的gpa是多少
* 对应的 hva是多少,对应那个memslot
* 目的就是好访问
*/
gpc->gpa = gpa;
gpc->generation = slots->generation;
gpc->memslot = __gfn_to_memslot(slots, gfn);
gpc->uhva = gfn_to_hva_memslot(gpc->memslot, gfn);
if (kvm_is_error_hva(gpc->uhva)) {
ret = -EFAULT;
goto out;
}
/*
* Even if the GPA and/or the memslot generation changed, the
* HVA may still be the same.
*/
if (gpc->uhva != old_uhva)
hva_change = true;
} else {
gpc->uhva = old_uhva;
}
}
/* Note: the offset must be correct before calling hva_to_pfn_retry() */
gpc->uhva += page_offset;
/*
* If the userspace HVA changed or the PFN was already invalid,
* drop the lock and do the HVA to PFN lookup again.
*/
if (!gpc->valid || hva_change) {
ret = hva_to_pfn_retry(gpc);
} else {
/*
* If the HVA鈫扨FN mapping was already valid, don't unmap it.
* But do update gpc->khva because the offset within the page
* may have changed.
*/
gpc->khva = old_khva + page_offset;
ret = 0;
goto out_unlock;
}
out:
/*
* Invalidate the cache and purge the pfn/khva if the refresh failed.
* Some/all of the uhva, gpa, and memslot generation info may still be
* valid, leave it as is.
*/
if (ret) {
gpc->valid = false;
gpc->pfn = KVM_PFN_ERR_FAULT;
gpc->khva = NULL;
}
/* Detect a pfn change before dropping the lock! */
unmap_old = (old_pfn != gpc->pfn);
out_unlock:
write_unlock_irq(&gpc->lock);
if (unmap_old)
gpc_unmap(old_pfn, old_khva);
return ret;
}
根据以上函数梳理kvm_gpc_activate函数的关键逻辑就是根据kvm memslot把gpa转化为hva并记录相关信息到vcpu->arch.pv_time中。目的是为了后面的使用方便。
接着就是vm-enter时对request请求处理,共有两个请求无论什么情况最后都要走到KVM_REQ_CLOCK_UPDATE请求的逻辑里
cpp
static int vcpu_enter_guest(struct kvm_vcpu *vcpu)
{
...
if (kvm_check_request(KVM_REQ_MASTERCLOCK_UPDATE, vcpu))
kvm_update_masterclock(vcpu->kvm);
if (kvm_check_request(KVM_REQ_GLOBAL_CLOCK_UPDATE, vcpu))
kvm_gen_kvmclock_update(vcpu);
if (kvm_check_request(KVM_REQ_CLOCK_UPDATE, vcpu)) {
r = kvm_guest_time_update(vcpu);
if (unlikely(r))
goto out;
}
...
}
----重申强调重点---
pv_time是guest端src->picv对应的gha与hva之间的对应关系为了好找src->pvti而做的变量
------------------
static int kvm_guest_time_update(struct kvm_vcpu *v)
{
unsigned long flags, tgt_tsc_khz;
unsigned seq;
struct kvm_vcpu_arch *vcpu = &v->arch;
struct kvm_arch *ka = &v->kvm->arch;
s64 kernel_ns;
u64 tsc_timestamp, host_tsc;
u8 pvclock_flags;
bool use_master_clock;
kernel_ns = 0;
host_tsc = 0;
/*
* If the host uses TSC clock, then passthrough TSC as stable
* to the guest.
*/
do {
seq = read_seqcount_begin(&ka->pvclock_sc);
use_master_clock = ka->use_master_clock;
if (use_master_clock) {
host_tsc = ka->master_cycle_now;
kernel_ns = ka->master_kernel_ns;
}
} while (read_seqcount_retry(&ka->pvclock_sc, seq));
/* Keep irq disabled to prevent changes to the clock */
local_irq_save(flags);
tgt_tsc_khz = get_cpu_tsc_khz();
if (unlikely(tgt_tsc_khz == 0)) {
local_irq_restore(flags);
kvm_make_request(KVM_REQ_CLOCK_UPDATE, v);
return 1;
}
if (!use_master_clock) {
host_tsc = rdtsc();
kernel_ns = get_kvmclock_base_ns();
}
tsc_timestamp = kvm_read_l1_tsc(v, host_tsc);
/*
* We may have to catch up the TSC to match elapsed wall clock
* time for two reasons, even if kvmclock is used.
* 1) CPU could have been running below the maximum TSC rate
* 2) Broken TSC compensation resets the base at each VCPU
* entry to avoid unknown leaps of TSC even when running
* again on the same CPU. This may cause apparent elapsed
* time to disappear, and the guest to stand still or run
* very slowly.
*/
if (vcpu->tsc_catchup) {
u64 tsc = compute_guest_tsc(v, kernel_ns);
if (tsc > tsc_timestamp) {
adjust_tsc_offset_guest(v, tsc - tsc_timestamp);
tsc_timestamp = tsc;
}
}
local_irq_restore(flags);
/* With all the info we got, fill in the values */
if (kvm_caps.has_tsc_control)
tgt_tsc_khz = kvm_scale_tsc(tgt_tsc_khz,
v->arch.l1_tsc_scaling_ratio);
if (unlikely(vcpu->hw_tsc_khz != tgt_tsc_khz)) {
kvm_get_time_scale(NSEC_PER_SEC, tgt_tsc_khz * 1000LL,
&vcpu->hv_clock.tsc_shift,
&vcpu->hv_clock.tsc_to_system_mul);
vcpu->hw_tsc_khz = tgt_tsc_khz;
kvm_xen_update_tsc_info(v);
}
vcpu->hv_clock.tsc_timestamp = tsc_timestamp;
vcpu->hv_clock.system_time = kernel_ns + v->kvm->arch.kvmclock_offset;
vcpu->last_guest_tsc = tsc_timestamp;
/* If the host uses TSC clocksource, then it is stable */
pvclock_flags = 0;
if (use_master_clock)
pvclock_flags |= PVCLOCK_TSC_STABLE_BIT;
vcpu->hv_clock.flags = pvclock_flags;
if (vcpu->pv_time.active)
kvm_setup_guest_pvclock(v, &vcpu->pv_time, 0, false);
kvm_hv_setup_tsc_page(v->kvm, &vcpu->hv_clock);
return 0;
}
static void kvm_setup_guest_pvclock(struct kvm_vcpu *v,
struct gfn_to_pfn_cache *gpc,
unsigned int offset,
bool force_tsc_unstable)
{
struct kvm_vcpu_arch *vcpu = &v->arch;
struct pvclock_vcpu_time_info *guest_hv_clock;
unsigned long flags;
read_lock_irqsave(&gpc->lock, flags);
while (!kvm_gpc_check(gpc, offset + sizeof(*guest_hv_clock))) {
read_unlock_irqrestore(&gpc->lock, flags);
if (kvm_gpc_refresh(gpc, offset + sizeof(*guest_hv_clock)))
return;
read_lock_irqsave(&gpc->lock, flags);
}
guest_hv_clock = (void *)(gpc->khva + offset);
/*
* This VCPU is paused, but it's legal for a guest to read another
* VCPU's kvmclock, so we really have to follow the specification where
* it says that version is odd if data is being modified, and even after
* it is consistent.
*/
guest_hv_clock->version = vcpu->hv_clock.version = (guest_hv_clock->version + 1) | 1;
smp_wmb();
/* retain PVCLOCK_GUEST_STOPPED if set in guest copy */
vcpu->hv_clock.flags |= (guest_hv_clock->flags & PVCLOCK_GUEST_STOPPED);
if (vcpu->pvclock_set_guest_stopped_request) {
vcpu->hv_clock.flags |= PVCLOCK_GUEST_STOPPED;
vcpu->pvclock_set_guest_stopped_request = false;
}
memcpy(guest_hv_clock, &vcpu->hv_clock, sizeof(*guest_hv_clock));
if (force_tsc_unstable)
guest_hv_clock->flags &= ~PVCLOCK_TSC_STABLE_BIT;
smp_wmb();
guest_hv_clock->version = ++vcpu->hv_clock.version;
/* 还不忘标记一下位图 */
kvm_gpc_mark_dirty_in_slot(gpc);
read_unlock_irqrestore(&gpc->lock, flags);
trace_kvm_pvclock_update(v->vcpu_id, &vcpu->hv_clock);
}
这里面主要关注一下数据是如何写入到指定地址的kvm_setup_guest_pvclock函数通过memcpy把vcpu->hv_clock里的内容拷贝到priv的hva地址里,这样两个关键的数据就从host端传输到了guest指定的地址里。
到了这里我们已经从代码端,梳理出了guest如何和跟hsot交互的,以及关键的两个变量wall_clock和per_cpu的pvti变量如何获取以及如何传给guest端的并且大致了解了这两个变量值的大致来源。
2.1.2.3 guest用户程序的获取
这时候guest系统对时钟的模拟已基本完成,guest系统的用户态程序通过指定的函数获取相应的值我们以clock_gettime(CLOCK_MONOTONIC,&ts)为例来说明,如何获取指定时间的。现在获取时间的方式一般是通过vDSO(virtual dynamic shared objec)的方式获取时间值,优点是不会触发系统调用只是内存的读取是一种很高效的读取时间的方式,其实现的基本原理是我们在编译好程序时会自动连接linux-vdos.so.1库,并执行里面的代码,这个其实是内核编译好的库,以供应用程序使用的。对程序的调试相关调用堆栈如下:
如果虚拟机使用了tsc时钟
bash
__arch_get_hw_counter (vd=0x7ffff7ff6080, clock_mode=1) at ./arch/x86/include/asm/vdso/gettimeofday.h:235
235 if (likely(clock_mode == VDSO_CLOCKMODE_TSC))
(gdb) n
236 return (u64)rdtsc_ordered();
(gdb) bt
#0 __arch_get_hw_counter (vd=0x7ffff7ff6080, clock_mode=1) at ./arch/x86/include/asm/vdso/gettimeofday.h:236
#1 do_hres (ts=0x7fffffffd990, clk=1, vd=0x7ffff7ff6080)
at arch/x86/entry/vdso/../../../../lib/vdso/gettimeofday.c:142
#2 __vdso_clock_gettime_common (vd=0x7ffff7ff6080, ts=0x7fffffffd990, clock=1)
at arch/x86/entry/vdso/../../../../lib/vdso/gettimeofday.c:249
#3 __vdso_clock_gettime_data (ts=0x7fffffffd990, clock=1, vd=<optimized out>)
at arch/x86/entry/vdso/../../../../lib/vdso/gettimeofday.c:256
#4 __vdso_clock_gettime (ts=0x7fffffffd990, clock=1)
at arch/x86/entry/vdso/../../../../lib/vdso/gettimeofday.c:266
#5 _vdso_clock_gettime (clock=1, ts=0x7fffffffd990) at arch/x86/entry/vdso/vclock_gettime.c:43
#6 000007ffff7af52ca in clock_gettime@GLIBC_2.2.5 () from /lib64/libc.so.6
#7 0x00000000004005d9 in main () at clock.c:10
如果是虚拟机使用kvm-clock时钟
bash
#0 vread_pvclock () at ./arch/x86/include/asm/vdso/gettimeofday.h:210
#1 __arch_get_hw_counter (clock_mode=2, vd=<optimized out>) at ./arch/x86/include/asm/vdso/gettimeofday.h:246
#2 __arch_get_hw_counter (vd=0x7ffff7ff6080)
clock_mode=<error reading variable: Cannot access memory at address 0x7ffff7ff6084>
at ./arch/x86/include/asm/vdso/gettimeofday.h:232
#3 do_hres (ts=0x7fffffffd990, clk=1, vd=0x7ffff7ff6080)
at arch/x86/entry/vdso/../../../../lib/vdso/gettimeofday.c:142
#4 __vdso_clock_gettime_common (vd=0x7ffff7ff6080, ts=0x7fffffffd990, clock=1)
at arch/x86/entry/vdso/../../../../lib/vdso/gettimeofday.c:249
#5 __vdso_clock_gettime_data (ts=0x7fffffffd990, clock=1, vd=<optimized out>)
at arch/x86/entry/vdso/../../../../lib/vdso/gettimeofday.c:256
#6 __vdso_clock_gettime (ts=0x7fffffffd990, clock=1)
at arch/x86/entry/vdso/../../../../lib/vdso/gettimeofday.c:266
#7 _vdso_clock_gettime (clock=<optimized out>, ts=0x7fffffffd990) at arch/x86/entry/vdso/vclock_gettime.c:43
#8 0x00007ffff7af52ca in clock_gettime@GLIBC_2.2.5 () from /lib64/libc.so.6
#9 0x00000000004005d9 in main () at clock.c:10
根据堆栈我们可以看到,不同的使用源最后调用的获取时间的接口是不一样的并且函数指令是在vdso库中我们看看其对应的具体函数:
cpp
static u64 vread_pvclock(void)
{
const struct pvclock_vcpu_time_info *pvti = &pvclock_page.pvti;
u32 version;
u64 ret;
/*
* Note: The kernel and hypervisor must guarantee that cpu ID
* number maps 1:1 to per-CPU pvclock time info.
*
* Because the hypervisor is entirely unaware of guest userspace
* preemption, it cannot guarantee that per-CPU pvclock time
* info is updated if the underlying CPU changes or that that
* version is increased whenever underlying CPU changes.
*
* On KVM, we are guaranteed that pvti updates for any vCPU are
* atomic as seen by *all* vCPUs. This is an even stronger
* guarantee than we get with a normal seqlock.
*
* On Xen, we don't appear to have that guarantee, but Xen still
* supplies a valid seqlock using the version field.
*
* We only do pvclock vdso timing at all if
* PVCLOCK_TSC_STABLE_BIT is set, and we interpret that bit to
* mean that all vCPUs have matching pvti and that the TSC is
* synced, so we can just look at vCPU 0's pvti.
*/
do {
version = pvclock_read_begin(pvti);
if (unlikely(!(pvti->flags & PVCLOCK_TSC_STABLE_BIT)))
return U64_MAX;
ret = __pvclock_read_cycles(pvti, rdtsc_ordered());
} while (pvclock_read_retry(pvti, version));
return ret & S64_MAX;
}
static __always_inline
u64 __pvclock_read_cycles(const struct pvclock_vcpu_time_info *src, u64 tsc)
{
u64 delta = tsc - src->tsc_timestamp;
u64 offset = pvclock_scale_delta(delta, src->tsc_to_system_mul,
src->tsc_shift);
return src->system_time + offset;
}
基本就是一些pvti和rdtsc_ordered的算数运算;
而对于tsc时钟源就是用户态执行
cpp
static __always_inline unsigned long long rdtsc_ordered(void)
{
DECLARE_ARGS(val, low, high);
/*
* The RDTSC instruction is not ordered relative to memory
* access. The Intel SDM and the AMD APM are both vague on this
* point, but empirically an RDTSC instruction can be
* speculatively executed before prior loads. An RDTSC
* immediately after an appropriate barrier appears to be
* ordered as a normal load, that is, it provides the same
* ordering guarantees as reading from a global memory location
* that some other imaginary CPU is updating continuously with a
* time stamp.
*
* Thus, use the preferred barrier on the respective CPU, aiming for
* RDTSCP as the default.
*/
asm volatile(ALTERNATIVE_2("rdtsc",
"lfence; rdtsc", X86_FEATURE_LFENCE_RDTSC,
"rdtscp", X86_FEATURE_RDTSCP)
: EAX_EDX_RET(val, low, high)
/* RDTSCP clobbers ECX with MSR_TSC_AUX. */
:: "ecx");
return EAX_EDX_VAL(val, low, high);
}
对比两种时钟模式,tsc时钟是只是读取寄存器信息,kvm-clock时钟读取的是内存数据加寄存器数据,其次在读取内存数据的时候也是有锁的操作,因此kvm-clock时钟再获取时间的速度上是低于tsc时钟的。
到了这里我们基本上已经了解到了数据传递的流程从host到guest-kernel,从guest-kernel到guest-user。但是还有几个问题要解决,(1)wall_clock的值到底值什么?如何计算的,(2)pvti的值是什么又是如何计算的?(3)guest执行clock_gettime时用户态是如何直接获取内核态的数据的。
再把vDSO的机制看一下,实际就是自定义的mmaping函数:
cpp
static int vdso_mremap(const struct vm_special_mapping *sm,
struct vm_area_struct *new_vma)
{
const struct vdso_image *image = current->mm->context.vdso_image;
vdso_fix_landing(image, new_vma);
current->mm->context.vdso = (void __user *)new_vma->vm_start;
return 0;
}
#ifdef CONFIG_TIME_NS
/*
* The vvar page layout depends on whether a task belongs to the root or
* non-root time namespace. Whenever a task changes its namespace, the VVAR
* page tables are cleared and then they will re-faulted with a
* corresponding layout.
* See also the comment near timens_setup_vdso_data() for details.
*/
int vdso_join_timens(struct task_struct *task, struct time_namespace *ns)
{
struct mm_struct *mm = task->mm;
struct vm_area_struct *vma;
VMA_ITERATOR(vmi, mm, 0);
mmap_read_lock(mm);
for_each_vma(vmi, vma) {
if (vma_is_special_mapping(vma, &vvar_mapping))
zap_vma_pages(vma);
}
mmap_read_unlock(mm);
return 0;
}
#endif
static vm_fault_t vvar_fault(const struct vm_special_mapping *sm,
struct vm_area_struct *vma, struct vm_fault *vmf)
{
const struct vdso_image *image = vma->vm_mm->context.vdso_image;
unsigned long pfn;
long sym_offset;
if (!image)
return VM_FAULT_SIGBUS;
sym_offset = (long)(vmf->pgoff << PAGE_SHIFT) +
image->sym_vvar_start;
/*
* Sanity check: a symbol offset of zero means that the page
* does not exist for this vdso image, not that the page is at
* offset zero relative to the text mapping. This should be
* impossible here, because sym_offset should only be zero for
* the page past the end of the vvar mapping.
*/
if (sym_offset == 0)
return VM_FAULT_SIGBUS;
if (sym_offset == image->sym_vvar_page) {
struct page *timens_page = find_timens_vvar_page(vma);
pfn = __pa_symbol(&__vvar_page) >> PAGE_SHIFT;
/*
* If a task belongs to a time namespace then a namespace
* specific VVAR is mapped with the sym_vvar_page offset and
* the real VVAR page is mapped with the sym_timens_page
* offset.
* See also the comment near timens_setup_vdso_data().
*/
if (timens_page) {
unsigned long addr;
vm_fault_t err;
/*
* Optimization: inside time namespace pre-fault
* VVAR page too. As on timens page there are only
* offsets for clocks on VVAR, it'll be faulted
* shortly by VDSO code.
*/
addr = vmf->address + (image->sym_timens_page - sym_offset);
err = vmf_insert_pfn(vma, addr, pfn);
if (unlikely(err & VM_FAULT_ERROR))
return err;
pfn = page_to_pfn(timens_page);
}
return vmf_insert_pfn(vma, vmf->address, pfn);
} else if (sym_offset == image->sym_pvclock_page) {
struct pvclock_vsyscall_time_info *pvti =
pvclock_get_pvti_cpu0_va();
if (pvti && vclock_was_used(VDSO_CLOCKMODE_PVCLOCK)) {
return vmf_insert_pfn_prot(vma, vmf->address,
__pa(pvti) >> PAGE_SHIFT,
pgprot_decrypted(vma->vm_page_prot));
}
} else if (sym_offset == image->sym_hvclock_page) {
pfn = hv_get_tsc_pfn();
if (pfn && vclock_was_used(VDSO_CLOCKMODE_HVCLOCK))
return vmf_insert_pfn(vma, vmf->address, pfn);
} else if (sym_offset == image->sym_timens_page) {
struct page *timens_page = find_timens_vvar_page(vma);
if (!timens_page)
return VM_FAULT_SIGBUS;
pfn = __pa_symbol(&__vvar_page) >> PAGE_SHIFT;
return vmf_insert_pfn(vma, vmf->address, pfn);
}
return VM_FAULT_SIGBUS;
}
static const struct vm_special_mapping vdso_mapping = {
.name = "[vdso]",
.fault = vdso_fault,
.mremap = vdso_mremap,
};
static const struct vm_special_mapping vvar_mapping = {
.name = "[vvar]",
.fault = vvar_fault,
};
这是内核回应的mmaping函数的具体实现,找到pvti然后返回其对应的物理地址。
2.1.3 内核时钟模型
在我们读取kvm-clock代码过程中遇到了很多变量和参数这些大部分跟timekeeper数据有关,这是一个全局的时钟模型。kernel内获取的各种时间都是直接或者间接的要读取该数据中的值,还有下一个kvm->arch或者vcpu->arch中的时钟变量都有相应的意义。系统时钟里有几个比较核心的时钟概念:
realtime、monotonic、monotonic_raw、boottime
偏差:
wall_to_monotonic、offs_boot、offs_real。
realtime就是walltime墙上时间,monotonic是一个单调递增的时间但不包括系统休眠时间,boottime是系统启动后经历的时间
对应的内核获取函数为:
realtime -àktime_get_real/ktime_get_real_ts
Monotonic Time -à ktime_get/ktime_get_ts
boot time -à ktime_get_boottime/get_monotonic_boottime
Raw Monotonic Time -à ktime_get_raw_ns()/ktime_get_raw()
wall_to_monotonic : monotonic - realtime墙上时钟与单调时钟的差值
offs_real:monotonic → real 的偏移即 ktime_get_real() - ktime_get()
offs_boot : monotonic → boottime 的偏移即累计 suspend 时间
offs_raw: RAW 时间基于独立硬件计数器,不与 MONOTONIC 共享基准
monotonic_to_boot: boottime - monotonic等同于 offs_boot 的 timespec 形式
把握住一个重点:所有的时间都是跟monotonic相关,即offs_boot + monotonic是boottime,offs_real+monotnonic是realtime,记住一点monotonic是时钟源衡量现实时间流逝的一个度量,经过了多长时间再加上相应的offs_xxx就可以得到对应的时间。
首先对于人类来说,我们想要的时间一般是RTC(real time clock)时间,就是我想知道现在是几点几分了,这个就是realtime,但是对于程序来来说,我想知道系统运行了多长时间,即内核运行了多长时间,这个就是monotonic,那系统自启动走了多少时间?就是boottime,这里说的是时间就是对标现实的RTC时间,但是计算机有他的计时方式就是时钟频率计时,1kHz的cpu认为震荡了1kHz就是一秒,但是cpu的频率又是不稳定的,但是我们不管,即便你2秒钟震荡了1kHz我也认为是一秒,这个时间就是monotonic raw时间。
cpp
/* 内核时钟数据结构模型 */
static struct {
seqcount_raw_spinlock_t seq;
struct timekeeper timekeeper;
} tk_core ____cacheline_aligned = {
.seq = SEQCNT_RAW_SPINLOCK_ZERO(tk_core.seq, &timekeeper_lock),
};
/**
* struct timekeeper - Structure holding internal timekeeping values.
* @tkr_mono: The readout base structure for CLOCK_MONOTONIC
* @tkr_raw: The readout base structure for CLOCK_MONOTONIC_RAW
* @xtime_sec: Current CLOCK_REALTIME time in seconds
* @ktime_sec: Current CLOCK_MONOTONIC time in seconds
* @wall_to_monotonic: CLOCK_REALTIME to CLOCK_MONOTONIC offset
* @offs_real: Offset clock monotonic -> clock realtime
* @offs_boot: Offset clock monotonic -> clock boottime
* @offs_tai: Offset clock monotonic -> clock tai
* @tai_offset: The current UTC to TAI offset in seconds
* @clock_was_set_seq: The sequence number of clock was set events
* @cs_was_changed_seq: The sequence number of clocksource change events
* @next_leap_ktime: CLOCK_MONOTONIC time value of a pending leap-second
* @raw_sec: CLOCK_MONOTONIC_RAW time in seconds
* @monotonic_to_boot: CLOCK_MONOTONIC to CLOCK_BOOTTIME offset
* @cycle_interval: Number of clock cycles in one NTP interval
* @xtime_interval: Number of clock shifted nano seconds in one NTP
* interval.
* @xtime_remainder: Shifted nano seconds left over when rounding
* @cycle_interval
* @raw_interval: Shifted raw nano seconds accumulated per NTP interval.
* @ntp_error: Difference between accumulated time and NTP time in ntp
* shifted nano seconds.
* @ntp_error_shift: Shift conversion between clock shifted nano seconds and
* ntp shifted nano seconds.
*
* Note: For timespec(64) based interfaces wall_to_monotonic is what
* we need to add to xtime (or xtime corrected for sub jiffy times)
* to get to monotonic time. Monotonic is pegged at zero at system
* boot time, so wall_to_monotonic will be negative, however, we will
* ALWAYS keep the tv_nsec part positive so we can use the usual
* normalization.
*
* wall_to_monotonic is moved after resume from suspend for the
* monotonic time not to jump. We need to add total_sleep_time to
* wall_to_monotonic to get the real boot based time offset.
*
* wall_to_monotonic is no longer the boot time, getboottime must be
* used instead.
*
* @monotonic_to_boottime is a timespec64 representation of @offs_boot to
* accelerate the VDSO update for CLOCK_BOOTTIME.
* VDSO的全称为: virtual dynamic share object
*/
struct timekeeper {
struct tk_read_base tkr_mono;
struct tk_read_base tkr_raw;
u64 xtime_sec;
unsigned long ktime_sec;
struct timespec64 wall_to_monotonic;
ktime_t offs_real;
ktime_t offs_boot;
ktime_t offs_tai;
s32 tai_offset;
unsigned int clock_was_set_seq;
u8 cs_was_changed_seq;
ktime_t next_leap_ktime;
u64 raw_sec;
struct timespec64 monotonic_to_boot;
/* The following members are for timekeeping internal use */
u64 cycle_interval;
u64 xtime_interval;
s64 xtime_remainder;
u64 raw_interval;
/* The ntp_tick_length() value currently being used.
* This cached copy ensures we consistently apply the tick
* length for an entire tick, as ntp_tick_length may change
* mid-tick, and we don't want to apply that new value to
* the tick in progress.
*/
u64 ntp_tick;
/* Difference between accumulated time and NTP time in ntp
* shifted nano seconds. */
s64 ntp_error;
u32 ntp_error_shift;
u32 ntp_err_mult;
/* Flag used to avoid updating NTP twice with same second */
u32 skip_second_overflow;
};
以上是tiemkeeper的数据结构我们从系统上电开始梳理各种时间的初始化时间点并说明对应timekeeper的那个变量,假设我们默认使用的是tsc时钟。当我们安装系统时会有一个时间设置的值,这个就是RTC时间,这个会设置在RTC的设备中无论系统启动不启动或者上不上电时间都会走,这个时间就是realtime就是我们说的几分几秒,当系统上点后,tsc开始跳动,开始累加数值这个值我们设置为T0,这时候系统还没启动什么都没有初始化,计时还未开始;到了biso阶段设置为T1,bios假设耗了5秒,tsc值继续增加这时候代码还没有运行,这时候仍然没有任何时间概念,开始kernel_init初始化内核代码运行设为T2,这时候还是什么都没有,因为还没初始化时间;T3时刻内核读取墙上时间,获取RTC时间,这时候时钟还没初始化,tsc继续走;到了执行了timekeep_init函数开始初始化各种时间我们设为T4假设现在是
// timekeeping_init 的核心逻辑
//日期转换为 Unix 时间戳(距离 1970 年的秒数),假设等于 1,790,000,000 秒
// 1. 读取刚才从 RTC 获取的墙上时间 (Unix 秒数)
struct timespec64 wall_time = { .tv_sec = 1790000000, .tv_nsec = 0 };
// 2. 赋值 xtime_sec (REALTIME 的整秒数)
tk->xtime_sec = wall_time.tv_sec;
// 【概念明确】xtime_sec 就是系统开机那一瞬间的 UTC 秒数。
// 3. 初始化 MONOTONIC 的基准 (tkr_mono.base)
// 因为 MONOTONIC 是从"系统启动"开始算的,所以开机这一瞬间,MONOTONIC = 0!
tk->tkr_mono.base = 0;
// 4. 计算并赋值 offs_real (最核心的偏移量!)
// 公式:offs_real = REALTIME - MONOTONIC
// 此时 REALTIME = 1790000000秒, MONOTONIC = 0
tk->offs_real = 1790000000 * 10^9; // 转换为纳秒
// 【概念明确】offs_real 的物理含义就是:系统开机那一刻,距离 1970 年过去了多少纳秒。
// 它是一个"常数"(除非用户手动改时间),把 MONOTONIC 平移到 REALTIME。
// 5. 赋值 wall_to_monotonic
// 公式:wall_to_monotonic = MONOTONIC - REALTIME = -offs_real
tk->wall_to_monotonic = -1790000000秒;
// 【概念明确】这是一个负数。含义是:从当前的 REALTIME 往回倒退多少时间,才能到达 MONOTONIC 的起点(0点)。
// 6. 初始化 BOOTTIME
tk->boottime = tk->offs_real; // 刚开机没休眠过,boot时间等于real偏移
t5时刻时钟接管,早期内核可能用 jiffies 或 acpi_pm 计时,现在内核发现 CPU 支持 invariant TSC(恒定频率 TSC),决定切换到 TSC,到了这里基本上就已经把时间都切换完了,TSC 此时的值:假设从开机(T0)到现在(T5)过去了 8 秒,TSC_now = 24,000,000,000 个 cycle。
内核动作:
- 调用 clocksource_register_hz("tsc", freq) 注册 TSC。
- 内核计算出 mult 和 shift(用于把 TSC cycle 转换为纳秒的魔法数字)。
- 记录 baseline:tk->tkr_mono.cycle_last = TSC_now。
注意:此时 tk->tkr_mono.base 依然是 0!就是说mono为0的时候我们就记录一下当时的TSC值,换言之记录一下当我们开始计时系统时间时当时的tsc值,这样我们才能计算,往后的任意时间到开机跑了多久。即monotonic过了多久,然后加上相应的offs_xx就可以得到相应的时间。
首先回答一个问题,为什么要用时钟源,直接读RTC不就行了?首先是TSC时钟精度更高,其次效率也更高,那么有了时钟源如何根据时钟源来计算各种时间那,首先时钟源能告诉我们的是过了多长时间,这个时间段加到那种时间的offs_xxx上,就是那个时间的累加,我们用一个例子来说明相应的时间计算方式:
假设现在距离T5(内核接管 TSC)又过去了10秒,此时TSC增加了10000000000个cycle,通过mult/shift转换,这个cycle增量等于10000000000纳秒(10秒)。我们称这个增量为 delta_ns。
- 获取 CLOCK_MONOTONIC (单调时间)
物理含义:系统开机后,真正"干活"了多久(不含休眠)
来源:完全来自 TSC 的增量!与 RTC 无关。
u64 mono_ns = tk->tkr_mono.base + delta_ns;
(初始值 0) + (TSC 转换来的 10秒)
结果:10,000,000,000 ns (10秒)
概念明确:MONOTONIC 就是 TSC 增量的累加。它不知道今天是几号,它只知道"内核接管后,CPU 跑了多久"。
- 获取 CLOCK_REALTIME (墙上时间)
物理含义:现在是哪年哪月哪日几点几分
来源:MONOTONIC + 开机那一刻的 RTC 时间种子
u64 real_ns = mono_ns + tk->offs_real;
(10秒) + (开机时的 1790000000秒)
结果:1790000010 秒 (时间往后走了10秒)
概念明确:REALTIME 就是把 MONOTONIC 的起点,从"开机瞬间"平移到"1970年"。offs_real 就是那个平移的距离。
- 获取 CLOCK_BOOTTIME (含休眠的单调时间)
物理含义:系统开机后,经历了多久(包含休眠睡着的时间)
来源:MONOTONIC + 开机RTC时间 + 历次休眠的时长
u64 boot_ns = mono_ns + tk->offs_boot;
假设期间休眠了 5 秒,resume 时 offs_boot 增加了 5秒
结果:10秒 + ( 5 + 5秒)
概念明确:BOOTTIME 和 REALTIME 在刚开机时是一样的。但当你合上笔记本盖子(Suspend),MONOTONIC 暂停了,但真实世界还在流逝。内核在唤醒(Resume)时,会测量睡了多久,把这个时长加到 offs_boot 里。
- 获取 CLOCK_MONOTONIC_RAW (原始硬件时间)
物理含义:纯粹的硬件振荡器时间,不受 NTP 软件微调影响
来源:TSC 增量,但使用不同的 mult/shift
u64 raw_ns = tk->tkr_raw.base + delta_raw_ns;
概念明确:NTP 会通过微调 tkr_mono.mult 来让 MONOTONIC 走得快一点或慢一点,以对齐互联网标准时间。但 tkr_raw 的 mult 永远不变,它反映的是你主板晶振最真实的物理频率。
代码说明:
cpp
static ktime_t *offsets[TK_OFFS_MAX] = {
[TK_OFFS_REAL] = &tk_core.timekeeper.offs_real,
[TK_OFFS_BOOT] = &tk_core.timekeeper.offs_boot,
[TK_OFFS_TAI] = &tk_core.timekeeper.offs_tai,
};
ktime_t ktime_get_with_offset(enum tk_offsets offs)
{
struct timekeeper *tk = &tk_core.timekeeper;
unsigned int seq;
ktime_t base, *offset = offsets[offs];
u64 nsecs;
WARN_ON(timekeeping_suspended);
do {
seq = read_seqcount_begin(&tk_core.seq);
/* 上一次记录时的时间 */
base = ktime_add(tk->tkr_mono.base, *offset);
/* 通过时钟源获取过了多长时间*/
nsecs = timekeeping_get_ns(&tk->tkr_mono);
} while (read_seqcount_retry(&tk_core.seq, seq));
return ktime_add_ns(base, nsecs);
}
static __always_inline u64 timekeeping_get_ns(const struct tk_read_base *tkr)
{
return timekeeping_cycles_to_ns (tkr, tk_clock_read(tkr));
}
static inline u64 tk_clock_read(const struct tk_read_base *tkr)
{
struct clocksource *clock = READ_ONCE(tkr->clock);
return clock->read(clock);
}
/* 因为host使用的是tsc时钟虚拟化因此 read函数对应的是 read_tsc()*/
static u64 read_tsc(struct clocksource *cs)
{
return (u64)rdtsc_ordered();
}
static __always_inline unsigned long long rdtsc_ordered(void)
{
DECLARE_ARGS(val, low, high);
/*
* The RDTSC instruction is not ordered relative to memory
* access. The Intel SDM and the AMD APM are both vague on this
* point, but empirically an RDTSC instruction can be
* speculatively executed before prior loads. An RDTSC
* immediately after an appropriate barrier appears to be
* ordered as a normal load, that is, it provides the same
* ordering guarantees as reading from a global memory location
* that some other imaginary CPU is updating continuously with a
* time stamp.
*
* Thus, use the preferred barrier on the respective CPU, aiming for
* RDTSCP as the default.
*/
asm volatile(ALTERNATIVE_2("rdtsc",
"lfence; rdtsc", X86_FEATURE_LFENCE_RDTSC,
"rdtscp", X86_FEATURE_RDTSCP)
: EAX_EDX_RET(val, low, high)
/* RDTSCP clobbers ECX with MSR_TSC_AUX. */
:: "ecx");
return EAX_EDX_VAL(val, low, high);
}
/* 把读取到的tsc值转化为时间*/
static inline u64 timekeeping_cycles_to_ns(const struct tk_read_base *tkr, u64 cycles)
{
/* Calculate the delta since the last update_wall_time() */
/* 就是读取到的tsc值减去设置tk->tkr_mono.base时的tsc值就是从那个时候
* 过了多少tsc,然后根据tsc的值和频率算出过了积分几秒。
*/
u64 mask = tkr->mask, delta = (cycles - tkr->cycle_last) & mask;
/*
* This detects both negative motion and the case where the delta
* overflows the multiplication with tkr->mult.
*/
if (unlikely(delta > tkr->clock->max_cycles)) {
/*
* Handle clocksource inconsistency between CPUs to prevent
* time from going backwards by checking for the MSB of the
* mask being set in the delta.
*/
if (delta & ~(mask >> 1))
return tkr->xtime_nsec >> tkr->shift;
return delta_to_ns_safe(tkr, delta);
}
return ((delta * tkr->mult) + tkr->xtime_nsec) >> tkr->shift;
}
这里面有一个转化公式来着,tsc的值是如何转化为ns的:首先我们知道如果cpu的频率如果是1秒2kHz那么我们认为tsc增加设为delta为4000时这时候就是2s,就是4000/2000 = 2,如果是纳秒那就是4000*10^9/2000=2000000000ns,但是计算机是不使用除法的,把除法变成乘法就是乘法和位移,要成的数就是mult要位移的数就是shift,就变成4000*mult>>shift=2000000000ns,计算出mult和shift的值,其次mult有缩放功能,当频率发生漂移的时候就可以增加或减少mult的值,使计算的值跟rtc对齐。
最后我们就可以把函数ktime_get_with_offset函数的意义确定了,就是设置offs_xx的时候时钟源又过了多长时间然后加上,就是现在的时间。
有了这些认识我们再看一下timekeeping_init就显得眉清目秀了
cpp
void __init timekeeping_init(void)
{
struct timespec64 wall_time, boot_offset, wall_to_mono;
struct timekeeper *tk = &tk_core.timekeeper;
struct clocksource *clock;
unsigned long flags;
/* 获取realtime 和monotonic时间 */
read_persistent_wall_and_boot_offset(&wall_time, &boot_offset);
if (timespec64_valid_settod(&wall_time) &&
timespec64_to_ns(&wall_time) > 0) {
persistent_clock_exists = true;
} else if (timespec64_to_ns(&wall_time) != 0) {
pr_warn("Persistent clock returned invalid value");
wall_time = (struct timespec64){0};
}
if (timespec64_compare(&wall_time, &boot_offset) < 0)
boot_offset = (struct timespec64){0};
/*
* We want set wall_to_mono, so the following is true:
* wall time + wall_to_mono = boot time
*/
/* 计算wall_to_mono */
wall_to_mono = timespec64_sub(boot_offset, wall_time);
raw_spin_lock_irqsave(&timekeeper_lock, flags);
write_seqcount_begin(&tk_core.seq);
ntp_init();
clock = clocksource_default_clock();
if (clock->enable)
clock->enable(clock);
tk_setup_internals(tk, clock);
/* 初始化real_time */
tk_set_xtime(tk, &wall_time);
tk->raw_sec = 0;
/* 初始化wall_to_mono */
tk_set_wall_to_mono(tk, wall_to_mono);
/* 更新tk里的变量值 */
timekeeping_update(tk, TK_MIRROR | TK_CLOCK_WAS_SET);
write_seqcount_end(&tk_core.seq);
raw_spin_unlock_irqrestore(&timekeeper_lock, flags);
}
void __weak __init
read_persistent_wall_and_boot_offset(struct timespec64 *wall_time,
struct timespec64 *boot_offset)
{
/* 从rtc中读取时间赋值给wall_time */
read_persistent_clock64(wall_time);
/* 从clock_source中读取过了多长时间,
*比如tsc时钟初始化后过了多长时间,实际就是monotonic
*/
*boot_offset = ns_to_timespec64(local_clock());
}
/* 就是monotonic -- real_time */
wall_to_mono = timespec64_sub(boot_offset, wall_time);
static void tk_setup_internals(struct timekeeper *tk, struct clocksource *clock)
{
u64 interval;
u64 tmp, ntpinterval;
struct clocksource *old_clock;
/* 初始化使用的时钟以及此时的tsc值赋值给cycle_last */
++tk->cs_was_changed_seq;
old_clock = tk->tkr_mono.clock;
tk->tkr_mono.clock = clock;
tk->tkr_mono.mask = clock->mask;
tk->tkr_mono.cycle_last = tk_clock_read(&tk->tkr_mono);
tk->tkr_raw.clock = clock;
tk->tkr_raw.mask = clock->mask;
tk->tkr_raw.cycle_last = tk->tkr_mono.cycle_last;
/* 下面就是cycle往ns转化需要的一些参数*/
/* Do the ns -> cycle conversion first, using original mult */
tmp = NTP_INTERVAL_LENGTH;
tmp <<= clock->shift;
ntpinterval = tmp;
tmp += clock->mult/2;
do_div(tmp, clock->mult);
if (tmp == 0)
tmp = 1;
interval = (u64) tmp;
tk->cycle_interval = interval;
/* Go back from cycles -> shifted ns */
tk->xtime_interval = interval * clock->mult;
tk->xtime_remainder = ntpinterval - tk->xtime_interval;
tk->raw_interval = interval * clock->mult;
/* if changing clocks, convert xtime_nsec shift units */
if (old_clock) {
int shift_change = clock->shift - old_clock->shift;
if (shift_change < 0) {
tk->tkr_mono.xtime_nsec >>= -shift_change;
tk->tkr_raw.xtime_nsec >>= -shift_change;
} else {
tk->tkr_mono.xtime_nsec <<= shift_change;
tk->tkr_raw.xtime_nsec <<= shift_change;
}
}
/*
* tkr_mono.shift是会变得,在update_wall_time=>timekeeping_advance=>
* timekeeping_adjust=> timekeeping_apply_adjustment中被修改,
* 但是tkr_raw.shift是不会变的,
* 这就是monotonic与monotonic_raw时间的区别
*/
tk->tkr_mono.shift = clock->shift;
tk->tkr_raw.shift = clock->shift;
tk->ntp_error = 0;
tk->ntp_error_shift = NTP_SCALE_SHIFT - clock->shift;
tk->ntp_tick = ntpinterval << tk->ntp_error_shift;
/*
* The timekeeper keeps its own mult values for the currently
* active clocksource. These value will be adjusted via NTP
* to counteract clock drifting.
*/
tk->tkr_mono.mult = clock->mult;
tk->tkr_raw.mult = clock->mult;
tk->ntp_err_mult = 0;
tk->skip_second_overflow = 0;
}
static void tk_set_xtime(struct timekeeper *tk, const struct timespec64 *ts)
{
/* tv_sec就是获取的rct时间*/
tk->xtime_sec = ts->tv_sec;
/* tv_nsec获取的时候是0 */
tk->tkr_mono.xtime_nsec = (u64)ts->tv_nsec << tk->tkr_mono.shift;
}
/* wtm 就是 wall_to_mono */
static void tk_set_wall_to_mono(struct timekeeper *tk, struct timespec64 wtm)
{
struct timespec64 tmp;
/*
* Verify consistency of: offset_real = -wall_to_monotonic
* before modifying anything
*/
/*
* 可以简单理解就是赋值函数,这是后wall_to_monotonic还未初始化*/
set_normalized_timespec64(&tmp, -tk->wall_to_monotonic.tv_sec,
-tk->wall_to_monotonic.tv_nsec);
WARN_ON_ONCE(tk->offs_real != timespec64_to_ktime(tmp));
tk->wall_to_monotonic = wtm;
/* 这时候把-wall_to_mono赋值给tmp ,tmp就是real_time -- monotonic */
set_normalized_timespec64(&tmp, -wtm.tv_sec, -wtm.tv_nsec);
/* 这时候就对上了,offs_real就是real_time --monotonic,如果想获取real_time 就计算
* monotonic时间然后再加上offs_real
*/
tk->offs_real = timespec64_to_ktime(tmp);
tk->offs_tai = ktime_add(tk->offs_real, ktime_set(tk->tai_offset, 0));
}
/* 比较重要的时钟更新函数,这个函数会一直被调用*/
static void timekeeping_update(struct timekeeper *tk, unsigned int action)
{
if (action & TK_CLEAR_NTP) {
tk->ntp_error = 0;
ntp_clear();
}
/* 根据时钟再次更新tk*/
tk_update_leap_state(tk);
tk_update_ktime_data(tk);
/* 把tk里的信息跟新到vdso的内存中*/
update_vsyscall(tk);
update_pvclock_gtod(tk, action & TK_CLOCK_WAS_SET);
/* 更新monotonic base_real的值,每次monotonic的base更新都要更新这个
* 这样ktime_get_with_offset的公式才能保持正确.
*/
tk->tkr_mono.base_real = tk->tkr_mono.base + tk->offs_real;
/*更新两个群居变量,快速访问使用*/
update_fast_timekeeper(&tk->tkr_mono, &tk_fast_mono);
update_fast_timekeeper(&tk->tkr_raw, &tk_fast_raw);
if (action & TK_CLOCK_WAS_SET)
tk->clock_was_set_seq++;
/*
* The mirroring of the data to the shadow-timekeeper needs
* to happen last here to ensure we don't over-write the
* timekeeper structure on the next update with stale data
*/
if (action & TK_MIRROR)
memcpy(&shadow_timekeeper, &tk_core.timekeeper,
sizeof(tk_core.timekeeper));
}
static inline void tk_update_ktime_data(struct timekeeper *tk)
{
u64 seconds;
u32 nsec;
/*
* The xtime based monotonic readout is:
* nsec = (xtime_sec + wtm_sec) * 1e9 + wtm_nsec + now();
* The ktime based monotonic readout is:
* nsec = base_mono + now();
* ==> base_mono = (xtime_sec + wtm_sec) * 1e9 + wtm_nsec
*/
/* 已知xtime_sec是wall_time,那么seconds就是monotonic时间
* wall_to_mono = monotonic --wall_time
*/
seconds = (u64)(tk->xtime_sec + tk->wall_to_monotonic.tv_sec);
nsec = (u32) tk->wall_to_monotonic.tv_nsec;
/*更新mono.bass为monotonic时间*/
tk->tkr_mono.base = ns_to_ktime(seconds * NSEC_PER_SEC + nsec);
/*
* The sum of the nanoseconds portions of xtime and
* wall_to_monotonic can be greater/equal one second. Take
* this into account before updating tk->ktime_sec.
*/
nsec += (u32)(tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift);
if (nsec >= NSEC_PER_SEC)
seconds++;
tk->ktime_sec = seconds;
/* Update the monotonic raw base */
tk->tkr_raw.base = ns_to_ktime(tk->raw_sec * NSEC_PER_SEC);
}
/* 更新tk信息到vdso的内存信息中供用户端快速获取*/
void update_vsyscall(struct timekeeper *tk)
{
struct vdso_data *vdata = __arch_get_k_vdso_data();
struct vdso_timestamp *vdso_ts;
s32 clock_mode;
u64 nsec;
/* copy vsyscall data */
vdso_write_begin(vdata);
clock_mode = tk->tkr_mono.clock->vdso_clock_mode;
vdata[CS_HRES_COARSE].clock_mode = clock_mode;
vdata[CS_RAW].clock_mode = clock_mode;
/* CLOCK_REALTIME also required for time() */
vdso_ts = &vdata[CS_HRES_COARSE].basetime[CLOCK_REALTIME];
vdso_ts->sec = tk->xtime_sec;
vdso_ts->nsec = tk->tkr_mono.xtime_nsec;
/* CLOCK_REALTIME_COARSE */
vdso_ts = &vdata[CS_HRES_COARSE].basetime[CLOCK_REALTIME_COARSE];
vdso_ts->sec = tk->xtime_sec;
vdso_ts->nsec = tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift;
/* CLOCK_MONOTONIC_COARSE */
vdso_ts = &vdata[CS_HRES_COARSE].basetime[CLOCK_MONOTONIC_COARSE];
vdso_ts->sec = tk->xtime_sec + tk->wall_to_monotonic.tv_sec;
nsec = tk->tkr_mono.xtime_nsec >> tk->tkr_mono.shift;
nsec = nsec + tk->wall_to_monotonic.tv_nsec;
vdso_ts->sec += __iter_div_u64_rem(nsec, NSEC_PER_SEC, &vdso_ts->nsec);
/*
* Read without the seqlock held by clock_getres().
* Note: No need to have a second copy.
*/
WRITE_ONCE(vdata[CS_HRES_COARSE].hrtimer_res, hrtimer_resolution);
/*
* If the current clocksource is not VDSO capable, then spare the
* update of the high resolution parts.
*/
if (clock_mode != VDSO_CLOCKMODE_NONE)
update_vdso_data(vdata, tk);
__arch_update_vsyscall(vdata, tk);
vdso_write_end(vdata);
__arch_sync_vdso_data(vdata);
}
回顾获取wall_clock时的赋值函数,其调用的堆栈为:
API层 ktime_get_real_ns()
└─> 转换层 ktime_to_ns()
└─> 核心层 ktime_get_real()
└─> 核心层 ktime_get_with_offset(TK_OFFS_REAL)
├─> 锁机制 read_seqcount_begin() / read_seqcount_retry()
├─> 计算层 timekeeping_get_ns()
│ ├─> 硬件读取 tk_clock_read()
│ │ └─> 时钟源 clock->read() (如 read_tsc 或 pvclock_clocksource_read)
│ │ └─>如果是tsc rdtsc_ordered()
│ │ └─> 汇编层 rdtsc / rdtscp 指令 (物理机) 或 读共享内存 (KVM)
│ └─> 数学换算 timekeeping_delta_to_ns()
└─> 语义投影 ktime_add(base, offset)
从函数调用关系可以明确获取的就是realtime就是host主机当时的时间。
2.1.4 KVM时钟模型
KVM时钟模型是建立在内核时钟模型的基础上的,里面也设计到了很多时间变量,有qemu传入的比如rct时间,也有设置的vcpu频率,也有从hsot获取的,这些时间变量都保存在kvm_arch的数据结构中,虚拟机内的时间跟host的时间息息相关。
cpp
struct kvm_arch {
...
/* case MSR_KVM_WALL_CLOCK_NEW: vcpu->kvm->arch.wall_clock = data;*/
gpa_t wall_clock;
bool mwait_in_guest;
bool hlt_in_guest;
bool pause_in_guest;
bool cstate_in_guest;
unsigned long irq_sources_bitmap;
/* arch_init
s64 kvmclock_offset;
/*
* This also protects nr_vcpus_matched_tsc which is read from a
* preemption-disabled region, so it must be a raw spinlock.
*/
raw_spinlock_t tsc_write_lock;
u64 last_tsc_nsec;
u64 last_tsc_write;
u32 last_tsc_khz;
u64 last_tsc_offset;
u64 cur_tsc_nsec;
u64 cur_tsc_write;
u64 cur_tsc_offset;
u64 cur_tsc_generation;
int nr_vcpus_matched_tsc;
u32 default_tsc_khz;
bool user_set_tsc;
u64 apic_bus_cycle_ns;
seqcount_raw_spinlock_t pvclock_sc;
bool use_master_clock;
u64 master_kernel_ns;
u64 master_cycle_now;
struct delayed_work kvmclock_update_work;
struct delayed_work kvmclock_sync_work;
...
}
struct kvm_vcpu_arch {
...
gpa_t time;
struct pvclock_vcpu_time_info hv_clock;
unsigned int hw_tsc_khz;
struct gfn_to_pfn_cache pv_time;
/* et guest stopped flag in pvclock flags field */
bool pvclock_set_guest_stopped_request;
struct {
u8 preempted;
u64 msr_val;
u64 last_steal;
struct gfn_to_hva_cache cache;
} st;
u64 l1_tsc_offset;
u64 tsc_offset; /* current tsc offset */
u64 last_guest_tsc;
u64 last_host_tsc;
u64 tsc_offset_adjustment;
u64 this_tsc_nsec;
u64 this_tsc_write;
u64 this_tsc_generation;
bool tsc_catchup;
bool tsc_always_catchup;
s8 virtual_tsc_shift;
u32 virtual_tsc_mult;
u32 virtual_tsc_khz;
s64 ia32_tsc_adjust_msr;
u64 msr_ia32_power_ctl;
u64 l1_tsc_scaling_ratio;
u64 tsc_scaling_ratio; /* current scaling ratio */
...
}
上面是两个跟时间相关的数据结构以及涉及到的变量,有些我们前面的章节已经提及到了,有些还没有涉及到,我们就从可能涉及到的变量进行罗列并说明这些值都是怎么设置的。
cpp
/* qemu 代码中设置tsc_khz的逻辑*/
static int kvm_arch_set_tsc_khz(CPUState *cs)
{
X86CPU *cpu = X86_CPU(cs);
CPUX86State *env = &cpu->env;
int r, cur_freq;
bool set_ioctl = false;
if (!env->tsc_khz) {
return 0;
}
...
r = set_ioctl ? kvm_vcpu_ioctl(cs, KVM_SET_TSC_KHZ, env->tsc_khz) :-ENOTSUP;
...
return 0;
}
/*内核中的回调 */
long kvm_arch_vcpu_ioctl(struct file *filp,
unsigned int ioctl, unsigned long arg)
{
struct kvm_vcpu *vcpu = filp->private_data;
void __user *argp = (void __user *)arg;
int r;
union {
struct kvm_sregs2 *sregs2;
struct kvm_lapic_state *lapic;
struct kvm_xsave *xsave;
struct kvm_xcrs *xcrs;
void *buffer;
} u;
vcpu_load(vcpu);
u.buffer = NULL;
switch (ioctl) {
...
case KVM_SET_TSC_KHZ: {
u32 user_tsc_khz;
r = -EINVAL;
/* user_tsc_khz 就是配置文件中设置的vcpu频率 */
user_tsc_khz = (u32)arg;
if (kvm_caps.has_tsc_control &&
user_tsc_khz >= kvm_caps.max_guest_tsc_khz)/* 频率也不能太大,也会报错*/
goto out;
if (user_tsc_khz == 0)
user_tsc_khz = tsc_khz;
if (!kvm_set_tsc_khz(vcpu, user_tsc_khz))
r = 0;
goto out;
}
case KVM_GET_TSC_KHZ: {
r = vcpu->arch.virtual_tsc_khz;
goto out;
}
...
case KVM_HAS_DEVICE_ATTR:
case KVM_GET_DEVICE_ATTR:
case KVM_SET_DEVICE_ATTR:
r = kvm_vcpu_ioctl_device_attr(vcpu, ioctl, argp);
break;
default:
r = -EINVAL;
}
out:
kfree(u.buffer);
out_nofree:
vcpu_put(vcpu);
return r;
}
static int kvm_set_tsc_khz(struct kvm_vcpu *vcpu, u32 user_tsc_khz)
{
u32 thresh_lo, thresh_hi;
int use_scaling = 0;
/* tsc_khz can be zero if TSC calibration fails */
/* 如果user_tsc_khz 为0 就使用host的tsc频率 ,此时vcpu的频率比host cpu的
* 频率为1*2^-48,即;
* kvm_caps.default_tsc_scaling_ratio = 1ULL << kvm_caps.tsc_scaling_ratio_frac_bits;
* kvm_caps.tsc_scaling_ratio_frac_bits = 48;
*/
if (user_tsc_khz == 0) {
/* set tsc_scaling_ratio to a safe value */
kvm_vcpu_write_tsc_multiplier(vcpu, kvm_caps.default_tsc_scaling_ratio);
return -1;
}
/* Compute a scale to convert nanoseconds in TSC cycles */
/* 根据user_tsc_khz 计算对应的shift和mult 即如何根据虚拟机的tsc计算时间*/
kvm_get_time_scale(user_tsc_khz * 1000LL, NSEC_PER_SEC,
&vcpu->arch.virtual_tsc_shift,
&vcpu->arch.virtual_tsc_mult);
/* 赋值到vcpu的arch成员变量virtual_tsc_khz 中 */
vcpu->arch.virtual_tsc_khz = user_tsc_khz;
/*
* Compute the variation in TSC rate which is acceptable
* within the range of tolerance and decide if the
* rate being applied is within that bounds of the hardware
* rate. If so, no scaling or compensation need be done.
*/
thresh_lo = adjust_tsc_khz(tsc_khz, -tsc_tolerance_ppm);
thresh_hi = adjust_tsc_khz(tsc_khz, tsc_tolerance_ppm);
if (user_tsc_khz < thresh_lo || user_tsc_khz > thresh_hi) {
pr_debug("requested TSC rate %u falls outside tolerance [%u,%u]\n",
user_tsc_khz, thresh_lo, thresh_hi);
use_scaling = 1;
}
return set_tsc_khz(vcpu, user_tsc_khz, use_scaling);
}
static int set_tsc_khz(struct kvm_vcpu *vcpu, u32 user_tsc_khz, bool scale)
{
u64 ratio;
/* Guest TSC same frequency as host TSC? */
if (!scale) {
kvm_vcpu_write_tsc_multiplier(vcpu, kvm_caps.default_tsc_scaling_ratio);
return 0;
}
/* TSC scaling supported? */
if (!kvm_caps.has_tsc_control) {
if (user_tsc_khz > tsc_khz) {
vcpu->arch.tsc_catchup = 1;
vcpu->arch.tsc_always_catchup = 1;
return 0;
} else {
pr_warn_ratelimited("user requested TSC rate below hardware speed\n");
return -1;
}
}
/* TSC scaling required - calculate ratio */
/* 计算guest tsc和host tsc的比率,单位是2^-48 */
ratio = mul_u64_u32_div(1ULL << kvm_caps.tsc_scaling_ratio_frac_bits,
user_tsc_khz, tsc_khz);
if (ratio == 0 || ratio >= kvm_caps.max_tsc_scaling_ratio) {
pr_warn_ratelimited("Invalid TSC scaling ratio - virtual-tsc-khz=%u\n",
user_tsc_khz);
return -1;
}
/* 把这个比率ratio赋值到vcpu的变量中 */
kvm_vcpu_write_tsc_multiplier(vcpu, ratio);
return 0;
}
static void kvm_vcpu_write_tsc_multiplier(struct kvm_vcpu *vcpu, u64 l1_multiplier)
{
/* 把 radio 赋值给l1_tsc_scaling_ratio */
vcpu->arch.l1_tsc_scaling_ratio = l1_multiplier;
/* Userspace is changing the multiplier while L2 is active */
if (is_guest_mode(vcpu))
vcpu->arch.tsc_scaling_ratio = kvm_calc_nested_tsc_multiplier(
l1_multiplier,
kvm_x86_call(get_l2_tsc_multiplier)(vcpu));
else
/* 把radio赋值给tsc_scaling_ratio */
vcpu->arch.tsc_scaling_ratio = l1_multiplier;
if (kvm_caps.has_tsc_control)
/* 写入到vmcs的变量中 供guest的rdtsc指令使用 */
kvm_x86_call(write_tsc_multiplier)(vcpu);
}
/* vmcs 写入函数 写到vmcs的指定地址处*/
void vmx_write_tsc_multiplier(struct kvm_vcpu *vcpu)
{
vmcs_write64(TSC_MULTIPLIER, vcpu->arch.tsc_scaling_ratio);
}
以前客户遇到过虚拟机长时间运行重启后虚拟机卡死的情况,当时的结论是tsc值过大产生的溢出,现在从代码端看看其对应的逻辑
cpp
/*qemu代码中设置env->tsc逻辑*/
static int kvm_put_msrs(X86CPU *cpu, int level)
{
...
kvm_msr_entry_add(cpu, MSR_IA32_TSC, env->tsc);
...
}
/* host端对应的逻辑 */
int kvm_set_msr_common(struct kvm_vcpu *vcpu, struct msr_data *msr_info)
{
u32 msr = msr_info->index;
u64 data = msr_info->data;
if (msr && msr == vcpu->kvm->arch.xen_hvm_config.msr)
return kvm_xen_write_hypercall_page(vcpu, data);
switch (msr) {
...
case MSR_IA32_TSC:
/* 重启的时候host_initialted已经为true */
if (msr_info->host_initiated) {
kvm_synchronize_tsc(vcpu, &data);
} else {
u64 adj = kvm_compute_l1_tsc_offset(vcpu, data) - vcpu->arch.l1_tsc_offset;
adjust_tsc_offset_guest(vcpu, adj);
vcpu->arch.ia32_tsc_adjust_msr += adj;
}
break;
case MSR_KVM_WALL_CLOCK_NEW:
if (!guest_pv_has(vcpu, KVM_FEATURE_CLOCKSOURCE2))
return 1;
/* 初始化arch.wall_clock为guest wall_clock的gha地址 */
vcpu->kvm->arch.wall_clock = data;
kvm_write_wall_clock(vcpu->kvm, data, 0);
break;
case MSR_KVM_SYSTEM_TIME_NEW:
if (!guest_pv_has(vcpu, KVM_FEATURE_CLOCKSOURCE2))
return 1;
/* 初始化vcpu->arch.time 为guest 中prvi的gpa地址 */
kvm_write_system_time(vcpu, data, false, msr_info->host_initiated);
break;
...
}
static void kvm_synchronize_tsc(struct kvm_vcpu *vcpu, u64 *user_value)
{
u64 data = user_value ? *user_value : 0;
struct kvm *kvm = vcpu->kvm;
u64 offset, ns, elapsed;
unsigned long flags;
bool matched = false;
bool synchronizing = false;
raw_spin_lock_irqsave(&kvm->arch.tsc_write_lock, flags);
/* 计算设置的env->tsc与host tsc的偏移 */
/* 该偏移的含义是以env->tsc为值,当guest的tsc为0时
* host端在guest的频率下tsc的值
* 比如: env->tsc 为20 ratio为 2
* host读取的tsc值为2000那么这个offset为:
* 2000/2 -20=980
* 即offset = 980
*/
offset = kvm_compute_l1_tsc_offset(vcpu, data);
/* 获取现在host对应的boottime */
ns = get_kvmclock_base_ns();
elapsed = ns - kvm->arch.last_tsc_nsec;
if (vcpu->arch.virtual_tsc_khz) {
if (data == 0) {
/*
* Force synchronization when creating a vCPU, or when
* userspace explicitly writes a zero value.
*/
synchronizing = true;
} else if (kvm->arch.user_set_tsc) {
u64 tsc_exp = kvm->arch.last_tsc_write +
nsec_to_cycles(vcpu, elapsed);
u64 tsc_hz = vcpu->arch.virtual_tsc_khz * 1000LL;
/*
* Here lies UAPI baggage: when a user-initiated TSC write has
* a small delta (1 second) of virtual cycle time against the
* previously set vCPU, we assume that they were intended to be
* in sync and the delta was only due to the racy nature of the
* legacy API.
*
* This trick falls down when restoring a guest which genuinely
* has been running for less time than the 1 second of imprecision
* which we allow for in the legacy API. In this case, the first
* value written by userspace (on any vCPU) should not be subject
* to this 'correction' to make it sync up with values that only
* come from the kernel's default vCPU creation. Make the 1-second
* slop hack only trigger if the user_set_tsc flag is already set.
*/
synchronizing = data < tsc_exp + tsc_hz &&
data + tsc_hz > tsc_exp;
}
}
if (user_value)
kvm->arch.user_set_tsc = true;
/*
* For a reliable TSC, we can match TSC offsets, and for an unstable
* TSC, we add elapsed time in this computation. We could let the
* compensation code attempt to catch up if we fall behind, but
* it's better to try to match offsets from the beginning.
*/
if (synchronizing &&
vcpu->arch.virtual_tsc_khz == kvm->arch.last_tsc_khz) {
if (!kvm_check_tsc_unstable()) {
offset = kvm->arch.cur_tsc_offset;
} else {
u64 delta = nsec_to_cycles(vcpu, elapsed);
data += delta;
offset = kvm_compute_l1_tsc_offset(vcpu, data);
}
matched = true;
}
/* 设置arch的各种变量 */
__kvm_synchronize_tsc(vcpu, offset, data, ns, matched);
raw_spin_unlock_irqrestore(&kvm->arch.tsc_write_lock, flags);
}
static u64 kvm_compute_l1_tsc_offset(struct kvm_vcpu *vcpu, u64 target_tsc)
{
u64 tsc;
/* 这个函数很好理解,rdtsc获取主机的tsc值,根据比率得到在
* guest的vcpu频率下对应的tsc值,
* 用我们设的env->tsc 减去这个值,实际就是在guest的频率标准下,当guest的tsc值为0时
* host端的tsc的值,这个值是负值。
*/
tsc = kvm_scale_tsc(rdtsc(), vcpu->arch.l1_tsc_scaling_ratio);
return target_tsc - tsc;
}
static struct pvclock_gtod_data pvclock_gtod_data;
static void update_pvclock_gtod(struct timekeeper *tk)
{
struct pvclock_gtod_data *vdata = &pvclock_gtod_data;
write_seqcount_begin(&vdata->seq);
vdata->clock.vclock_mode = tk->tkr_mono.clock->vdso_clock_mode;
vdata->clock.cycle_last = tk->tkr_mono.cycle_last;
vdata->clock.mask = tk->tkr_mono.mask;
vdata->clock.mult = tk->tkr_mono.mult;
vdata->clock.shift = tk->tkr_mono.shift;
vdata->clock.base_cycles = tk->tkr_mono.xtime_nsec;
vdata->clock.offset = tk->tkr_mono.base;
vdata->raw_clock.vclock_mode = tk->tkr_raw.clock->vdso_clock_mode;
vdata->raw_clock.cycle_last = tk->tkr_raw.cycle_last;
vdata->raw_clock.mask = tk->tkr_raw.mask;
vdata->raw_clock.mult = tk->tkr_raw.mult;
vdata->raw_clock.shift = tk->tkr_raw.shift;
vdata->raw_clock.base_cycles = tk->tkr_raw.xtime_nsec;
vdata->raw_clock.offset = tk->tkr_raw.base;
vdata->wall_time_sec = tk->xtime_sec;
vdata->offs_boot = tk->offs_boot;
write_seqcount_end(&vdata->seq);
}
static s64 get_kvmclock_base_ns(void)
{
/* Count up from boot time, but with the frequency of the raw clock. */
/* ktime_get_raw是获取自启动时到现在的monotonic raw时间包括暂停,
* 再加上offs_boot就是主机的boottime。
* 所有get_kvm_clock_base_ns实际就是以host的boottime为时间轴,取一个锚点。
*/
return ktime_to_ns(ktime_add(ktime_get_raw(), pvclock_gtod_data.offs_boot));
}
/* 把这些设置的值放到该放的kvm->arch和vcpu->arch中 */
static void __kvm_synchronize_tsc(struct kvm_vcpu *vcpu, u64 offset, u64 tsc,
u64 ns, bool matched)
{
struct kvm *kvm = vcpu->kvm;
lockdep_assert_held(&kvm->arch.tsc_write_lock);
/*
* We also track th most recent recorded KHZ, write and time to
* allow the matching interval to be extended at each write.
*/
/* 对应host的boottime */
kvm->arch.last_tsc_nsec = ns;
/* 对应 qemu中evn->tsc的值 */
kvm->arch.last_tsc_write = tsc;
/* 对应vcpu的频率 */
kvm->arch.last_tsc_khz = vcpu->arch.virtual_tsc_khz;
/* 对应虚拟机启动时即tsc为0时,host在vcpu的频率下的频率*/
kvm->arch.last_tsc_offset = offset;
vcpu->arch.last_guest_tsc = tsc;
/* 把这些值往vcpu->arch中再写一份,以及把offset写到vmcs中*/
kvm_vcpu_write_tsc_offset(vcpu, offset);
if (!matched) {
/*
* We split periods of matched TSC writes into generations.
* For each generation, we track the original measured
* nanosecond time, offset, and write, so if TSCs are in
* sync, we can match exact offset, and if not, we can match
* exact software computation in compute_guest_tsc()
*
* These values are tracked in kvm->arch.cur_xxx variables.
*/
kvm->arch.cur_tsc_generation++;
kvm->arch.cur_tsc_nsec = ns;
kvm->arch.cur_tsc_write = tsc;
kvm->arch.cur_tsc_offset = offset;
kvm->arch.nr_vcpus_matched_tsc = 0;
} else if (vcpu->arch.this_tsc_generation != kvm->arch.cur_tsc_generation) {
kvm->arch.nr_vcpus_matched_tsc++;
}
/* Keep track of which generation this VCPU has synchronized to */
/* 就是保持vcpu中的变量跟kvm中的变量能保持一致 */
vcpu->arch.this_tsc_generation = kvm->arch.cur_tsc_generation;
vcpu->arch.this_tsc_nsec = kvm->arch.cur_tsc_nsec;
vcpu->arch.this_tsc_write = kvm->arch.cur_tsc_write;
kvm_track_tsc_matching(vcpu, !matched);
}
static void kvm_vcpu_write_tsc_offset(struct kvm_vcpu *vcpu, u64 l1_offset)
{
trace_kvm_write_tsc_offset(vcpu->vcpu_id,
vcpu->arch.l1_tsc_offset,
l1_offset);
/* offset 复制到 vpcu的arch中 */
vcpu->arch.l1_tsc_offset = l1_offset;
/*
* If we are here because L1 chose not to trap WRMSR to TSC then
* according to the spec this should set L1's TSC (as opposed to
* setting L1's offset for L2).
*/
if (is_guest_mode(vcpu))
vcpu->arch.tsc_offset = kvm_calc_nested_tsc_offset(
l1_offset,
kvm_x86_call(get_l2_tsc_offset)(vcpu),
kvm_x86_call(get_l2_tsc_multiplier)(vcpu));
else
/* offset 复制到 vpcu的arch中 */
vcpu->arch.tsc_offset = l1_offset;
/* 写入到vmcs中*/
kvm_x86_call(write_tsc_offset)(vcpu);
}
/* 这里面有点说法的:
* vmcs是vcpu可以直接读取的变量,vcpu访问的时候是不用产生vm-exit的,而在guest的tsc的值是虚拟的
* 他是通过用host里的值进行变换出来的,变换公式为:rdtsc()/radio -- tsc_offset
* 而这个算数操作在cpu的指令中用汇编已经实现,就是说在虚拟机中执行rdtsc获取的就是guest的tsc值。
*
void vmx_write_tsc_offset(struct kvm_vcpu *vcpu)
{
vmcs_write64(TSC_OFFSET, vcpu->arch.tsc_offset);
}
static int vcpu_enter_guest(struct kvm_vcpu *vcpu)
{
int r;
bool req_int_win =
dm_request_for_irq_injection(vcpu) &&
kvm_cpu_accept_dm_intr(vcpu);
fastpath_t exit_fastpath;
bool req_immediate_exit = false;
if (kvm_request_pending(vcpu)) {
...
/* 在时钟初始化更新prvi的时候会发出该请求*/
if (kvm_check_request(KVM_REQ_MASTERCLOCK_UPDATE, vcpu))
kvm_update_masterclock(vcpu->kvm);
...
}
static void kvm_update_masterclock(struct kvm *kvm)
{
kvm_hv_request_tsc_page_update(kvm);
kvm_start_pvclock_update(kvm);
pvclock_update_vm_gtod_copy(kvm);
kvm_end_pvclock_update(kvm);
}
static void pvclock_update_vm_gtod_copy(struct kvm *kvm)
{
#ifdef CONFIG_X86_64
struct kvm_arch *ka = &kvm->arch;
int vclock_mode;
bool host_tsc_clocksource, vcpus_matched;
lockdep_assert_held(&kvm->arch.tsc_write_lock);
vcpus_matched = (ka->nr_vcpus_matched_tsc + 1 ==
atomic_read(&kvm->online_vcpus));
/*
* If the host uses TSC clock, then passthrough TSC as stable
* to the guest.
*/
/* 初始化 master_kernel_ns 和 master_cycle_now */
host_tsc_clocksource = kvm_get_time_and_clockread(
&ka->master_kernel_ns,
&ka->master_cycle_now);
ka->use_master_clock = host_tsc_clocksource && vcpus_matched
&& !ka->backwards_tsc_observed
&& !ka->boot_vcpu_runs_old_kvmclock;
if (ka->use_master_clock)
atomic_set(&kvm_guest_has_master_clock, 1);
vclock_mode = pvclock_gtod_data.clock.vclock_mode;
trace_kvm_update_master_clock(ka->use_master_clock, vclock_mode,
vcpus_matched);
#endif
}
/*
* Calculates the kvmclock_base_ns (CLOCK_MONOTONIC_RAW + boot time) and
* reports the TSC value from which it do so. Returns true if host is
* using TSC based clocksource.
*/
static bool kvm_get_time_and_clockread(s64 *kernel_ns, u64 *tsc_timestamp)
{
/* checked again under seqlock below */
if (!gtod_is_based_on_tsc(pvclock_gtod_data.clock.vclock_mode))
return false;
/* 核心函数do_kvmclock_base获取kernel_ns(master_kernel_ns)和tsc_timestamp(master_cycle_now)*/
return gtod_is_based_on_tsc(do_kvmclock_base(kernel_ns,
tsc_timestamp));
}
/*
* As with get_kvmclock_base_ns(), this counts from boot time, at the
* frequency of CLOCK_MONOTONIC_RAW (hence adding gtos->offs_boot).
*/
static int do_kvmclock_base(s64 *t, u64 *tsc_timestamp)
{
struct pvclock_gtod_data *gtod = &pvclock_gtod_data;
unsigned long seq;
int mode;
u64 ns;
/* 可以看到用到了全局变量pvclock_gtod_data,该变量在update_pvclock_gtod中更
* 用到的变量与tk里的变量对应关系为:
* vdata->raw_clock.base_cycles = tk->tkr_raw.xtime_nsec;
* vdata->raw_clock.shift = tk->tkr_raw.shift;
* vdata->raw_clock.offset = tk->tkr_raw.base;
* vdata->offs_boot = tk->offs_boot;
* 根据以上变量我们进行对下面的公式进行转换:
* tk->tkr_raw.xtime_nsec + monotonic_raw(度过) + k->tkr_raw.base + tk->offs_boot
* 换个位置:
* tk->offs_boot + (tk->tkr_raw.xtime_nsec + k->tkr_raw.base + monotonic_raw(度过))
* = 休眠时间 + monotonic_raw(总共) = 主机的boottime (从开机到现在的时间)
*/
do {
seq = read_seqcount_begin(>od->seq);
ns = gtod->raw_clock.base_cycles;
/* host 直接在vsdo中直接拿取clock信息 */
ns += vgettsc(>od->raw_clock, tsc_timestamp, &mode);
ns >>= gtod->raw_clock.shift;
ns += ktime_to_ns(ktime_add(gtod->raw_clock.offset, gtod->offs_boot));
} while (unlikely(read_seqcount_retry(>od->seq, seq)));
*t = ns;
return mode;
}
/* 读取host的tsc值并返回,经过的monotonic raw时间 */
static inline u64 vgettsc(struct pvclock_clock *clock, u64 *tsc_timestamp,
int *mode)
{
u64 tsc_pg_val;
long v;
switch (clock->vclock_mode) {
case VDSO_CLOCKMODE_HVCLOCK:
if (hv_read_tsc_page_tsc(hv_get_tsc_page(),
tsc_timestamp, &tsc_pg_val)) {
/* TSC page valid */
*mode = VDSO_CLOCKMODE_HVCLOCK;
v = (tsc_pg_val - clock->cycle_last) &
clock->mask;
} else {
/* TSC page invalid */
*mode = VDSO_CLOCKMODE_NONE;
}
break;
case VDSO_CLOCKMODE_TSC:
*mode = VDSO_CLOCKMODE_TSC;
*tsc_timestamp = read_tsc();
v = (*tsc_timestamp - clock->cycle_last) &
clock->mask;
break;
default:
*mode = VDSO_CLOCKMODE_NONE;
}
if (*mode == VDSO_CLOCKMODE_NONE)
*tsc_timestamp = v = 0;
return v * clock->mult;
}
static __always_inline bool
hv_read_tsc_page_tsc(const struct ms_hyperv_tsc_page *tsc_pg,
u64 *cur_tsc, u64 *time)
{
u64 scale, offset;
u32 sequence;
/*
* The protocol for reading Hyper-V TSC page is specified in Hypervisor
* Top-Level Functional Specification ver. 3.0 and above. To get the
* reference time we must do the following:
* - READ ReferenceTscSequence
* A special '0' value indicates the time source is unreliable and we
* need to use something else. The currently published specification
* versions (up to 4.0b) contain a mistake and wrongly claim '-1'
* instead of '0' as the special value, see commit c35b82ef0294.
* - ReferenceTime =
* ((RDTSC() * ReferenceTscScale) >> 64) + ReferenceTscOffset
* - READ ReferenceTscSequence again. In case its value has changed
* since our first reading we need to discard ReferenceTime and repeat
* the whole sequence as the hypervisor was updating the page in
* between.
*/
do {
sequence = READ_ONCE(tsc_pg->tsc_sequence);
if (!sequence)
return false;
/*
* Make sure we read sequence before we read other values from
* TSC page.
*/
smp_rmb();
scale = READ_ONCE(tsc_pg->tsc_scale);
offset = READ_ONCE(tsc_pg->tsc_offset);
/* #define hv_get_raw_timer() rdtsc_ordered()*/
*cur_tsc = hv_get_raw_timer();
/*
* Make sure we read sequence after we read all other values
* from TSC page.
*/
smp_rmb();
} while (READ_ONCE(tsc_pg->tsc_sequence) != sequence);
*time = mul_u64_u64_shr(*cur_tsc, scale, 64) + offset;
return true;
}
qemu中设置kvm_clock的逻辑:
cpp
static void kvmclock_vm_state_change(void *opaque, bool running,
RunState state)
{
KVMClockState *s = opaque;
CPUState *cpu;
int cap_clock_ctrl = kvm_check_extension(kvm_state, KVM_CAP_KVMCLOCK_CTRL);
int ret;
/* 虚拟机恢复时设置该值 */
if (running) {
struct kvm_clock_data data = {};
/*
* If the host where s->clock was read did not support reliable
* KVM_GET_CLOCK, read kvmclock value from memory.
*/
if (!s->clock_is_reliable) {
uint64_t pvclock_via_mem = kvmclock_current_nsec(s);
/* We can't rely on the saved clock value, just discard it */
if (pvclock_via_mem) {
s->clock = pvclock_via_mem;
}
}
s->clock_valid = false;
data.clock = s->clock;
ret = kvm_vm_ioctl(kvm_state, KVM_SET_CLOCK, &data);
if (ret < 0) {
fprintf(stderr, "KVM_SET_CLOCK failed: %s\n", strerror(-ret));
abort();
}
if (!cap_clock_ctrl) {
return;
}
CPU_FOREACH(cpu) {
run_on_cpu(cpu, do_kvmclock_ctrl, RUN_ON_CPU_NULL);
}
/* 虚拟机暂停时获取 clock值 */
} else {
if (s->clock_valid) {
return;
}
s->runstate_paused = runstate_check(RUN_STATE_PAUSED);
kvm_synchronize_all_tsc();
/* 获取 clock值 */
kvm_update_clock(s);
/*
* If the VM is stopped, declare the clock state valid to
* avoid re-reading it on next vmsave (which would return
* a different value). Will be reset when the VM is continued.
*/
s->clock_valid = true;
}
}
static void kvm_update_clock(KVMClockState *s)
{
struct kvm_clock_data data;
int ret;
ret = kvm_vm_ioctl(kvm_state, KVM_GET_CLOCK, &data);
if (ret < 0) {
fprintf(stderr, "KVM_GET_CLOCK failed: %s\n", strerror(-ret));
abort();
}
s->clock = data.clock;
/* If kvm_has_adjust_clock_stable() is false, KVM_GET_CLOCK returns
* essentially CLOCK_MONOTONIC plus a guest-specific adjustment. This
* can drift from the TSC-based value that is computed by the guest,
* so we need to go through kvmclock_current_nsec(). If
* kvm_has_adjust_clock_stable() is true, and the flags contain
*KVM_CLOCK_TSC_STABLE, then KVM_GET_CLOCK returns a TSC-based value
* and kvmclock_current_nsec() is not necessary.
*
* Here, however, we need not check KVM_CLOCK_TSC_STABLE. This is because
* - if the host has disabled the kvmclock master clock, the guest already
* has protection against time going backwards. This "safety net" is only
* absent when kvmclock is stable;
*
* - therefore, we can replace a check like
*
* if last KVM_GET_CLOCK was not reliable then
* read from memory
*
* with
*
* if last KVM_GET_CLOCK was not reliable && masterclock is enabled
* read from memory
*
* However:
* - if kvm_has_adjust_clock_stable() returns false, the left side is
* always true (KVM_GET_CLOCK is never reliable), and the right side is
* unknown (because we don't have data.flags). We must assume it's true
* and read from memory.
*
* - if kvm_has_adjust_clock_stable() returns true, the result of the &&
* is always false (masterclock is enabled iff KVM_GET_CLOCK is reliable)
*
* So we can just use this instead:
*
* if !kvm_has_adjust_clock_stable() then
* read from memory
*/
s->clock_is_reliable = kvm_has_adjust_clock_stable();
}
kernel中的回调:
cpp
int kvm_arch_vm_ioctl(struct file *filp, unsigned int ioctl, unsigned long arg)
{
struct kvm *kvm = filp->private_data;
void __user *argp = (void __user *)arg;
int r = -ENOTTY;
/*
* This union makes it completely explicit to gcc-3.x
* that these two variables' stack usage should be
* combined, not added together.
*/
union {
struct kvm_pit_state ps;
struct kvm_pit_state2 ps2;
struct kvm_pit_config pit_config;
} u;
switch (ioctl) {
...
case KVM_SET_CLOCK:
r = kvm_vm_ioctl_set_clock(kvm, argp);
break;
case KVM_GET_CLOCK:
r = kvm_vm_ioctl_get_clock(kvm, argp);
break;
case KVM_SET_TSC_KHZ: {
...
}
static int kvm_vm_ioctl_get_clock(struct kvm *kvm, void __user *argp)
{
struct kvm_clock_data data = { 0 };
get_kvmclock(kvm, &data);
if (copy_to_user(argp, &data, sizeof(data)))
return -EFAULT;
return 0;
}
static void get_kvmclock(struct kvm *kvm, struct kvm_clock_data *data)
{
struct kvm_arch *ka = &kvm->arch;
unsigned seq;
do {
seq = read_seqcount_begin(&ka->pvclock_sc);
__get_kvmclock(kvm, data);
} while (read_seqcount_retry(&ka->pvclock_sc, seq));
}
static void __get_kvmclock(struct kvm *kvm, struct kvm_clock_data *data)
{
struct kvm_arch *ka = &kvm->arch;
struct pvclock_vcpu_time_info hv_clock;
/* both __this_cpu_read() and rdtsc() should be on the same cpu */
get_cpu();
data->flags = 0;
if (ka->use_master_clock &&
(static_cpu_has(X86_FEATURE_CONSTANT_TSC) || __this_cpu_read(cpu_tsc_khz))) {
#ifdef CONFIG_X86_64
struct timespec64 ts;
if (kvm_get_walltime_and_clockread(&ts, &data->host_tsc)) {
data->realtime = ts.tv_nsec + NSEC_PER_SEC * ts.tv_sec;
data->flags |= KVM_CLOCK_REALTIME | KVM_CLOCK_HOST_TSC;
} else
#endif
data->host_tsc = rdtsc();
data->flags |= KVM_CLOCK_TSC_STABLE;
hv_clock.tsc_timestamp = ka->master_cycle_now;
/* master_kernel_ns是主机的boottime,kvmcloc_offset是guest与host的boot_time的
* 偏移,kvm->arch.kvmclock_offset = -get_kvmclock_base_ns() = -boottime_raw;
* 但是在kvm_vm_ioctl_set_clock中kvmclock_offset会把guest的休眠时间添加进去
* 那么system_time就是guest系统的运行时间去掉guest休眠的
*/
hv_clock.system_time = ka->master_kernel_ns + ka->kvmclock_offset;
kvm_get_time_scale(NSEC_PER_SEC, get_cpu_tsc_khz() * 1000LL,
&hv_clock.tsc_shift,
&hv_clock.tsc_to_system_mul);
data->clock = __pvclock_read_cycles(&hv_clock, data->host_tsc);
} else {
/* 根据kvm_arch_init_vm 函数中对kvmclock_offset的初始化:
* kvm->arch.kvmclock_offset = -get_kvmclock_base_ns();
* get_kvmclock_base_ns()表示此时的boottime,那么kvmclock_offset就是表示虚拟机启动时
* 相比较主机的偏移
* 那么data->clock 就是此时主机的monotonic_raw_time
* 或者说是此时虚拟机运行的时间不包括休眠时间,ka->kvmclock_offset把休眠时间包括进去了
* kvmclock_offset是负值。
*/
data->clock = get_kvmclock_base_ns() + ka->kvmclock_offset;
}
put_cpu();
}
static int kvm_vm_ioctl_set_clock(struct kvm *kvm, void __user *argp)
{
struct kvm_arch *ka = &kvm->arch;
struct kvm_clock_data data;
u64 now_raw_ns;
if (copy_from_user(&data, argp, sizeof(data)))
return -EFAULT;
/*
* Only KVM_CLOCK_REALTIME is used, but allow passing the
* result of KVM_GET_CLOCK back to KVM_SET_CLOCK.
*/
if (data.flags & ~KVM_CLOCK_VALID_FLAGS)
return -EINVAL;
kvm_hv_request_tsc_page_update(kvm);
kvm_start_pvclock_update(kvm);
/* 更新ka中的时间变量特别是ka->master_kernel_ns (host的boottime)*/
pvclock_update_vm_gtod_copy(kvm);
/*
* This pairs with kvm_guest_time_update(): when masterclock is
* in use, we use master_kernel_ns + kvmclock_offset to set
* unsigned 'system_time' so if we use get_kvmclock_ns() (which
* is slightly ahead) here we risk going negative on unsigned
* 'system_time' when 'data.clock' is very small.
*/
if (data.flags & KVM_CLOCK_REALTIME) {
/* 获取主机的realtime */
u64 now_real_ns = ktime_get_real_ns();
/*
* Avoid stepping the kvmclock backwards.
*/
if (now_real_ns > data.realtime)
data.clock += now_real_ns - data.realtime;
}
if (ka->use_master_clock)
/* 现在的monotonic_raw时间包括暂停的 */
now_raw_ns = ka->master_kernel_ns;
else
now_raw_ns = get_kvmclock_base_ns();
/* data.clock是虚拟机的monotonic_raw时间,每次暂停要从新设置
* 就是说now_raw_ns一直在走,而data.clock却不变
* kvmclock_offset本神就是赋值,这样绝对值更大了,
* 如果guest暂停的时候data.clock是不变的那么他就是guest的monotonic时间
*
ka->kvmclock_offset = data.clock - now_raw_ns;
kvm_end_pvclock_update(kvm);
return 0;
}
这时候再看get_kvmclock_ns的调用拓扑为
API层 get_kvmclock_ns(struct kvm *kvm)
└─> 核心层 get_kvmclock(struct kvm *kvm, struct kvm_clock_data *data)
├─> 锁机制 spin_lock_irqsave(&ka->pvclock_gtod_sync_lock)
├─> 状态判断 检查 ka->use_master_clock (是否使用主时钟/TSC)
├─> 计算层 (如果 use_master_clock 为真)
│ ├─> 获取Host TSC rdtsc() / rdpmc()
│ ├─> 应用偏移 guest_tsc = host_tsc + kvm->arch. kvmclock_offset
│ └─> 数学换算 pvclock_scale_delta() (将 Guest TSC 转换为纳秒)
└─> 回退层 (如果 use_master_clock 为假)
└─> 直接读取 Host 的 ktime_get_ns() 并加上固定偏移
具体函数为:
cpp
u64 get_kvmclock_ns(struct kvm *kvm)
{
struct kvm_clock_data data;
get_kvmclock(kvm, &data);
return data.clock;
}
实际就是虚拟机的monotonic_raw_boottime,虚拟机live状态的时间。
现在二刷kvm_guest_time_update,在此之前我们还要明确一些变量之间的关系:
内核中kvm初始化时,把用slab分配的方式把vcpu中的arch到stats_id这一段内存暴漏给了用户态,可被qemu访问使用。
cpp
int kvm_init(unsigned vcpu_size, unsigned vcpu_align, struct module *module)
{
int r;
int cpu;
/* A kmem cache lets us meet the alignment requirements of fx_save. */
if (!vcpu_align)
vcpu_align = __alignof__(struct kvm_vcpu);
kvm_vcpu_cache =
kmem_cache_create_usercopy("kvm_vcpu", vcpu_size, vcpu_align,
SLAB_ACCOUNT,
offsetof(struct kvm_vcpu, arch),
offsetofend(struct kvm_vcpu, stats_id)
- offsetof(struct kvm_vcpu, arch),
NULL);
...
}
qemu中也有pvclock_vcpu_time_info,获取方式是kvmclock_current_nsec,以直接读取的方式获取,因为guest的内存都是qemu虚拟化的,只要知道gpa,qemu就可以转化为hva,然后直接访问。
cpp
struct pvclock_vcpu_time_info {
uint32_t version;
uint32_t pad0;
uint64_t tsc_timestamp;
uint64_t system_time;
uint32_t tsc_to_system_mul;
int8_t tsc_shift;
uint8_t flags;
uint8_t pad[2];
} __attribute__((__packed__)); /* 32 bytes */
static uint64_t kvmclock_current_nsec(KVMClockState *s)
{
CPUState *cpu = first_cpu;
CPUX86State *env = cpu->env_ptr;
hwaddr kvmclock_struct_pa;
uint64_t migration_tsc = env->tsc;
struct pvclock_vcpu_time_info time;
uint64_t delta;
uint64_t nsec_lo;
uint64_t nsec_hi;
uint64_t nsec;
cpu_synchronize_state(cpu);
if (!(env->system_time_msr & 1ULL)) {
/* KVM clock not active */
return 0;
}
/* env中保存的有pvti的gpa地址 通过msr的方式获取的 */
kvmclock_struct_pa = env->system_time_msr & ~1ULL;
cpu_physical_memory_read(kvmclock_struct_pa, &time, sizeof(time));
assert(time.tsc_timestamp <= migration_tsc);
delta = migration_tsc - time.tsc_timestamp;
if (time.tsc_shift < 0) {
delta >>= -time.tsc_shift;
} else {
delta <<= time.tsc_shift;
}
mulu64(&nsec_lo, &nsec_hi, delta, time.tsc_to_system_mul);
nsec = (nsec_lo >> 32) | (nsec_hi << 32);
return nsec + time.system_time;
}
static int kvm_get_msrs(X86CPU *cpu)
{
CPUX86State *env = &cpu->env;
struct kvm_msr_entry *msrs = cpu->kvm_msr_buf->entries;
...
for (i = 0; i < ret; i++) {
uint32_t index = msrs[i].index;
switch (index) {
...
case MSR_KVM_SYSTEM_TIME:
env->system_time_msr = msrs[i].data;
break;
case MSR_KVM_WALL_CLOCK:
env->wall_clock_msr = msrs[i].data;
break;
...
}
内核态的回调:
cpp
int kvm_get_msr_common(struct kvm_vcpu *vcpu, struct msr_data *msr_info)
{
switch (msr_info->index) {
...
case MSR_KVM_WALL_CLOCK:
if (!guest_pv_has(vcpu, KVM_FEATURE_CLOCKSOURCE))
return 1;
msr_info->data = vcpu->kvm->arch.wall_clock;
break;
case MSR_KVM_SYSTEM_TIME:
if (!guest_pv_has(vcpu, KVM_FEATURE_CLOCKSOURCE))
return 1;
msr_info->data = vcpu->arch.time;
break;
...
}
这些逻辑说明一个问题就是qemu host guest读取的都是同一块数据。而host的vcpu->arch.hv_clock则是为了往该内存中写入的一个副本。
cpp
static int kvm_guest_time_update(struct kvm_vcpu *v)
{
unsigned long flags, tgt_tsc_khz;
unsigned seq;
struct kvm_vcpu_arch *vcpu = &v->arch;
struct kvm_arch *ka = &v->kvm->arch;
s64 kernel_ns;
u64 tsc_timestamp, host_tsc;
u8 pvclock_flags;
bool use_master_clock;
kernel_ns = 0;
host_tsc = 0;
/*
* If the host uses TSC clock, then passthrough TSC as stable
* to the guest.
*/
do {
seq = read_seqcount_begin(&ka->pvclock_sc);
use_master_clock = ka->use_master_clock;
if (use_master_clock) {
host_tsc = ka->master_cycle_now;
kernel_ns = ka->master_kernel_ns;
}
} while (read_seqcount_retry(&ka->pvclock_sc, seq));
/* Keep irq disabled to prevent changes to the clock */
local_irq_save(flags);
tgt_tsc_khz = get_cpu_tsc_khz();
if (unlikely(tgt_tsc_khz == 0)) {
local_irq_restore(flags);
kvm_make_request(KVM_REQ_CLOCK_UPDATE, v);
return 1;
}
if (!use_master_clock) {
host_tsc = rdtsc();
kernel_ns = get_kvmclock_base_ns();
}
/* 计算一下这时候guest对应的tsc值,就是vcpu对应的tsc值 */
tsc_timestamp = kvm_read_l1_tsc(v, host_tsc);
/*
* We may have to catch up the TSC to match elapsed wall clock
* time for two reasons, even if kvmclock is used.
* 1) CPU could have been running below the maximum TSC rate
* 2) Broken TSC compensation resets the base at each VCPU
* entry to avoid unknown leaps of TSC even when running
* again on the same CPU. This may cause apparent elapsed
* time to disappear, and the guest to stand still or run
* very slowly.
*/
if (vcpu->tsc_catchup) {
u64 tsc = compute_guest_tsc(v, kernel_ns);
if (tsc > tsc_timestamp) {
adjust_tsc_offset_guest(v, tsc - tsc_timestamp);
tsc_timestamp = tsc;
}
}
local_irq_restore(flags);
/* With all the info we got, fill in the values */
if (kvm_caps.has_tsc_control)
tgt_tsc_khz = kvm_scale_tsc(tgt_tsc_khz,
v->arch.l1_tsc_scaling_ratio);
if (unlikely(vcpu->hw_tsc_khz != tgt_tsc_khz)) {
kvm_get_time_scale(NSEC_PER_SEC, tgt_tsc_khz * 1000LL,
&vcpu->hv_clock.tsc_shift,
&vcpu->hv_clock.tsc_to_system_mul);
vcpu->hw_tsc_khz = tgt_tsc_khz;
kvm_xen_update_tsc_info(v);
}
vcpu->hv_clock.tsc_timestamp = tsc_timestamp;
vcpu->hv_clock.system_time = kernel_ns + v->kvm->arch.kvmclock_offset;
vcpu->last_guest_tsc = tsc_timestamp;
/* If the host uses TSC clocksource, then it is stable */
pvclock_flags = 0;
if (use_master_clock)
pvclock_flags |= PVCLOCK_TSC_STABLE_BIT;
vcpu->hv_clock.flags = pvclock_flags;
if (vcpu->pv_time.active)
kvm_setup_guest_pvclock(v, &vcpu->pv_time, 0, false);
kvm_hv_setup_tsc_page(v->kvm, &vcpu->hv_clock);
return 0;
}
static void kvm_setup_guest_pvclock(struct kvm_vcpu *v,
struct gfn_to_pfn_cache *gpc,
unsigned int offset,
bool force_tsc_unstable)
{
struct kvm_vcpu_arch *vcpu = &v->arch;
struct pvclock_vcpu_time_info *guest_hv_clock;
unsigned long flags;
read_lock_irqsave(&gpc->lock, flags);
while (!kvm_gpc_check(gpc, offset + sizeof(*guest_hv_clock))) {
read_unlock_irqrestore(&gpc->lock, flags);
if (kvm_gpc_refresh(gpc, offset + sizeof(*guest_hv_clock)))
return;
read_lock_irqsave(&gpc->lock, flags);
}
guest_hv_clock = (void *)(gpc->khva + offset);
/*
* This VCPU is paused, but it's legal for a guest to read another
* VCPU's kvmclock, so we really have to follow the specification where
* it says that version is odd if data is being modified, and even after
* it is consistent.
*/
guest_hv_clock->version = vcpu->hv_clock.version = (guest_hv_clock->version + 1) | 1;
smp_wmb();
/* retain PVCLOCK_GUEST_STOPPED if set in guest copy */
vcpu->hv_clock.flags |= (guest_hv_clock->flags & PVCLOCK_GUEST_STOPPED);
if (vcpu->pvclock_set_guest_stopped_request) {
vcpu->hv_clock.flags |= PVCLOCK_GUEST_STOPPED;
vcpu->pvclock_set_guest_stopped_request = false;
}
memcpy(guest_hv_clock, &vcpu->hv_clock, sizeof(*guest_hv_clock));
if (force_tsc_unstable)
guest_hv_clock->flags &= ~PVCLOCK_TSC_STABLE_BIT;
smp_wmb();
guest_hv_clock->version = ++vcpu->hv_clock.version;
/* 还不忘标记一下位图 */
kvm_gpc_mark_dirty_in_slot(gpc);
read_unlock_irqrestore(&gpc->lock, flags);
trace_kvm_pvclock_update(v->vcpu_id, &vcpu->hv_clock);
}
static inline void kvm_gpc_mark_dirty_in_slot(struct gfn_to_pfn_cache *gpc)
{
lockdep_assert_held(&gpc->lock);
if (!gpc->memslot)
return;
mark_page_dirty_in_slot(gpc->kvm, gpc->memslot, gpa_to_gfn(gpc->gpa));
}
void mark_page_dirty_in_slot(struct kvm *kvm,
const struct kvm_memory_slot *memslot,
gfn_t gfn)
{
struct kvm_vcpu *vcpu = kvm_get_running_vcpu();
#ifdef CONFIG_HAVE_KVM_DIRTY_RING
if (WARN_ON_ONCE(vcpu && vcpu->kvm != kvm))
return;
WARN_ON_ONCE(!vcpu && !kvm_arch_allow_write_without_running_vcpu(kvm));
#endif
if (memslot && kvm_slot_dirty_track_enabled(memslot)) {
unsigned long rel_gfn = gfn - memslot->base_gfn;
u32 slot = (memslot->as_id << 16) | memslot->id
if (kvm->dirty_ring_size && vcpu)
kvm_dirty_ring_push(vcpu, slot, rel_gfn);
else if (memslot->dirty_bitmap)
set_bit_le(rel_gfn, memslot->dirty_bitmap);
}
}
void kvm_dirty_ring_push(struct kvm_vcpu *vcpu, u32 slot, u64 offset)
{
struct kvm_dirty_ring *ring = &vcpu->dirty_ring;
struct kvm_dirty_gfn *entry;
/* It should never get full */
WARN_ON_ONCE(kvm_dirty_ring_full(ring));
entry = &ring->dirty_gfns[ring->dirty_index & (ring->size - 1)];
entry->slot = slot;
entry->offset = offset;
/*
* Make sure the data is filled in before we publish this to
* the userspace program. There's no paired kernel-side reader.
*/
smp_wmb();
kvm_dirty_gfn_set_dirtied(entry);
ring->dirty_index++;
trace_kvm_dirty_ring_push(ring, slot, offset);
if (kvm_dirty_ring_soft_full(ring))
kvm_make_request(KVM_REQ_DIRTY_RING_SOFT_FULL, vcpu);
}
/*就是把地址段写入dirty_gfns*/
static inline void kvm_dirty_gfn_set_dirtied(struct kvm_dirty_gfn *gfn)
{
gfn->flags = KVM_DIRTY_GFN_F_DIRTY;
}
/* 或者是就是设置bitmap位 */
static inline void set_bit_le(int nr, void *addr)
{
set_bit(nr ^ BITOP_LE_SWIZZLE, addr);
}
/* 读取l1_tsc的函数*/
u64 kvm_read_l1_tsc(struct kvm_vcpu *vcpu, u64 host_tsc)
{
return vcpu->arch.l1_tsc_offset +
kvm_scale_tsc(host_tsc, vcpu->arch.l1_tsc_scaling_ratio);
}
EXPORT_SYMBOL_GPL(kvm_read_l1_tsc);
2.3 时钟实践
时钟问题是影响性能的一个重要方面,在支持客户的实践中也有碰到过各种问题。
2.3.1 时钟支持问题
案例:客户使用的是window 2008 虚拟机,时钟配置为
<timer name="hpet" tickpolicy="catchup"/>
<timer name="hypervclock" persent="yes"/>
虚拟机内部fio加压,查看虚拟机时钟并计时,发现ntp始终跑1分钟,虚拟机时钟一分钟,即虚拟机时钟慢了,经过对比msr寄存器配置,发现window 2008并不支持hypervclock因此会降级使用hpet始终。由于加压导致cpu比较繁忙hpet时钟产生的大量中断无法及时处理,导致时钟变慢。解决方案就是不使用hpet时钟设置
<timer name="hpet" persent="no"/>
<timer name="hypervclock" persent="no"/>
代码验证分析:
根据qemu对hpet时钟的虚拟化,当配置
<timer name="hpet" tickpolicy="catchup"/>
<timer name="hypervclock" persent="yes"/>
时,qemu的hpet时钟代码 /hw/rtc/hpet.c 会调用相应的函数,比如:hpet_get_ticks , hpte_timer等,
当配置为:
<timer name="hpet" persent="no"/>
<timer name="hypervclock" persent="no"/>
qemu代码中的/hw/rtc/mc1146818rtc.c文件就会被调用 rtc_periodic_timer等接口,即降级到用rtc时钟。
案例2:客户ntp时钟不准
客户使用的是rhel5.8虚拟机系统,ntp时间校准时,总是不准,
原因是rhel5.8本身是不支持kvm-clock的,因此红帽把kvm-clock驱动自行补上的,但是由于兼容性问题,如果有hv-xxx的feature就会导致cpuid的返回值有区别,让系统无法识别到在kvm环境中运行,无法使用kvm-clock,因此需要把相关的hv-feature都删掉。
2.3.2 tsc和kvm_clock时钟效率问题
在同一linux虚拟机下执行以下代码:
#include <time.h>
main()
{
int rc;
long i;
struct timespec ts;
for(i=0; i<500000000; i++) {
rc = clock_gettime(CLOCK_MONOTONIC, &ts);
}
发现tsc时钟和kvm-clock时钟获取monotonic时间指定次数消耗的时间不同kvm-clock大约消耗16秒,tsc大约消耗13秒,当然这个平均到每次上大约3000000000ns /500000000 = 60ns
为什么会出现这种情况,我们在第3章节已经分析了
我们再看一下到底是怎么计算的:kvm-clock
cpp
static u64 vread_pvclock(void)
{
const struct pvclock_vcpu_time_info *pvti = &pvclock_page.pvti;
u32 version;
u64 ret;
do {
/* kvm-clock获取pvti锁 */s
version = pvclock_read_begin(pvti);
if (unlikely(!(pvti->flags & PVCLOCK_TSC_STABLE_BIT)))
return U64_MAX;
/* guest中rdtsc_ordered 返回的是guest视角下的tsc值,用到了
* vmcs里的offset和radio
* 计算公式为host_tsc /radio -- offset
* 这个也是tsc时钟获取的值
*/
ret = __pvclock_read_cycles(pvti, rdtsc_ordered());
} while (pvclock_read_retry(pvti, version));
return ret & S64_MAX;
}
static __always_inline
u64 __pvclock_read_cycles(const struct pvclock_vcpu_time_info *src, u64 tsc)
{
/* tsc_timestmap是上一次记录的vcpu的频率,那么delta就是过了过长时间 monotoonic raw */
u64 delta = tsc - src->tsc_timestamp;
/* 根据delta 计算换为ns是多长时间 */
u64 offset = pvclock_scale_delta(delta, src->tsc_to_system_mul,
src->tsc_shift);
/* system_time是guest运行的时间,不包括休眠时间, 再加上offset就是现在的monotonic时间*/
return src->system_time + offset;
}
那这样做有什么好处?比较直接rdtsc_ordered又多算了几步,有什么好处那?
rdtsc_ordered用的是vmcs未回的offset和radio,而这个用的是虚拟机内部维护的mul和shift,现在看没啥区别。。。可能就是博客里说的,就是迁移的时候如果vcpu的频率不一样,vcpu的频率也会变,这样在线线迁移的时候会有问题。或者prvi中透传的信息多点。。
2.1.4 hpet时钟
hpet时钟还没有详细的调研后期可以补上,不过也有共识就是hpet是通过中断注入时间值的,这个会导致vm频繁的退出。因为需要频繁的执行qemu中的hpet_ram_read/hpet_ram_write
#kvm-clock监督10秒的中断调用
root@allinone-62112-25155 \~# timeout 10 bpftrace -e 'kprobe:kvm_lapic_* {@probe=count();}' Attaching 18 probes...
@kprobe:kvm_lapic_restart_hv_timer: 1
@kprobe:kvm_lapic_get_cr8: 1048
@kprobe:kvm_lapic_expired_hv_timer: 2015
@kprobe:kvm_lapic_switch_to_sw_timer: 20147
@kprobe:kvm_lapic_switch_to_hv_timer: 20182
@kprobe:kvm_lapic_hv_timer_in_use: 20644
@kprobe:kvm_lapic_reg_write: 36622
@kprobe:kvm_lapic_find_highest_irr: 458446
#hpet监督的10秒中断调用
root@allinone-62112-25155 \~# timeout 10 bpftrace -e 'kprobe:kvm_lapic_* {@probe=count();}' Attaching 18 probes...
@kprobe:kvm_lapic_restart_hv_timer: 1
@kprobe:kvm_lapic_sync_to_vapic: 130
@kprobe:kvm_lapic_expired_hv_timer: 2878
@kprobe:kvm_lapic_hv_timer_in_use: 20773
@kprobe:kvm_lapic_switch_to_hv_timer: 20864
@kprobe:kvm_lapic_switch_to_sw_timer: 20977
@kprobe:kvm_lapic_reg_write: 37871
@kprobe:kvm_lapic_get_cr8: 158262
@kprobe:kvm_lapic_find_highest_irr: 505076
可以看到guest使用hpet时钟时kvm_lapic_get_cr8被调用的此时增加很多。
2.1.5 时钟kick
对于keeptimeing模型,我们可以看到实际获取的时候,是不会陷入系统调用的,基本上对内核任务影响轻微,唯一的影响也就是需要获取锁后占用了,因此对于虚拟机的性能的影响可以忽略不计,即无论是kvm-clock还是tsc,虽然调用快慢有区别但是,对cpu或者说虚拟机的性能影响是非常有限的,并且有一个现象:当vcpu比较繁忙的时候tick的会更加频繁,当vcpu空闲比较高的时候tick就会降低。
root@allinone-50765-72774 \~# timeout 10 bpftrace -e 'kprobe:*timekeep* {@probe = count();}' Attaching 15 probes...
@kprobe:timekeeping_update: 5888
@kprobe:timekeeping_advance: 6525
@kprobe:timekeeping_max_deferment: 10365
@kprobe:update_fast_timekeeper: 11644
部分调用栈:
timekeeping_update+1
timekeeping_advance+834
update_wall_time+12
tick_sched_do_timer+74
tick_sched_timer+39
__hrtimer_run_queues+256
hrtimer_interrupt+256
smp_apic_timer_interrupt+106
apic_timer_interrupt+15
native_safe_halt+14
__sched_text_end+10
default_idle_call+68
do_idle+531
cpu_startup_entry+111
start_kernel+1314
secondary_startup_64
_no_verify+194
1 kvmclock代码学习 - EwanHai - 博客园
2 大话计算机:计算机系统底层架构原理极限剖析