What a context switch is

A CPU core runs one thread at a time. Linux gives every runnable thread (the kernel calls each one a task) a turn, so the core keeps stopping one task and starting another. That handover is a context switch. The context is everything the CPU holds for the task that is not in memory already: the general-purpose registers, the instruction pointer (RIP), the stack pointer (RSP), the FS base register that points at thread-local storage, the FPU/SIMD registers, and CR3, the register that selects the page table and with it the whole address space.

The switch must save that context so that later the task continues exactly where it stopped, as if it had never left the CPU. In the animation, T1's loop counter in RAX goes on counting from where it was.

When does the kernel switch?

A switch only ever happens inside the kernel, in schedule(). A task gets there in one of these ways:

  • It blocks (voluntary). It calls read() on a socket with no data, waits for a lock, sleep()s, or waits for a page to come from disk. It marks itself asleep (state S), puts itself on a wait queue, and calls schedule().
  • Its time slice ends (involuntary). The timer interrupt fires every tick (4 ms at HZ=250). If the running task has had its share, the tick sets the TIF_NEED_RESCHED flag, and on the way back out of the interrupt the kernel calls schedule() instead of returning to the task: it is preempted.
  • A wake-up preempts it (involuntary). When a sleeping task is woken (data arrived, a lock was released) and it deserves the CPU more than the running one, the same flag is set.
  • It yields. sched_yield() gives the CPU to another runnable task. The task is still runnable, so Linux counts it as involuntary.

Linux counts both kinds per task: grep ctxt /proc/<pid>/status shows voluntary_ctxt_switches (nvcsw) and nonvoluntary_ctxt_switches (nivcsw). Many voluntary switches mean the task mostly waits for I/O; many involuntary ones mean it competes for CPU.

Which task runs next

Each CPU has a run queue. For normal tasks, the Completely Fair Scheduler (CFS) keeps them in a red-black tree ordered by vruntime: the CPU time a task has used, scaled by its weight. A nice-0 task's vruntime grows as fast as the clock; a nice 5 task's grows about 3 times faster, so it falls behind in the queue and gets less CPU. pick_next_task takes the leftmost task, the one that has had the least. A task that wakes up after a long sleep is placed near min_vruntime (with a small bonus), so it cannot hoard the CPU with credit saved while asleep.

The animation uses simple CFS numbers: a 24 ms latency shared among the runnable tasks by weight, a 4 ms minimum slice, a 4 ms wake-up granularity. Since Linux 6.6 the fair class uses EEVDF, which picks among eligible tasks by an earliest virtual deadline instead of the plain leftmost vruntime. The switch itself, which is what this page is about, is the same.

See also: CPU scheduling, which runs the same processes through FCFS, SJF, Round Robin, priority and CFS and compares their waiting and response times. Concurrency vs parallelism shows what context switches cost when one core is time-sliced between threads. From a Java thread to a CPU core follows a Java thread down to its kernel task, the per-CPU run queues and the SMT hardware threads of a core. The interrupt-driven I/O cycle shows the device interrupt that wakes a task blocked in read(), from the IRQ line to iret.

The switch, step by step (x86-64)

entry (syscall / interrupt)   save user RIP, RSP, RAX, ... in pt_regs,
                              at the top of the task's kernel stack
__schedule()
    put prev back in the run queue (if still runnable)
    next = pick_next_task(rq)
    context_switch(rq, prev, next)
        switch_mm_irqs_off()  if next->mm != the loaded mm: load CR3
        switch_to(prev, next)
            __switch_to_asm:  push rbx, rbp, r12-r15      (on prev's stack)
                              prev->thread.sp = rsp
                              rsp = next->thread.sp       <-- the switch
                              pop r15-r12, rbp, rbx       (from next's stack)
            __switch_to:      FS/GS bases, current = next, TSS.sp0,
                              FPU marked to reload before user mode
    return ...                on NEXT's stack, into next's own earlier call
exit to user mode             restore next's pt_regs, sysret / iret

The key idea is in the middle: each task has its own kernel stack, and the only thing switch_to really changes is the stack pointer. Every task that is not running is frozen inside its own earlier call to schedule(), with its callee-saved registers pushed on its kernel stack and its user registers in pt_regs above them. Once RSP points at next's stack, the ordinary function returns go back up through next's call chain: out of schedule(), out of the read() or interrupt it was in, and back to user space. Only the callee-saved registers need an explicit push, because the C calling convention already lets the call to switch_to clobber all the others.

T1 (prev, process A) and T3 (next, process B) each have a kernel stack with pt_regs on top, the read() and schedule() frames, and the pushed rbx, rbp, r12–r15 at the bottom; the CPU's RSP pointed into T1's stack before and points into T3's stack after, and CR3 changes from mm A to mm B. Steps: push registers and save RSP in T1->thread.sp, load RSP from T3->thread.sp, pop T3's registers and return up T3's calls
A context switch is mostly one instruction: RSP moves from the old task's kernel stack to the new task's, and the returns then run the new task's code.

Threads vs processes: CR3 and the TLB

Threads of one process share one mm_struct, so they share one page table. Switching between them leaves CR3 alone, and the TLB, the CPU's cache of address translations, stays valid. Switching to a different process loads its page table into CR3, and without further help that flushes the TLB: the new task's first memory accesses all miss and walk the page table. That is why a thread switch is cheaper than a process switch.

Modern x86 CPUs soften this with PCID (process-context identifiers): TLB entries are tagged with an address-space ID, so Linux can switch CR3 without a full flush and find the old entries still there when it switches back. The page-table isolation added against Meltdown (KPTI) makes PCID all the more important, since it switches CR3 on every entry to and exit from the kernel as well. Linux gives each CPU a handful of dynamic ASIDs (6 on x86) and hands them to the address spaces that ran there most recently.

The TLB row in the animation shows this. Each tick the running task touches three pages: its code, its data and its own stack. A hit is green; a miss (yellow) costs a page walk, four memory reads through the page-table levels. With PCID off, every CR3 load empties the TLB. With PCID on, entries are tagged P1 (mm A) or P2 (mm B) and survive the switch. Compare the "TLB misses" counter after Demo: time slices and Demo: time slices with PCID. T1 and T2 share their code and data pages, so a thread switch misses only on the new stack page. A real TLB has hundreds to thousands of entries, but the pattern is the same.

Three cases. T1 to T2, a thread switch: CR3 unchanged, TLB keeps A code, A data and T1 stack, so T2 hits code and data and misses only its stack: 1 page walk. T1 to T3 with PCID off: the CR3 load empties the TLB, so code, data and stack all miss: 3 page walks. T1 to T3 with PCID on: entries tagged P1 and P2 survive, T3 hits all three: 0 page walks
A thread switch keeps the TLB warm; a process switch empties it unless PCID tags let both processes' entries stay.

See also: paging and page replacement, which shows what the page table and the TLB contain, a page walk, and what a cold TLB costs.

The idle task and lazy TLB

When the run queue is empty the CPU runs its idle task (swapper/N), which halts the core until an interrupt. The idle task, like every kernel thread, has no user address space (mm == NULL). Instead of loading a page table it will never use, it keeps the previous task's one loaded and borrows it as active_mm: lazy TLB mode. If the next task to run belongs to that same process, switching to it needs no CR3 load at all. Demo 3 shows this: T3 blocks, idle borrows mm B, T3 wakes and runs again with CR3 untouched.

What a switch costs

  • Direct cost: entering the kernel, the scheduler's bookkeeping, saving and restoring registers (the FPU/AVX state is the biggest part), and the CR3 load. Around 1–2 µs on current hardware, more with KPTI and other mitigations.
  • Indirect cost: the next task finds the TLB, the CPU caches and the branch predictors full of the previous task's data. Refilling them can cost far more than the switch itself, especially for memory-heavy tasks.

This is why servers avoid one thread per connection (see the epoll page): a few threads looping on epoll switch far less than thousands of blocking threads. It is also why many runtimes switch between their own lightweight tasks in user space: Go's goroutines, Java's virtual threads, async/await in Rust, C#, Python and JavaScript. A user-level switch saves a handful of registers and never enters the kernel or touches CR3.

How to see it on a real system

vmstat 1                              # column "cs": context switches per second, all CPUs
pidstat -w 1                          # cswch/s and nvcswch/s per task
grep ctxt /proc/<pid>/status          # voluntary / nonvoluntary counts of one task
perf stat -e context-switches,cpu-migrations ./program
perf sched record ./program; perf sched latency   # who waited how long to run

What the animation leaves out

  • Only one CPU: no migrations between CPUs, no load balancing, no inter-processor interrupts to wake a task on another core.
  • Only three user registers are drawn in pt_regs (the real one holds all 21 saved values), and the callee-saved registers are drawn as one box.
  • PCIDs are fixed per process (P1, P2) instead of Linux's per-CPU dynamic ASIDs; the TLB has 8 entries and only three pages per task; no KPTI, no FPU state, no real-time or deadline scheduling classes.
  • The scheduler follows classic CFS with round numbers, not the exact kernel defaults or EEVDF.