The idea: five ways into the kernel

A CPU runs a program until something pulls it into the kernel. Linux calls most of these things interrupts, but they come from three different places:

SourceOn this pageWhen it happens
Hardwaredevice IRQ (eth0), timer, NMIat any moment, unrelated to the instruction that is running
The instruction itself (an exception)page faultwhen an instruction can't finish, here because its page isn't mapped
The program (a system call)syscallwhen the program asks the kernel for something

Every one of them works the same way. The CPU saves where it was, jumps to a handler in the kernel, and goes back when the handler is done. The box on the CPU's stack shows what is running now. The IDT (interrupt descriptor table) is the list that tells the CPU which handler belongs to which vector number.

Five sources and their IDT vectors: eth0 packet arrived, vector 0x22, device IRQ handler; APIC timer every 4 ms, vector 0xec, timer handler; perf counter or watchdog, vector 2, NMI handler; *p = 1 on an unmapped page, vector 14, page fault handler; a program calling getpid() goes through syscall and MSR_LSTAR straight to the system call handler, skipping the IDT
Hardware, the running instruction and the program itself can all pull the CPU into the kernel; all but the system call go through an IDT vector.

Device interrupt

eth0 copies a packet into memory, then sends vector 0x22 to the CPU's local APIC, the small interrupt controller next to each core. The CPU finishes its current instruction, saves RIP and the flags, turns interrupts off (IF = 0) and runs the handler. Linux splits the work in two:

  • The top half is the real handler. It runs with interrupts off, so it only acknowledges the card and raises the NET_RX softirq.
  • The bottom half, the softirq, runs right after with interrupts back on. It does the slow part: the packet goes up the TCP/IP stack and the program waiting for it is woken.
A timeline for one packet from eth0: the program runs, IRQ vector 0x22 arrives, the short top half runs with IF = 0 to acknowledge the card and raise NET_RX, the softirq runs with IF = 1 to push the packet up TCP/IP and wake the reader, then the program continues
A device interrupt is split: a short top half with interrupts off, then the slow work in a softirq with interrupts back on.

Timer interrupt

The local APIC timer fires every tick: HZ = 250 times a second on many systems, so every 4 ms. The handler adds 1 to jiffies, the kernel's tick counter, and charges the running program one tick of its time slice. When the slice is used up, the kernel switches to another task on the way back to user mode. This is how one CPU shares its time between programs. (Press Timer tick three times.)

NMI

A non-maskable interrupt (vector 2) can't be turned off. Linux uses it for perf profiling (a hardware counter overflows) and for the watchdog that detects a CPU stuck with interrupts off.

Page fault: an exception

No device is involved. The instruction *p = 1 touches a page that has no entry in the page table yet, so the CPU raises a page fault (vector 14) and saves the address in CR2. The kernel maps a page and returns to the same instruction, which now works. That is what makes it a fault. A trap, such as the int3 breakpoint, returns to the next instruction. An abort, such as a machine check, doesn't return to the program at all. See Linux Virtual Memory for page faults in detail.

Three panels with instructions 1, 2 and 3, where instruction 2 enters the kernel: after a fault such as a page fault the CPU returns to instruction 2 again; after a trap such as the int3 breakpoint it returns to instruction 3; after an abort such as a machine check it never returns to the program
A fault re-runs the instruction that failed, a trap continues after it, and an abort does not go back at all.

System call

A program asks the kernel for something, here getpid(). 64-bit Linux uses the syscall instruction, which doesn't go through the IDT: the CPU jumps straight to the address stored in the register MSR_LSTAR. sysret brings it back to the next instruction. The old 32-bit way, int $0x80, is a real software interrupt through the IDT, and slower. See A Linux System Call, Step by Step.

Turning interrupts off

Kernel code that shares data with an interrupt handler turns interrupts off while it holds the lock (spin_lock_irqsave() sets IF = 0). A device interrupt that arrives then is not lost. It waits in the APIC and is taken as soon as interrupts are back on. An NMI does not wait. (Demo: IRQs off.)

KindVectorWaits while IF = 0?Counted in
device IRQ0x22 (given out per device)yes/proc/interrupts (one line per device)
timer0xecyes/proc/interrupts LOC
NMI2no/proc/interrupts NMI
page fault14can't be delayed: the instruction needs it/proc/vmstat pgfault
system callnone (MSR_LSTAR)not an interruptnot counted

What the page leaves out

Real machines have many CPUs. They interrupt each other with IPIs (inter-processor interrupts), for example to wake a task on another CPU or to flush another CPU's TLB, and each device interrupt is routed to a chosen CPU (/proc/irq/N/smp_affinity). Also left out: ksoftirqd, which takes over softirq work under heavy load; a hardirq nesting on top of a running softirq; threaded IRQs and workqueues; and the cost of an interrupt, which is about a microsecond.

See also The Interrupt-Driven I/O Cycle for one device interrupt in full detail, DMA for how the card puts the packet in memory, and Linux Context Switch for what happens when the time slice runs out.