The idea: five ways into the kernel
A CPU runs a program until something pulls it into the kernel. Linux calls most of these things interrupts, but they come from three different places:
| Source | On this page | When it happens |
|---|---|---|
| Hardware | device IRQ (eth0), timer, NMI | at any moment, unrelated to the instruction that is running |
| The instruction itself (an exception) | page fault | when an instruction can't finish, here because its page isn't mapped |
| The program (a system call) | syscall | when the program asks the kernel for something |
Every one of them works the same way. The CPU saves where it was, jumps to a handler in the kernel, and goes back when the handler is done. The box on the CPU's stack shows what is running now. The IDT (interrupt descriptor table) is the list that tells the CPU which handler belongs to which vector number.
Device interrupt
eth0 copies a packet into memory, then sends vector 0x22 to the CPU's local APIC, the small interrupt controller next to each core. The CPU finishes its current instruction, saves RIP and the flags, turns interrupts off (IF = 0) and runs the handler. Linux splits the work in two:
- The top half is the real handler. It runs with interrupts off, so it only acknowledges the card and raises the NET_RX softirq.
- The bottom half, the softirq, runs right after with interrupts back on. It does the slow part: the packet goes up the TCP/IP stack and the program waiting for it is woken.
Timer interrupt
The local APIC timer fires every tick: HZ = 250 times a second on many systems, so every 4 ms. The handler adds 1 to jiffies, the kernel's tick counter, and charges the running program one tick of its time slice. When the slice is used up, the kernel switches to another task on the way back to user mode. This is how one CPU shares its time between programs. (Press Timer tick three times.)
NMI
A non-maskable interrupt (vector 2) can't be turned off. Linux uses it for perf profiling (a hardware counter overflows) and for the watchdog that detects a CPU stuck with interrupts off.
Page fault: an exception
No device is involved. The instruction *p = 1 touches a page that has no entry in the page table yet, so the CPU raises a page fault (vector 14) and saves the address in CR2. The kernel maps a page and returns to the same instruction, which now works. That is what makes it a fault. A trap, such as the int3 breakpoint, returns to the next instruction. An abort, such as a machine check, doesn't return to the program at all. See Linux Virtual Memory for page faults in detail.
System call
A program asks the kernel for something, here getpid(). 64-bit Linux uses the syscall instruction, which doesn't go through the IDT: the CPU jumps straight to the address stored in the register MSR_LSTAR. sysret brings it back to the next instruction. The old 32-bit way, int $0x80, is a real software interrupt through the IDT, and slower. See A Linux System Call, Step by Step.
Turning interrupts off
Kernel code that shares data with an interrupt handler turns interrupts off while it holds the lock (spin_lock_irqsave() sets IF = 0). A device interrupt that arrives then is not lost. It waits in the APIC and is taken as soon as interrupts are back on. An NMI does not wait. (Demo: IRQs off.)
| Kind | Vector | Waits while IF = 0? | Counted in |
|---|---|---|---|
| device IRQ | 0x22 (given out per device) | yes | /proc/interrupts (one line per device) |
| timer | 0xec | yes | /proc/interrupts LOC |
| NMI | 2 | no | /proc/interrupts NMI |
| page fault | 14 | can't be delayed: the instruction needs it | /proc/vmstat pgfault |
| system call | none (MSR_LSTAR) | not an interrupt | not counted |
What the page leaves out
Real machines have many CPUs. They interrupt each other with IPIs (inter-processor interrupts), for example to wake a task on another CPU or to flush another CPU's TLB, and each device interrupt is routed to a chosen CPU (/proc/irq/N/smp_affinity). Also left out: ksoftirqd, which takes over softirq work under heavy load; a hardirq nesting on top of a running softirq; threaded IRQs and workqueues; and the cost of an interrupt, which is about a microsecond.
See also The Interrupt-Driven I/O Cycle for one device interrupt in full detail, DMA for how the card puts the packet in memory, and Linux Context Switch for what happens when the time slice runs out.