Why virtual memory
For page replacement (FIFO, clock, LRU, Belady's anomaly), swap and copy-on-write in detail, on a smaller 32-bit machine with a two-level table, see Paging and Page Replacement.
What a C program keeps in its stack and heap mappings, and how malloc gets them with brk and mmap: Stack vs Heap in C.
Devices don't go through this MMU: how a network card gets bus addresses from dma_map_single() and how the IOMMU translates and checks them is on Direct Memory Access (DMA).
Every process sees its own private address space: the same address 0x602010 in two processes refers to two different bytes of RAM, and a process cannot even name memory that belongs to another. The CPU translates every address a program uses, the virtual address, into a physical address in RAM, using a table the kernel keeps for each process: the page table. Memory is handled in pages of 4 KiB; a page of RAM is a frame, numbered by its PFN (page frame number).
Because the kernel controls the table, it can give memory lazily (map a page only when it is first used), share pages between processes (libraries, the zero page, pages after fork), and refuse accesses (writes to code, execution of data).
The address split and the page walk (x86-64, 4 levels)
A user address has 48 significant bits (bits 63–48 are copies of bit 47, so user addresses go up to 0x7fffffffffff). The MMU cuts them into four 9-bit indices and a 12-bit offset:
47 39 38 30 29 21 20 12 11 0
+----------+----------+----------+----------+------------+
| PGD idx | PUD idx | PMD idx | PT idx | offset |
+----------+----------+----------+----------+------------+
CR3 ──► PGD page ──[PGD idx]──► PUD page ──[PUD idx]──► PMD page ──[PMD idx]──► PT page ──[PT idx]──► frame
+ offset
Each table is itself one 4 KiB page holding 512 entries of 8 bytes. CR3 holds the physical address of the top table (the PGD). Every entry has a present bit, a writable bit, a user bit, a no-execute (NX) bit and the physical address of the next table or of the frame. A walk therefore costs four memory reads before the real access can happen. With 5-level paging (LA57) there is one more level and 57-bit addresses; the idea is the same.
The TLB, and PCID
Four extra reads for every load and store would be ruinous, so the CPU caches finished translations in the TLB (translation lookaside buffer): a small, very fast table from virtual page to frame plus permission bits. Real CPUs have a few dozen first-level entries and one to two thousand second-level ones; the animation has 8. A hit costs nothing extra; a miss costs a page walk. Programs that touch memory all over the place miss a lot, which is one reason data locality matters.
The TLB caches one address space's translations. When the kernel switches to another process it loads a new CR3, and without further help that flushes the TLB. PCID (process-context identifiers) tags every entry with a 12-bit address-space ID, so entries of several processes can live in the TLB together and survive switches. When the kernel changes a page-table entry it must also drop the stale TLB entry (invlpg, or an inter-processor "TLB shootdown" on other CPUs). The Context Switch page shows where the CR3 load happens in a switch.
Memory is given lazily
malloc, mmap or a growing stack only create a VMA (a vm_area_struct): a range with permissions, recorded in the process's mm_struct. No frame and not even page tables are allocated. The first access finds no entry, the CPU raises a page fault, and the kernel fills in the page table then:
- a read of an untouched anonymous page maps the zero page, one read-only frame of zeros shared by everybody. Reading a huge zeroed array costs almost no memory;
- a write allocates a zeroed frame and maps it read-write (and a later write to a zero-page mapping replaces it the same way);
- missing page-table pages (PUD, PMD, PT) are allocated on the way.
These are minor faults: no disk I/O. Pages of files (program code, libraries, mmaped files) are faulted in from the page cache or the disk instead; that is demand paging of files and is left out here. When the fault is fixed, the CPU retries the instruction, which now finds a valid entry.
The page-fault handler
do_user_addr_fault(regs, error_code, address = CR2):
vma = find_vma(mm, address)
if no vma contains address: SIGSEGV, SEGV_MAPERR (e.g. NULL pointer)
if access_error(error_code, vma): SIGSEGV, SEGV_ACCERR (write to code, exec of heap)
handle_mm_fault(vma, address, flags):
allocate missing PUD / PMD / PT pages
THP vma and empty PMD -> do_huge_pmd_anonymous_page (a 2 MiB page)
PTE empty -> do_anonymous_page (zero page, or a new frame)
write to read-only PTE -> do_wp_page (copy-on-write)
return and retry the instruction
So a segmentation fault is simply a page fault the kernel refuses to fix: either no mapping covers the address, or the mapping does not allow that kind of access.
fork() and copy-on-write
fork() must give the child a copy of the parent's memory, but copying gigabytes would be slow, and usually the child calls execve right away. So copy_page_range copies only the page tables: both processes point at the same frames, and every private writable page is marked read-only in both. The first write by either process faults; do_wp_page sees the VMA is writable and the page is shared, copies the 4 KiB into a new frame, and maps the copy writable (wp_page_copy). When the other process later writes the same page it is the only one left mapping the old frame, and it simply gets the page back writable without a copy (wp_page_reuse).
Huge pages
A PMD entry with the PS bit set maps a whole 2 MiB page directly: the walk stops one level early, and one TLB entry covers what would take 512 entries of 4 KiB pages. For programs with large working sets (databases, JVM heaps, VMs) that means far fewer TLB misses. Linux can use them explicitly (hugetlbfs) or automatically: transparent huge pages (THP) back suitable anonymous regions with 2 MiB pages, depending on /sys/kernel/mm/transparent_hugepage/enabled and madvise(MADV_HUGEPAGE). The costs are memory (a 2 MiB page is allocated even if only one byte is used) and the work to find 2 MiB of contiguous physical memory. When a shared huge page is written after fork, Linux does not copy 2 MiB: it splits the PMD into 512 ordinary PTEs and copies only the 4 KiB that was written.
How to see it on a real system
cat /proc/<pid>/maps # the VMAs: ranges, permissions, backing file grep -e Rss -e AnonHugePages /proc/<pid>/smaps_rollup /usr/bin/time -v ./program # minor / major page faults perf stat -e page-faults,dTLB-loads,dTLB-load-misses ./program cat /sys/kernel/mm/transparent_hugepage/enabled
What the animation leaves out
- File-backed pages, the page cache and major faults; swap; stack growth below the stack VMA.
- The huge zero page: a read of an untouched THP area allocates the 2 MiB page here, where Linux may map a shared huge zero page first.
- Reference counts are drawn as "processes mapping the frame"; the kernel's real rules (mapcount, refcount, the anon-exclusive flag) decide the same way for these cases.
- One CPU, so no TLB shootdowns; a single 8-entry TLB instead of separate instruction/data and first/second-level TLBs; no paging-structure caches; accessed and dirty bits; KPTI.
- A SIGSEGV would normally kill the process; the model lets it continue so you can try more accesses.