What happens when you type ./hello world
A program on disk is just a file. To run it, the shell asks the kernel for a new process (fork) and then asks the kernel to replace that process's program with the file (execve). The surprise is how little execve actually loads: it reads the file's headers and sets up mappings, and the code and data come off the disk later, one page at a time, when the CPU first touches them. For a dynamically linked program, the first instruction that runs is not even in the program: it is in the dynamic loader, ld-linux-x86-64.so.2, which has to find and link libc before your _start can run.
The animation follows this program, built with Debian's gcc 14 on x86-64. The segment layout, entry points, symbol offsets and file sizes on the canvas come from readelf and nm on the real binaries, and the system calls follow a real strace.
How hello is built in the first place (the preprocessor, the compiler, the assembler, object files with holes, and the linker that fills them in) is on How a C Program Is Compiled, Linked, Loaded and Started.
// hello.c gcc -O2 -o hello hello.c (a PIE, dynamically linked)
#include <stdio.h>
int counter = 1; // .data
int buf[4096]; // .bss: 16 KiB of zeros, takes no space in the file
int main(int argc, char **argv) { printf("hello %s\n", argv[1]); return 0; }
Reading the canvas
- Left, top: bash, the child it forks, and the terminal.
- Kernel column: the kernel function running now, the
linux_binprmthatexecvebuilds (with the first bytes of the file), the list of binary-format handlers, the program headers, the CPU registers and a log of system calls. - Middle: the new process's initial stack, as the kernel builds it.
- Right: the address space of pid 813, like
/proc/813/maps: one row per VMA (virtual memory area), highest address on top. The colour says which kinds of pages are present in it (the pages column counts the ones drawn): none yet (white), shared page-cache pages (green), private copies (orange), zero-filled anonymous memory (blue). ▶ marks the VMA holding the instruction the CPU is running. - Bottom: the page cache: pages of files kept in memory and shared by every process, one box per page drawn, green when cached. Below it is the disk, and a line of counters.
Next Stage runs one stage and Run to main() runs up to main's first instruction. A program tab starts that program at once, and a change of linking, page cache or ASLR starts the open program again with it. Click the open tab to run it again: with ASLR on, that picks new random addresses.
fork, then execve
fork() copies the description of bash's memory, not the memory: both processes share every page, marked copy-on-write, so a fork costs microseconds. The child then calls execve(path, argv, envp). The kernel:
- copies
argvandenvpout of the old image (they would be destroyed with it) onto the first page of the new stack; - opens the file: path walk,
xpermission, not on anoexecmount, not open for writing (ETXTBSY); - reads the first 256 bytes and asks each binfmt handler whether it recognises them:
#!scripts (binfmt_script), ELF (binfmt_elf), and user-registered formats (binfmt_misc, used for qemu-user, Wine, Java jars). A script's handler rewrites the call toexecve("/bin/sh", ["/bin/sh", "./hello.sh", "world"])and searches again.
ELF: segments are what the loader sees
An ELF file has two views. Sections (.text, .data, .bss, .got …) are for the linker and debuggers. Segments (program headers) are for loading: each PT_LOAD says "map this byte range of the file at this address, with these permissions". Our hello has four:
$ readelf -lW hello
Type Offset VirtAddr FileSiz MemSiz Flg
INTERP 0x000394 0x0394 0x00001c 0x00001c R [/lib64/ld-linux-x86-64.so.2]
LOAD 0x000000 0x0000 0x000628 0x000628 R headers, .dynsym, .rela.dyn
LOAD 0x001000 0x1000 0x000165 0x000165 R E .plt, .text (main, _start)
LOAD 0x002000 0x2000 0x000104 0x000104 R .rodata ("hello %s\n")
LOAD 0x002dd0 0x3dd0 0x00024c 0x004270 RW .init_array .dynamic .got .data | .bss
GNU_RELRO 0x3dd0 0x000230 R read-only after relocation
The last segment is 0x24c bytes in the file but 0x4270 in memory: the difference is .bss, which the kernel provides as zero-filled anonymous memory. It also has to zero the rest of the last file page (after .data), and that write is the only page fault during execve itself.
The point of no return
Up to begin_new_exec, any error (EACCES, ENOEXEC, ENOENT for a missing interpreter) makes execve return −1 and the caller carries on. Then the old memory is thrown away and the new, almost empty, one installed. After that there is nothing to return to, so a failure kills the process with SIGSEGV (try the broken PT_LOAD tab). This is also when close-on-exec files are closed, caught signal handlers are reset (their code is gone), other threads are killed and the process name changes.
"Loading" means mapping
For each PT_LOAD, the kernel creates a VMA. A VMA is only a record, "these addresses show that file from this offset", and no page-table entry exists yet. A position-independent executable (ET_DYN, the default in every major distribution) gets a random load bias added to all its addresses (ASLR); an old-style ET_EXEC runs at the address it was linked for (0x400000). The interpreter is mapped the same way in the mmap area near the top of user space, then the kernel maps the vDSO, a tiny library that lets clock_gettime run without a system call.
The initial stack and the auxiliary vector
The only thing a new program receives is its stack pointer. At RSP the ABI puts:
RSP → argc
argv[0] … argv[argc-1], NULL
envp[0] … NULL
auxv: (AT_PHDR, …) (AT_ENTRY, …) (AT_BASE, …) (AT_RANDOM, …) … (AT_NULL, 0)
16 random bytes, then the argument and environment strings (higher addresses)
The auxiliary vector is how the kernel talks to the loader: where the program's headers are (AT_PHDR), where it starts (AT_ENTRY), where the loader itself was put (AT_BASE), the vDSO, the page size, the user IDs, CPU features, and 16 random bytes that seed the stack canary. See it with LD_SHOW_AUXV=1 ./hello. Then start_thread overwrites the user registers the syscall saved: RIP to the entry point, RSP to argc. execve "returns" 0 into a different program.
Demand paging and the page cache
The first instruction fetch finds no page-table entry: a page fault. The handler looks up the VMA, sees it maps a file, and asks the page cache for that page. If it is there, the fault is minor: install a PTE pointing at the cached page and retry the instruction. If not, it is major: the block layer reads it from disk, together with a readahead window of the neighbouring pages (32 here), because programs rarely touch just one page.
- Code and read-only data are mapped shared: every process running
libcuses the same physical pages. This is whyld.soandlibcare never read from disk in the animation: some other process always has them cached. - A write to a private file mapping (relocating a GOT, zeroing the bss tail) triggers copy-on-write: the process gets its own copy, and the file and the cache are untouched.
- A 16 KiB program is read whole by the first readahead. For a 780 KiB static binary (static-pie, cold), faults on pages far from the header are major faults. Choose warm to see a second run with no disk access at all.
The animation draws a handful of the faults. A real run of this hello takes about 80 minor faults (about 55 for the static build), measured with /usr/bin/time -f "%R minor %F major".
See also: paging and page replacement, which draws the page tables and the TLB, and follows faults, swapping and copy-on-write in detail.
What happens to the stack and the brk heap once main runs: Stack vs Heap in C.
The dynamic loader
ld-linux-x86-64.so.2 starts with nothing relocated, not even itself. In order it:
- relocates itself (
_dl_start), using only position-independent code; - reads the program's
PT_DYNAMICthroughAT_PHDR:DT_NEEDED libc.so.6, relocation tables, flags; - finds the libraries:
LD_PRELOADand/etc/ld.so.preload, thenDT_RPATH,LD_LIBRARY_PATH,DT_RUNPATH,/etc/ld.so.cache(built byldconfig) and the default directories.LD_DEBUG=libs ./helloshows the search; - maps them with
mmap: one reservation for the whole library, then one fixed mapping per segment; - applies relocations, dependencies first:
R_X86_64_RELATIVEadds the load bias,GLOB_DATandJUMP_SLOTlook symbols up by name (with GNU hash tables) and write their addresses into the GOT; - sets up TLS (thread-local storage, e.g.
errno) and points the FS register at the thread descriptor; - applies RELRO:
mprotectmakes the GOT and other startup-only data read-only; - runs the libraries' initialisers and jumps to
AT_ENTRY, the program's_start.
GOT, PLT and lazy binding
Code in a shared object cannot contain absolute addresses of other libraries' functions, because it is shared and read-only. A call to printf goes to a small stub in the PLT, which jumps through a slot in the GOT, a writable table the loader fills. With lazy binding the slot at first points back into the PLT, so the first call lands in _dl_runtime_resolve, which looks printf up and patches the slot; later calls go straight to libc. With -z now (BIND_NOW, the default on Ubuntu and Fedora; Debian's gcc leaves it off) every symbol is resolved before main. Startup costs a little more, missing symbols fail at start instead of mid-run, and together with RELRO the whole GOT can be made read-only ("full RELRO").
_start, __libc_start_main, main
_start comes from crt1.o: it clears the frame pointer, takes argc and argv off the stack and calls __libc_start_main(main, argc, argv, …). That function saves the environment, initialises stdio, runs the program's constructors (.init_array), registers the destructors, and finally calls main. When main returns, exit() runs atexit handlers and destructors, flushes stdio (our printf became a write(1, "hello world\n", 12)), and calls exit_group.
Static vs dynamic
| dynamic (PIE) | static-pie | static (no PIE) | |
|---|---|---|---|
| File size (this program) | 16 KiB | 777 KiB | 737 KiB |
| First user instruction | in ld.so | the program's _start | the program's _start |
| Who relocates | ld.so | the program itself | nobody (fixed addresses) |
| ASLR of the program | yes | yes | no |
System calls before main (strace) | ≈ 25 | ≈ 12 | ≈ 12 |
| Shares libc pages with other processes | yes | no | no |
| Gets a libc security fix | when libc is updated | only when rebuilt | only when rebuilt |
See it on your own machine
strace -f ./hello world # every system call, from execve to exit_group LD_SHOW_AUXV=1 ./hello world # the auxiliary vector LD_DEBUG=libs,bindings ./hello world # library search and symbol binding LD_DEBUG=statistics ./hello world # relocation counts and time spent in ld.so readelf -hlW hello; readelf -dW hello # header, program headers, dynamic section cat /proc/self/maps # the address space of cat itself setarch -R cat /proc/self/maps # the same with ASLR off: same addresses every time /usr/bin/time -f "%R minor %F major" ./hello world
Common surprises
- "No such file or directory" for a file that exists: the missing file is the interpreter in
PT_INTERP, e.g. a glibc binary copied into an Alpine (musl) container. Newer bash says "cannot execute: required file not found". ETXTBSY, "Text file busy": writing a binary that is running, or running one that is open for writing.noexecmounts (often/tmp) refuseexecvewithEACCESand refusemmapwithPROT_EXEC, butsh script.shstill runs a script from there: the shell only reads it as data.- setuid programs run in "secure mode" (
AT_SECURE): ld.so ignoresLD_PRELOADandLD_LIBRARY_PATH. - A crash before
mainis usually in ld.so (a missing symbol, "version GLIBC_2.38 not found") or a constructor.LD_DEBUGshows which.