The idea: four programs and a driver

Typing gcc -O2 main.c util.c -o sum runs four programs, and running ./sum runs two more before your main. The page follows a two-file program the whole way, with the real output of gcc 14.2, GNU ld 2.44 and glibc 2.41 on x86-64 Linux:

// util.h                    // util.c                         // main.c
#ifndef UTIL_H               #include "util.h"                 #include <stdio.h>
#define UTIL_H               int calls;                        #include "util.h"
#define SQUARE(x) ((x) * (x))  int sum_squares(int n) {        int limit = 3;
extern int calls;                calls++;                      int main(void) {
int sum_squares(int n);          int s = 0;                        int r = sum_squares(limit);
#endif                           for (int i = 1; i <= n; i++)      printf("sum = %d, calls = %d\n", r, calls);
                                     s += SQUARE(i);               return 0;
                                 return s;                     }
                             }

It prints sum = 14, calls = 1. On the way:

  1. Preprocess (cc1 -E): #include and #define are text operations. main.c grows from 8 lines to 824.
  2. Compile (cc1): one file at a time, to assembly. The compiler only needs declarations of what the file uses from elsewhere.
  3. Assemble (as): assembly to machine code in an object file, with a hole wherever an address is not known yet.
  4. Link (ld): put the object files and the C library together, decide every address, fill every hole.
  5. Load (execve, then ld.so): map the file into memory and, for a dynamic program, link in the shared libraries.
  6. Start: _start → __libc_start_main → main.
Build time: main.c and util.c each go through cc1 -E to .i, cc1 to .s and as to .o; ld joins main.o, util.o, the start files and libc.so.6 into the executable sum. Run time: execve maps sum, ld.so maps libc.so.6, then _start, __libc_start_main and main run
Each .c file is preprocessed, compiled and assembled on its own; ld is the first tool that sees the whole program, and ld.so finishes the linking at run time.

gcc itself is only the driver that runs the others; gcc -v prints their command lines. You can stop it after any step: gcc -E (preprocessed .i), gcc -S (assembly .s), gcc -c (object .o). gcc -save-temps keeps all the intermediate files.

This is the top of the stack of abstraction levels: the compiler and linker turn a program into instructions of the architecture. Levels of abstraction follows one statement from C all the way down to transistors and electrons.

Reading the canvas

The strip at the top is the pipeline: one row per source file (main.c, util.c), the tools between the files they read and write, and on the right the run time. A box turns yellow while it works, green when done, red when it fails. Below it three panels show what the current stage reads, what it writes and the tables it keeps (symbol tables, relocations, the address space). The colours stay the same everywhere: blue for what comes from main.c, orange for util.c, grey for the start files gcc adds, purple for libc, light green for what the linker makes itself, pink for a hole, green for a filled hole or a resolved symbol.

Step moves one small step, Next Stage a whole stage, Run to ./sum stops when the executable is written. The program tabs load the correct program and five broken ones; link switches between a dynamic and a static (-static) link. Reset starts the open program again.

1. Preprocess: text in, text out

The preprocessor knows nothing about C. It follows the # lines:

  • #include <stdio.h> pastes the file in. Angle brackets search the system directories (/usr/lib/gcc/x86_64-linux-gnu/14/include, /usr/local/include, /usr/include/x86_64-linux-gnu, /usr/include); quotes look next to the including file first. stdio.h includes 25 more headers, 27 in all. Of all that, main.c uses one line: extern int printf (const char *__restrict __format, ...);
  • #define lines are remembered and disappear. A later SQUARE(i) is replaced by ((i) * (i)). The parentheses matter: with #define SQUARE(x) x * x, SQUARE(a + 1) would become a + 1 * a + 1.
  • #ifndef UTIL_H … #endif is a header guard: a second #include "util.h" in the same file adds nothing. It does not help across files: util.c reads util.h again from scratch.
  • Lines like # 3 "main.c" 2 are line markers. They let the compiler report main.c:4 for an error it found on line 820 of main.i.

The result for main.c: 824 lines, 21 474 bytes, 314 of them code. A missing header is fatal right here, before any C is parsed: Demo: missing header → preprocessor error.

2. Compile: one file at a time

cc1 parses the .i file into a syntax tree and checks the types. Then it lowers the tree to GIMPLE, simple statements with at most one operation each (calls.0_1 = calls; _2 = calls.0_1 + 1; calls = _2;), puts it in SSA form (every temporary is assigned once; PHI nodes merge the values coming from different branches), and runs about 200 optimisation passes over it. Last, it picks x86-64 instructions and registers and writes assembly. gcc -fdump-tree-all shows every pass.

Each .c file is a separate translation unit. When cc1 compiles main.c it has never seen util.c. It knows from util.h that sum_squares takes an int and returns an int, so it can pass limit in %edi and take the result from %eax (the System V calling convention). But it does not know where sum_squares will be. So it writes a name: call sum_squares@PLT, movl calls(%rip), %edx.

The optimiser changes a lot. Here is sum_squares with and without -O2:

-O0 (variables on the stack)          -O2 (variables in registers)
sum_squares:                          sum_squares:
    pushq   %rbp                          addl    $1, calls(%rip)
    movq    %rsp, %rbp                    testl   %edi, %edi
    movl    %edi, -20(%rbp)               jle     .L4
    movl    calls(%rip), %eax             addl    $1, %edi
    addl    $1, %eax                      movl    $1, %eax
    movl    %eax, calls(%rip)             xorl    %edx, %edx
    movl    $0, -4(%rbp)              .L3:
    movl    $1, -8(%rbp)                  movl    %eax, %ecx
    jmp     .L2                           imull   %eax, %ecx
.L3:                                      addl    $1, %eax
    movl    -8(%rbp), %eax                addl    %ecx, %edx
    imull   %eax, %eax                    cmpl    %edi, %eax
    addl    %eax, -4(%rbp)                jne     .L3
    addl    $1, -8(%rbp)                  movl    %edx, %eax
.L2:                                      ret
    movl    -8(%rbp), %eax            .L4:
    cmpl    -20(%rbp), %eax               xorl    %edx, %edx
    jle     .L3                           movl    %edx, %eax
    movl    -4(%rbp), %eax                ret
    popq    %rbp
    ret

Errors in the C itself are found here: a missing semicolon by the parser, an undeclared function by the type checker (an error since GCC 14; older versions only warned and the mistake turned into a linker error later). Demo: missing ; → compiler error, Demo: typo → implicit declaration error. gcc still compiles the other files, but it does not link.

The listings on the page were made with -fno-asynchronous-unwind-tables -fcf-protection=none, which only removes the .cfi_* unwind directives and the endbr64 instructions, to keep them short.

3. Assemble: bytes, symbols and holes

as encodes each instruction and writes an object file (main.o, ELF type REL). It has three parts that matter:

  • Sections: .text (code), .data (initialised variables), .bss (zeroed variables: only a size, no bytes), .rodata (constants, strings). Every section starts at offset 0.
  • The symbol table (nm, readelf -s): each name and where it is, like main = .text.startup + 0. Names the file uses but does not define are UND: in main.o, sum_squares, calls and printf.
  • Relocations (readelf -r, objdump -dr): wherever an instruction needs an address the assembler cannot know, it writes 00 00 00 00 and an entry like "at offset 0x6, R_X86_64_PC32, limit − 4". main.o has 5 such holes, util.o 1.

Why − 4? x86-64 code addresses data relative to the instruction pointer, and when the CPU adds the 4-byte displacement, %rip already points at the next instruction, 4 bytes after the start of the hole. In addl $1, calls(%rip) one more byte (the constant 1) follows the hole, so the addend is − 5. Jumps inside one section (jle .L4, jne .L3) need no relocation: the assembler knows the distance, and the linker never splits a section.

The linker reads its inputs in command-line order. Besides your objects, gcc passes the start files: Scrt1.o (with _start, the real entry point), crti.o and crtbeginS.o before your code, crtendS.o and crtn.o after, and the libraries -lgcc -lgcc_s -lc. -lc finds /usr/lib/x86_64-linux-gnu/libc.so, which is a small text file, a linker script: GROUP ( libc.so.6 libc_nonshared.a AS_NEEDED ( ld-linux-x86-64.so.2 ) ).

ld keeps one global symbol table. A definition resolves the earlier uses of its name; a use adds an undefined entry. Two rules:

  • When the inputs are used up, every strong undefined symbol is an error: undefined reference to `sum_squares'. ld even says where: main.c:(.text.startup+0xb) is the hole at offset 0xb. Demo: forgot util.c → undefined reference. (Weak symbols, like the profiler hook __gmon_start__, may stay undefined: they get the value 0.)
  • Two strong definitions of one name is an error: multiple definition of `calls'. The classic cause is a definition in a header, int calls = 0; instead of extern int calls;: every file that includes it defines its own calls. Demo: variable defined in a header → multiple definition. Since GCC 10 (-fno-common by default) this is also an error for int calls; without a value.

Order matters for archives (.a files). An archive is a bundle of object files with an index; ld takes only the members that define something undefined at that point, and each member it takes can make more symbols undefined. That is why -lm must come after the files that use it.

ld concatenates the input sections of each kind into output sections, in the order of its built-in linker script (ld --verbose prints it), and gives them addresses. -Wl,-Map=sum.map writes the result: main.o's .text.startup at 0x1050 (the script puts .text.startup first), _start at 0x1080, sum_squares at 0x1170, the format string at 0x2004, limit at 0x4018, calls at 0x4020. Code, read-only data and writable data go on separate pages so that each can get its own permissions.

Now every address is known and ld fills each hole with S + A − P: S the symbol's address, A the addend from the relocation, P the address of the hole. For mov limit(%rip),%edi at 0x1054: S = 0x4018, A = −4, P = 0x1056, so the hole gets 0x2fbe, stored little-endian as be 2f 00 00. The CPU later computes 0x105a + 0x2fbe = 0x4018. Demo: the holes and how ld fills them works out all six; the page computes them from the tables and they match objdump -d sum byte for byte.

In main.o the instruction mov limit(%rip),%edi is 8b 3d followed by a 4-byte hole of zeros with relocation R_X86_64_PC32 limit − 4; ld computes 0x4018 − 4 − 0x1056 = 0x2fbe and writes be 2f 00 00 at 0x1056 in sum; at run time 0x105a + 0x2fbe reaches limit at 0x4018
The assembler leaves zeros where an address goes; ld writes the distance S + A − P there, so the CPU lands on the variable.
Hole (in sum)InstructionSAPS + A − PBytes
0x1056mov limit(%rip),%edi0x4018−40x10560x2fbebe 2f 00 00
0x105bcall sum_squares0x1170−40x105b0x11111 01 00 00
0x1061mov calls(%rip),%edx0x4020−40x10610x2fbbbb 2f 00 00
0x1068lea .LC0(%rip),%rdi0x2004−40x10680xf9898 0f 00 00
0x1071call printf@plt0x1030−40x1071−0x45bb ff ff ff
0x1172addl $1,calls(%rip)0x4020−50x11720x2ea9a9 2e 00 00

sum_squares has a PLT32 relocation, but it is defined in the link, so the call goes straight to it. printf is in a shared library, so the call goes to a small stub, printf@plt, and ld adds what ld.so needs to finish the job at run time: NEEDED libc.so.6 in the .dynamic section and 9 dynamic relocations (3 RELATIVE for pointers stored in data, 5 GLOB_DAT GOT slots, 1 JUMP_SLOT for printf). Last, it writes the program headers: four LOAD segments (R, R E, R, RW), INTERP, DYNAMIC, GNU_RELRO. The kernel reads only those; sections are for the linker and the debugger.

Static vs dynamic linking

dynamic (the default)static (-static)
libc comes fromlibc.so.6, at run timelibc.a: 431 archive members copied in
Size of sum16 088 bytes754 416 bytes
AddressPIE: anywhere (load bias, ASLR)fixed at 0x400000
Call to printfvia printf@plt and a GOT slotdirect: call 4047a0
Start-upkernel → ld.so (loads libc, 90 relocations) → _startkernel → _start
Fix a bug in libcupdate libc.so.6, every program gets itrebuild every program
Memorylibc's code pages shared by all processeseach program has its own copy

Demo: static link shows the archive at work: printf pulls in printf.o, which needs vfprintf-internal.o, which needs memcpy, strlen, the locale code… until nothing is undefined.

6. Load and start

./sum: the shell forks and the child calls execve. The kernel reads the program headers, maps each LOAD segment at the load bias (0x555555554000 with ASLR off; random otherwise), zero-fills .bss, sees PT_INTERP and maps /lib64/ld-linux-x86-64.so.2 as well. It builds the stack (argc, argv, the environment and the auxiliary vector with AT_ENTRY, AT_BASE …) and returns to user mode in ld.so. Nothing of the file is read yet; pages come in on first use. How Linux Loads a Program shows this part in detail: the system calls, the page faults, the page cache.

ld.so is the last step of linking. It reads NEEDED libc.so.6, finds the file through /etc/ld.so.cache, maps it, applies the relocations of libc and of sum (90 in all; LD_DEBUG=statistics prints the count), makes the RELRO pages read-only, runs the initialisers and jumps to _start. _start, a few lines of assembly from Scrt1.o, takes argc and argv from the stack and calls __libc_start_main(main, …), which calls your main. When main returns, __libc_start_main calls exit: the atexit handlers run, stdout is flushed, and exit_group ends the process.

7. The first printf call: PLT and GOT

Debian's gcc does not pass -z now, so printf is bound lazily. At start-up ld.so sets printf's GOT slot (0x4000) to an address inside printf@plt. The first call runs:

call printf@plt          ; main
jmp  *GOT[0x4000]        ; printf@plt: the slot points at the next line
push $0                  ; which JUMP_SLOT relocation
jmp  PLT0
push GOT[1]              ; PLT0: ld.so's handle for sum
jmp  *GOT[2]             ; _dl_runtime_resolve → _dl_fixup
                         ; finds printf in libc.so.6, writes it into GOT[0x4000], jumps there
First call: main calls printf@plt, which jumps through GOT[0x4000]; the slot still points back into printf@plt, which pushes 0 and jumps to PLT0, ld.so _dl_fixup finds printf, writes its address into GOT[0x4000] and jumps there. Later calls: printf@plt jumps through GOT[0x4000] straight to printf in libc.so.6
Only the first printf call goes through ld.so; it patches the GOT slot so every later call is one indirect jump.

Every later call is a single indirect jump to libc. LD_BIND_NOW=1 or linking with -z now resolves everything at start-up instead; with -z relro -z now (full RELRO) the whole GOT is then made read-only. Demo: the first printf call (lazy binding).

Which stage reports which error

MistakeFound byMessage (gcc 14)
#include "utils.h" (no such file)preprocessorfatal error: utils.h: No such file or directory
int limit = 3 (no ;)compiler, parserexpected ',' or ';' before 'int'
sum_square(limit) (typo)compiler, type checkimplicit declaration of function 'sum_square'
gcc main.c (forgot util.c)linkerundefined reference to `sum_squares'
int calls = 0; in util.hlinkermultiple definition of `calls'
a library missing at run timeld.soerror while loading shared libraries: libfoo.so.1: cannot open shared object file

See it on your own machine

gcc -v -save-temps -O2 main.c util.c -o sum   # the commands; keeps main.i, main.s, main.o
objdump -dr main.o                             # machine code with the relocations
readelf -s main.o ; nm main.o                  # symbol tables (U = undefined)
gcc -O2 main.c util.c -o sum -Wl,-Map=sum.map  # where ld put every section
readelf -lW sum ; readelf -d sum ; readelf -r sum   # segments, NEEDED, dynamic relocations
ldd ./sum                                      # which shared libraries it will get
LD_DEBUG=libs,bindings,statistics ./sum        # ld.so at work

What the page leaves out

Debug information (-g, DWARF sections), link-time optimisation (-flto: the compiler runs again inside the linker), C++ name mangling, constructors and templates, building a shared library (-fPIC -shared) and symbol visibility, thread-local storage, IFUNCs and the IRELATIVE relocations of static programs, linker relaxation, .eh_frame and stack unwinding, the separate cpp program (modern gcc does the preprocessing inside cc1), and other linkers (gold, lld, mold), which do the same three jobs in a different order. Run-time addresses are those of a run with ASLR off.