The idea: two programs, two times
Running a Java program takes two programs. At compile time, javac reads Calc.java and writes Calc.class: not machine code, but bytecode for a virtual stack machine, plus a constant pool of names. At run time, java starts a JVM, which loads the class file, checks it, and interprets the bytecode one instruction at a time (and compiles the hot parts to machine code later).
The same road for C, where the compiler writes machine code and a linker joins the files: How a C Program Is Compiled, Linked, Loaded and Started.
The animation follows this program through both. The token counts, tree node kinds, phase names, the bytecode, the constant pool numbers, the class file sizes, the class counts and the error messages come from real runs of javac and java 21.0.8.
public class Calc {
static int square(int x) {
return x * x;
}
public static void main(String[] args) {
int sum = 0;
for (int i = 1; i <= 3; i++) {
sum += square(i);
}
System.out.println(sum); // 14
}
}
Reading the canvas
- Top: the pipeline. javac's phases, then the JVM's. The current phase is yellow, finished ones green, and the phase that finds an error turns red.
- Left: the source file, with the current token or line highlighted. During desugar it shows the tree printed back as source, and while the program runs it shows the bytecode of the running method with the interpreter's position (▶).
- Middle: what the current phase works on: the tokens, the syntax tree (with types after attribute, ✓ after flow), the list of desugarings, the bytecode being generated with the stack depth after each instruction, the class file layout, and at run time the JVM stack: one frame per active method, each with its local variables and its operand stack.
- Right: the symbol table, then the constant pool (red: just added; grey: symbolic; green ✓: resolved at run time), the heap and the terminal.
Step does one token, tree node, instruction or bytecode. Next Stage finishes the current phase. Run to .class stops when javac is done. Each tab loads one program: Calc, Shapes, Sugar, and four broken versions of Calc. Its Demo buttons run it to the end, and Reset starts it again from its source.
1. Scan: characters to tokens
The scanner (lexer) groups characters into tokens: keywords (public, int, for), identifiers (Calc, sum, System), literals (0, "total ") and operators and separators ({, <=, +=). Whitespace and comments are dropped. <= is one token, LTEQ, because the scanner always takes the longest match. Calc is 68 tokens. The scanner knows nothing about grammar: int sum = 0 without a semicolon scans fine.
2. Parse: tokens to a syntax tree
The parser reads the tokens by the grammar and builds the abstract syntax tree (AST): a class contains methods, a method a block of statements, a statement expressions. Precedence decides the shape (x * x + y * y is a PLUS of two MULTIPLYs), and then parentheses and semicolons are no longer needed. The node kinds on the canvas are javac's own (com.sun.source.tree.Tree.Kind). A missing semicolon is found here: the Syntax error tab.
3. Enter: symbols
Enter walks the declarations and creates a symbol for every class, field and method, in a scope, before any method body is checked. That is why main can call square no matter which comes first in the file. A class without a constructor gets the default constructor here: public Calc() { super(); }. It is in the class file even though nobody wrote it.
4. Attribute: types and names
Attribute is the type checker. It gives every expression a type and every name a meaning: x is a parameter, System is the class java.lang.System, out is a static field of type PrintStream, and println(sum) is println(int), chosen from 10 overloads by the argument's type. It infers generic type arguments and lambda parameter types (x -> … gets its type from what forEach expects). To do this it reads the JDK's classes: javac -verbose Calc.java lists 96 of them.
| Edit to Calc | Found by | javac says |
|---|---|---|
int sum = 0 (no ;) | parser | Calc.java:7: error: ';' expected |
sqare(i) | attribute | error: cannot find symbol / symbol: method sqare(int) |
int sum = "0"; | attribute | error: incompatible types: String cannot be converted to int |
int sum; | flow | error: variable sum might not have been initialized (twice) |
With -XDverboseCompilePolicy, javac prints each phase as it starts: [attribute Calc], [flow Calc], [desugar Calc], [generate code Calc]. After an error it stops at the end of that phase: a file with a type error never reaches flow, and no class file is written.
5. Flow: every path
Flow analysis follows every path through each method. A local variable must be definitely assigned before it is read (fields get default values, locals never do), every statement must be reachable, a non-void method must return on every path, a final field must be assigned exactly once, and checked exceptions must be caught or declared. The Uninit var tab.
6. Desugar: what the JVM does not have
Much of the Java language is not in the JVM. javac rewrites it into plain classes, methods and fields (the Sugar tab):
| You write | The class file has | javac pass |
|---|---|---|
List<Integer> xs | raw List, and checkcast Integer where values come out | TransTypes (erasure) |
x -> System.out.println(x * 2) | a method private static synthetic lambda$main$0(Integer) and an invokedynamic (bootstrap LambdaMetafactory) | LambdaToMethod |
List.of(1, 2, 3), x * 2 | Integer.valueOf(1) …, x.intValue() * 2 | Lower (boxing) |
for (int x : xs) | iterator(), hasNext(), next(), cast, intValue() | Lower |
static class Point inside Shapes | a separate Shapes$Point.class with NestHost/NestMembers | Lower |
"total " + total | invokedynamic makeConcatWithConstants, recipe "total \u0001" | Gen (Java 9+) |
javac -printsource -d out Sugar.java prints the tree after erasure and lambda translation as source; javap -c -p shows the rest.
7. Gen: bytecode for a stack machine
The JVM is a stack machine: an instruction takes its operands from the operand stack and pushes its result. sum += square(i) becomes:
9: iload_1 // push sum 10: iload_2 // push i 11: invokestatic #7 // Calc.square:(I)I pops i, pushes the result 14: iadd // pops two, pushes the sum 15: istore_1 // pop into sum
While emitting, javac computes each method's max_stack (2 for main) and max_locals (3: args, sum, i; local variables are numbered slots, their names are not kept unless you compile with -g). A loop condition becomes a jump on its opposite (i <= 3 → if_icmpgt 22, leave when i > 3); the target is not known when the jump is emitted, so javac leaves a hole and patches it. Branch targets get a StackMapTable entry, which lets the verifier check the method in one pass.
Instructions do not contain addresses. invokestatic #7 points at constant pool entry 7, Methodref Calc.square:(I)I, which is made of more entries: a Class, a NameAndType and Utf8 strings. javac numbers entries in the order it first needs them. Calc's pool has 31 entries; Sugar's has 102.
8. The class file
| Bytes | Calc.class (516 bytes) |
|---|---|
CA FE BA BE | magic number |
00 00 00 41 | minor 0, major 65 = Java 21 (an older JVM refuses it: UnsupportedClassVersionError) |
00 20 … | constant pool count 32: 31 entries, 295 bytes |
00 21 00 08 00 02 | flags public + super, this class #8, superclass #2 |
00 00 00 00 | no interfaces, no fields |
00 03 … | 3 methods (<init>, square, main), 191 bytes, each with a Code attribute |
00 01 … | 1 attribute: SourceFile "Calc.java" |
One class file per class: Shapes produces Shapes.class (522 bytes) and Shapes$Point.class (400 bytes). The same file runs on every OS and CPU.
9. Running: load, link, interpret
java Calc starts a JVM (How a Java Application Starts follows that in detail), which loads about 650 JDK classes and then Calc: the application class loader finds Calc.class on the class path, the JVM parses and verifies it (a class file may come from anywhere, not only from javac), and calls main. Classes load when first used: Shapes$Point is loaded only when new first runs.
Each method call gets a frame on the thread's stack, with max_locals slots and room for max_stack values. invokestatic pops the arguments into the new frame's first locals; ireturn pops the frame and pushes the result on the caller's operand stack. The first time an instruction uses a constant pool entry, the JVM resolves it (finds the class, the method, the field's offset) and caches the result.
Calc interprets 53 bytecodes: 41 in main and 12 in three calls to square. The JIT compilers wait for a method to be called about 200 times (HotSpot JIT Tiers), so nothing here is compiled. Compiling costs more than running for a program this small: javac Calc.java takes about 260 ms, java Calc about 44 ms. java Calc.java (Java 11+) does both in one command, compiling in memory.
See it on your own machine
javac -XDverboseCompilePolicy Calc.java # the phases: [attribute Calc] [flow Calc] … javac -verbose Calc.java # every JDK class it reads, [wrote Calc.class] javac -printsource -d out Sugar.java # the tree after erasure and lambdas, as source javap -c -p -v Calc.class # constant pool, bytecode, max_stack, StackMapTable java -Xlog:class+load Calc # classes as they load java Calc.java # compile in memory and run
What the page leaves out
Annotation processing (which runs between enter and attribute and can generate new sources), modules and module-info.java, incremental and parallel compilation, javac's error recovery (it keeps parsing after a syntax error to report more), constant folding and the finer points of overload resolution and type inference. At run time, the JDK methods (println, List.of, the iterator, the invokedynamic bootstraps) are shown as single steps, and the interpreter is shown as a loop over instructions; HotSpot's is a template interpreter, machine code generated at start-up for every bytecode.