The idea: two programs, two times

Running a Java program takes two programs. At compile time, javac reads Calc.java and writes Calc.class: not machine code, but bytecode for a virtual stack machine, plus a constant pool of names. At run time, java starts a JVM, which loads the class file, checks it, and interprets the bytecode one instruction at a time (and compiles the hot parts to machine code later).

Compile time: javac runs Calc.java through scan, parse, enter, attribute, flow, desugar and gen and writes Calc.class with bytecode and a constant pool. Run time: java starts the JVM, the class loader reads Calc.class, it is verified, the bytecode is interpreted and the program prints 14
javac turns source into portable bytecode once; the JVM loads, verifies and interprets that bytecode every time the program runs.

The same road for C, where the compiler writes machine code and a linker joins the files: How a C Program Is Compiled, Linked, Loaded and Started.

The animation follows this program through both. The token counts, tree node kinds, phase names, the bytecode, the constant pool numbers, the class file sizes, the class counts and the error messages come from real runs of javac and java 21.0.8.

public class Calc {
    static int square(int x) {
        return x * x;
    }

    public static void main(String[] args) {
        int sum = 0;
        for (int i = 1; i <= 3; i++) {
            sum += square(i);
        }
        System.out.println(sum);          // 14
    }
}

Reading the canvas

  • Top: the pipeline. javac's phases, then the JVM's. The current phase is yellow, finished ones green, and the phase that finds an error turns red.
  • Left: the source file, with the current token or line highlighted. During desugar it shows the tree printed back as source, and while the program runs it shows the bytecode of the running method with the interpreter's position (▶).
  • Middle: what the current phase works on: the tokens, the syntax tree (with types after attribute, ✓ after flow), the list of desugarings, the bytecode being generated with the stack depth after each instruction, the class file layout, and at run time the JVM stack: one frame per active method, each with its local variables and its operand stack.
  • Right: the symbol table, then the constant pool (red: just added; grey: symbolic; green ✓: resolved at run time), the heap and the terminal.

Step does one token, tree node, instruction or bytecode. Next Stage finishes the current phase. Run to .class stops when javac is done. Each tab loads one program: Calc, Shapes, Sugar, and four broken versions of Calc. Its Demo buttons run it to the end, and Reset starts it again from its source.

1. Scan: characters to tokens

The scanner (lexer) groups characters into tokens: keywords (public, int, for), identifiers (Calc, sum, System), literals (0, "total ") and operators and separators ({, <=, +=). Whitespace and comments are dropped. <= is one token, LTEQ, because the scanner always takes the longest match. Calc is 68 tokens. The scanner knows nothing about grammar: int sum = 0 without a semicolon scans fine.

2. Parse: tokens to a syntax tree

The parser reads the tokens by the grammar and builds the abstract syntax tree (AST): a class contains methods, a method a block of statements, a statement expressions. Precedence decides the shape (x * x + y * y is a PLUS of two MULTIPLYs), and then parentheses and semicolons are no longer needed. The node kinds on the canvas are javac's own (com.sun.source.tree.Tree.Kind). A missing semicolon is found here: the Syntax error tab.

3. Enter: symbols

Enter walks the declarations and creates a symbol for every class, field and method, in a scope, before any method body is checked. That is why main can call square no matter which comes first in the file. A class without a constructor gets the default constructor here: public Calc() { super(); }. It is in the class file even though nobody wrote it.

4. Attribute: types and names

Attribute is the type checker. It gives every expression a type and every name a meaning: x is a parameter, System is the class java.lang.System, out is a static field of type PrintStream, and println(sum) is println(int), chosen from 10 overloads by the argument's type. It infers generic type arguments and lambda parameter types (x -> … gets its type from what forEach expects). To do this it reads the JDK's classes: javac -verbose Calc.java lists 96 of them.

Edit to CalcFound byjavac says
int sum = 0 (no ;)parserCalc.java:7: error: ';' expected
sqare(i)attributeerror: cannot find symbol / symbol: method sqare(int)
int sum = "0";attributeerror: incompatible types: String cannot be converted to int
int sum;flowerror: variable sum might not have been initialized (twice)

With -XDverboseCompilePolicy, javac prints each phase as it starts: [attribute Calc], [flow Calc], [desugar Calc], [generate code Calc]. After an error it stops at the end of that phase: a file with a type error never reaches flow, and no class file is written.

5. Flow: every path

Flow analysis follows every path through each method. A local variable must be definitely assigned before it is read (fields get default values, locals never do), every statement must be reachable, a non-void method must return on every path, a final field must be assigned exactly once, and checked exceptions must be caught or declared. The Uninit var tab.

6. Desugar: what the JVM does not have

Much of the Java language is not in the JVM. javac rewrites it into plain classes, methods and fields (the Sugar tab):

You writeThe class file hasjavac pass
List<Integer> xsraw List, and checkcast Integer where values come outTransTypes (erasure)
x -> System.out.println(x * 2)a method private static synthetic lambda$main$0(Integer) and an invokedynamic (bootstrap LambdaMetafactory)LambdaToMethod
List.of(1, 2, 3), x * 2Integer.valueOf(1) …, x.intValue() * 2Lower (boxing)
for (int x : xs)iterator(), hasNext(), next(), cast, intValue()Lower
static class Point inside Shapesa separate Shapes$Point.class with NestHost/NestMembersLower
"total " + totalinvokedynamic makeConcatWithConstants, recipe "total \u0001"Gen (Java 9+)

javac -printsource -d out Sugar.java prints the tree after erasure and lambda translation as source; javap -c -p shows the rest.

7. Gen: bytecode for a stack machine

The JVM is a stack machine: an instruction takes its operands from the operand stack and pushes its result. sum += square(i) becomes:

 9: iload_1                 // push sum
10: iload_2                 // push i
11: invokestatic  #7        // Calc.square:(I)I   pops i, pushes the result
14: iadd                    // pops two, pushes the sum
15: istore_1                // pop into sum
Five snapshots of the operand stack with sum = 1 and i = 2: iload_1 pushes 1, iload_2 pushes 2, invokestatic #7 replaces 2 by square(2) = 4, iadd leaves 5, istore_1 pops 5 into local sum
Every bytecode takes its operands from the top of the operand stack and pushes its result back; locals live in numbered slots.

While emitting, javac computes each method's max_stack (2 for main) and max_locals (3: args, sum, i; local variables are numbered slots, their names are not kept unless you compile with -g). A loop condition becomes a jump on its opposite (i <= 3 → if_icmpgt 22, leave when i > 3); the target is not known when the jump is emitted, so javac leaves a hole and patches it. Branch targets get a StackMapTable entry, which lets the verifier check the method in one pass.

Instructions do not contain addresses. invokestatic #7 points at constant pool entry 7, Methodref Calc.square:(I)I, which is made of more entries: a Class, a NameAndType and Utf8 strings. javac numbers entries in the order it first needs them. Calc's pool has 31 entries; Sugar's has 102.

invokestatic #7 points to constant pool entry #7 Methodref, made of #8 Class (pointing to Utf8 Calc at #10) and #9 NameAndType (pointing to Utf8 square at #11 and (I)I at #12)
Bytecode refers to methods by constant-pool number; the names are only turned into an address at run time.

8. The class file

BytesCalc.class (516 bytes)
CA FE BA BEmagic number
00 00 00 41minor 0, major 65 = Java 21 (an older JVM refuses it: UnsupportedClassVersionError)
00 20 …constant pool count 32: 31 entries, 295 bytes
00 21 00 08 00 02flags public + super, this class #8, superclass #2
00 00 00 00no interfaces, no fields
00 03 …3 methods (<init>, square, main), 191 bytes, each with a Code attribute
00 01 …1 attribute: SourceFile "Calc.java"

One class file per class: Shapes produces Shapes.class (522 bytes) and Shapes$Point.class (400 bytes). The same file runs on every OS and CPU.

9. Running: load, link, interpret

java Calc starts a JVM (How a Java Application Starts follows that in detail), which loads about 650 JDK classes and then Calc: the application class loader finds Calc.class on the class path, the JVM parses and verifies it (a class file may come from anywhere, not only from javac), and calls main. Classes load when first used: Shapes$Point is loaded only when new first runs.

Each method call gets a frame on the thread's stack, with max_locals slots and room for max_stack values. invokestatic pops the arguments into the new frame's first locals; ireturn pops the frame and pushes the result on the caller's operand stack. The first time an instruction uses a constant pool entry, the JVM resolves it (finds the class, the method, the field's offset) and caches the result.

Calc interprets 53 bytecodes: 41 in main and 12 in three calls to square. The JIT compilers wait for a method to be called about 200 times (HotSpot JIT Tiers), so nothing here is compiled. Compiling costs more than running for a program this small: javac Calc.java takes about 260 ms, java Calc about 44 ms. java Calc.java (Java 11+) does both in one command, compiling in memory.

See it on your own machine

javac -XDverboseCompilePolicy Calc.java   # the phases: [attribute Calc] [flow Calc] …
javac -verbose Calc.java                  # every JDK class it reads, [wrote Calc.class]
javac -printsource -d out Sugar.java      # the tree after erasure and lambdas, as source
javap -c -p -v Calc.class                 # constant pool, bytecode, max_stack, StackMapTable
java -Xlog:class+load Calc                # classes as they load
java Calc.java                            # compile in memory and run

What the page leaves out

Annotation processing (which runs between enter and attribute and can generate new sources), modules and module-info.java, incremental and parallel compilation, javac's error recovery (it keeps parsing after a syntax error to report more), constant folding and the finer points of overload resolution and type inference. At run time, the JDK methods (println, List.of, the iterator, the invokedynamic bootstraps) are shown as single steps, and the interpreter is shown as a loop over instructions; HotSpot's is a template interpreter, machine code generated at start-up for every bytecode.