The idea: four layers between your code and the silicon

When you call new Thread(task).start(), the code in task goes through four layers before it runs:

  1. a Java thread, the java.lang.Thread object;
  2. a Linux thread, a kernel task with its own TID;
  3. a logical CPU (a hardware thread), which is what Linux calls cpu0, cpu1, …;
  4. a physical core, which holds the execution units and the caches.

Each layer maps onto the next in a different way:

  • Java thread to Linux thread is 1:1, and fixed for the thread's whole life.
  • Linux thread to logical CPU is many to few, and the kernel changes it all the time.
  • Logical CPU to physical core is 2:1 with SMT and 1:1 without it, and fixed by the hardware.
Six Java threads, main and worker-1 to worker-5, each map 1:1 to Linux tasks TID 4201 to 4206; the scheduler puts them on four logical CPUs, cpu0, cpu2, cpu1, cpu3, each running one task with two tasks waiting in queues; cpu0 and cpu2 are the two hardware threads of core 0, cpu1 and cpu3 of core 1
A Java thread is fixed to one Linux task, the kernel moves that task between logical CPUs, and two logical CPUs share one physical core.

The canvas draws the layers as stacked bands. Each thread keeps the same colour in every band.

Java thread to Linux thread: start(), pthread_create, clone

new Thread(...) only makes an object on the heap. Its state is NEW and there is no kernel thread yet.

start() goes into the JVM (JVM_StartThread → os::create_thread in HotSpot). The JVM allocates a JavaThread and an OSThread, picks a stack size (-Xss, 1 MiB by default on Linux x64), and calls glibc's pthread_create.

glibc then makes the clone() system call with these flags:

  • CLONE_VM: share the address space;
  • CLONE_FILES: share the open files;
  • CLONE_SIGHAND: share the signal handlers;
  • CLONE_THREAD: join the java process's thread group.

The kernel creates a new task_struct with its own TID. The kernel treats a thread as a task like any other, and so does its scheduler.

The TID is what jstack prints as nid=0x… (in hex) and what top -H -p <pid> lists as a row. To find a hot Java thread, run top -H, convert its TID to hex and search for that nid in the thread dump. (Demo: one thread, all the way down)

The JVM's own threads are tasks too: main, GC Thread#0, C2 CompilerThread0, VM Thread, and others. A "hello world" JVM already has about 20 of them. (Demo: the JVM's own threads are Linux threads)

The kernel scheduler: run queues, time slices, migration

The JVM does not choose where or when a platform thread runs. The Linux scheduler does. Each logical CPU has its own run queue. On the canvas:

  • A new or woken task is placed by select_task_rq. It tries a fully idle core first, then the task's previous CPU (its caches may still be warm), then any idle CPU, then the shortest queue.
  • A task that has used up its time slice is preempted if another task is waiting. This is an involuntary context switch.
  • A CPU whose queue runs dry pulls work from a busy queue. This is migration: the thread continues on another CPU, with cold caches there.

With more runnable threads than logical CPUs, the threads take turns. The timeline shows the slices. (Demo: more threads than CPUs)

The real scheduler (CFS, and EEVDF since Linux 6.6) orders tasks by virtual runtime and balances load across scheduling domains. The page keeps only the shape of that. See CPU Scheduling and Context Switch for more detail.

Logical CPU vs physical core: SMT / Hyper-Threading

A physical core has the execution units: ALUs, FPUs, load/store units, the branch predictor, and the L1 and L2 caches.

With SMT (simultaneous multithreading, which Intel calls Hyper-Threading), each core has two sets of architectural state: registers, a program counter, and so on. Each set is a hardware thread. Linux shows each hardware thread as a logical CPU.

The two hardware threads of one core are siblings. On Intel they are often numbered far apart, like cpu0 and cpu2 here. You can check with lscpu -e or /sys/devices/system/cpu/cpu0/topology/thread_siblings_list.

Siblings share the core's execution units. When one thread stalls on a cache miss, the other can use the idle units. So SMT gives roughly 10–30 % more throughput, not 100 %.

The page models this as 10 work units per tick for a thread alone on its core and 6 for each of two siblings. That makes 12 per core instead of 10. This is also why the scheduler spreads threads over idle cores before it doubles up on siblings. (Demos: SMT on, siblings share a core · SMT off, 4 threads on 2 CPUs)

Threads busySMT on (4 logical CPUs)SMT off (2 logical CPUs)
2cpu0 + cpu1 (different cores): 2 × 10 = 20 units/tickcpu0 + cpu1: 20 units/tick
4all 4 hardware threads: 4 × 6 = 24 units/tick2 run, 2 wait in queues: 20 units/tick

Java thread states are not Linux thread states

Thread.getState()Linux task stateWhere the task is
NEWno task yetonly the Java object exists
RUNNABLERrunning on a CPU, or waiting in a run queue
RUNNABLE (in native read())Ssleeping on a socket's wait queue: the JVM cannot tell
BLOCKED / WAITING / TIMED_WAITINGSsleeping on a futex or a timer, in no run queue
TERMINATEDtask has exitednone

So a thread dump full of RUNNABLE threads does not mean the CPUs are busy. A sleeping thread costs memory (its stack) but no CPU time. When worker-2 calls Thread.sleep, its CPU is free at once, and an idle CPU pulls a waiting thread. (Demo: blocking (Thread.sleep) frees the CPU)

Virtual threads: M:N on carrier threads

A virtual thread (Java 21, JEP 444) is not a kernel task. It is a Java object plus a stack that the JVM can copy to and from the heap.

The virtual-thread scheduler is a ForkJoinPool. Its parallelism defaults to availableProcessors(). It mounts a virtual thread on a carrier, which is an ordinary platform thread named ForkJoinPool-1-worker-N (shortened to carrier-N on the canvas). Only the carriers have TIDs.

When a virtual thread blocks (sleep, socket I/O, a ReentrantLock), it unmounts: its frames are saved to the heap. The carrier then mounts the next virtual thread without any kernel context switch. Eight virtual threads run on four kernel tasks. (Demo: 8 virtual threads on 4 carriers)

Four carrier threads, each a kernel task on cpu0, cpu2, cpu1 or cpu3, have VT#1 to VT#4 mounted; VT#5 and VT#6 sleep unmounted with their stacks in the heap, VT#7 and VT#8 wait in the ForkJoinPool queue
Virtual threads are mounted on a few carrier threads; a blocked one is moved to the heap and the carrier picks up the next, with no kernel context switch.

The catch is pinning. In JDK 21–23, a virtual thread that blocks inside a synchronized block cannot unmount. Its carrier blocks with it. With every carrier pinned, other virtual threads wait in the queue even though CPUs are idle. (Demo: pinning)

JDK 24 (JEP 491) removed pinning for synchronized. Native frames still pin. For how a ForkJoinPool shares work between its workers, see Java ForkJoinPool and Work Stealing.

How many threads?

Runtime.availableProcessors() counts logical CPUs. It also respects CPU affinity (taskset) and container limits (cgroup cpu.max, cpuset). It sizes the defaults for the common pool, the virtual-thread scheduler and ParallelGCThreads.

For CPU-bound work, more runnable threads than logical CPUs only adds context switches. For I/O-bound work, threads spend most of their time in S, so many more threads (or virtual threads) fit. See Concurrency vs Parallelism and Latency vs Throughput.

See also Stack vs Heap in Java (each thread's stack), How a Java Application Starts (where main and the JVM threads come from) and How a CPU Cache Works (why migration and sibling sharing cost cache hits).

What the page leaves out

The page leaves out these parts of the real system:

  • NUMA and multiple sockets;
  • real vruntime and EEVDF deadlines, priorities and nice;
  • interrupts and softirqs;
  • cgroup CPU quota throttling;
  • real-time scheduling classes;
  • CPU frequency scaling and turbo;
  • JVM safepoints that stop all Java threads;
  • ForkJoinPool work stealing between carriers, and its compensation threads;
  • the cost of the context switch itself.

Time slices here are 2 ticks. Real slices are a few milliseconds.