The problem: serve 10 000 connections without 10 000 threads
The classic web server gives each connection a process or a thread. It is simple to program: the code reads the request, waits for the disk or the backend, writes the response, all top to bottom. But each thread costs memory for its stack and time for context switches, and most of the time these threads are just waiting: for a slow client on a mobile network, for a backend, for the next request on a keep-alive connection. At ten thousand connections (the "C10K problem") this falls over.
nginx turns this around. Each worker is a single-threaded event loop built on epoll, and every piece of work is a short handler that never blocks. When a request has to wait for something, its handler arranges to be called back when that something is ready and returns to the loop. One worker then serves thousands of connections, and a handful of workers use every CPU core.
Master and workers
- The master process runs as root. It reads and checks
nginx.conf, opens the listening sockets (so port 80 can be bound), writes the pid file and forks the workers. It never touches a request. Onnginx -s reloadit reads the new config, starts new workers and tells the old ones to finish their current requests and exit, so a reload drops no connection. - The workers (
worker_processes auto: one per CPU core) run as an unprivileged user and do all the work. Each has its own epoll instance, its own connections and its own timers; they share nothing but the listening sockets. Withreuseporteach worker gets its own listening socket and the kernel spreads new connections among them; otherwise all workers watch the same socket andEPOLLEXCLUSIVE(or the olderaccept_mutex) keeps a new connection from waking them all.
The event loop
for (;;) { // ngx_worker_process_cycle
ngx_process_events_and_timers():
timeout = time until the nearest timer // timers live in a red-black tree
n = epoll_wait(ep, events, 512, timeout)
for each event: call its handler // c->read->handler or c->write->handler
expire every timer that is due // e.g. client_header_timeout
run posted events
}
Every connection is an ngx_connection_t taken from a pool of worker_connections slots (512 by default; the page uses 6). Each has a read event and a write event, and each event has a handler, a function pointer that nginx changes as the connection moves on: ngx_http_wait_request_handler while waiting for the first bytes, ngx_http_process_request_headers while reading the header, ngx_http_request_handler while the request runs, ngx_http_keepalive_handler between requests. The animation shows the current read handler under each slot. A connection to a backend takes a slot too, so a proxy uses two slots per request.
Reading the request
After accept4(), ngx_http_init_connection adds the socket to epoll and arms client_header_timeout, but allocates almost nothing: a connection that never sends a byte costs little. When the bytes arrive, nginx reads them into a small buffer (client_header_buffer_size, 1 KB; long headers move to large_client_header_buffers), creates the ngx_http_request_t with its own memory pool, and parses the request line and headers with hand-written state machines that work on the buffer in place. If the header is not complete yet, the handler simply returns and the next read event continues where it stopped.
The phase engine
Once the header is complete, ngx_http_core_run_phases walks the request through eleven phases. Modules register handlers in the phases they care about, and each handler returns "go on", "go to the next phase", "I finished the request" or "I will be called back" (NGX_AGAIN).
| Phase | What runs there |
|---|---|
| POST_READ | realip (take the client address from X-Forwarded-For) |
| SERVER_REWRITE | rewrite, return, set at server level |
| FIND_CONFIG | choose the location: exact =, then the longest prefix, then regexes in order (internal, no modules) |
| REWRITE | rewrite / return inside the location |
| POST_REWRITE | if the uri changed, go back to FIND_CONFIG (at most 10 times) |
| PREACCESS | limit_req, limit_conn |
| ACCESS | allow/deny, auth_basic, auth_request |
| POST_ACCESS | combine the access results (satisfy all | any) |
| PRECONTENT | try_files, mirror |
| CONTENT | exactly one content handler: static files, proxy_pass, fastcgi_pass, return, … |
| LOG | access_log, after the response is finished |
This order explains many nginx surprises: return in a location beats allow/deny because REWRITE runs before ACCESS, and try_files runs after access checks but before the content handler.
The output filters
A content handler does not write to the socket. It hands a header and a chain of buffers to two filter chains: header filters (headers_more, the final ngx_http_header_filter that renders the status line and headers) and body filters (gzip, sub_filter, SSI, chunked, and last the write filter). Each filter looks at the buffers, may change them, and passes them on. Buffers can point to memory or to a region of a file; the static module's buffer points to the file, so the write filter can use sendfile() and the file's bytes never enter user space. If the socket's send buffer is full, the write filter gets EAGAIN, keeps the rest, arms a write event and returns to the loop; it continues when epoll reports the socket writable.
Proxying: where the event loop pays off
With proxy_pass, the content handler starts an upstream: it opens a non-blocking connection to the backend, which returns EINPROGRESS, and returns. The request is parked. When epoll reports the backend socket writable, nginx sends the request; when it reports it readable, nginx parses the backend's header and streams the body back through the filters (buffering it in memory or a temp file if the client is slower than the backend, proxy_buffering). Run Demo: concurrency: while client A's request waits three ticks for the backend, the same single thread serves client B's static file from start to finish.
To see the same kind of proxied request from the operating system's side (memory pools, every system call, socket buffers and the copies between kernel and user space) open nginx Proxy Request: Memory, System Calls and Kernel.
Node.js runs JavaScript on the same kind of single-threaded loop; see The Node.js Event Loop.
Timers, keep-alive and slow clients
Every wait has a limit, kept in a red-black tree of timers ordered by expiry time; the loop's epoll_wait timeout is the time to the nearest one. client_header_timeout and client_body_timeout (60 s) protect against clients that connect and send nothing, or trickle bytes (Slowloris): nginx answers 408 and closes. keepalive_timeout (75 s) closes an idle keep-alive connection; send_timeout and proxy_read_timeout limit waits on the client and on the backend. Because an idle connection is just a slot and a timer, nginx can hold tens of thousands of them at little cost; a thread-per-connection server pays for a whole thread each time.
What must never happen: blocking the worker
All of this works only while every handler returns quickly. A handler that blocks stalls every connection of that worker: a slow disk read that is not in the page cache, a blocking DNS lookup, a heavy regex, CPU-bound Lua or Perl code. nginx has escapes for the usual cases: aio threads moves file reads to a thread pool, resolver is an asynchronous DNS client, and third-party modules are expected to be non-blocking. The same rule holds for every event-loop program, from Node.js to Redis.