The same request, seen from the operating system
The nginx page follows a request through nginx's own logic: phases, filters, handlers. This page looks one level lower, at what the operating system sees. It shows which memory the worker allocates and frees, which system calls its thread makes (the right column is what strace -p <worker pid> prints), what the kernel keeps in its socket buffers, and every time bytes are copied between kernel and user memory. The scenario is deliberately small: one request, proxy_pass to a backend on the same machine, and Connection: close on both sides.
location /api/ {
proxy_pass http://127.0.0.1:8080; # proxy_buffer_size 4k; proxy_buffers 8 4k (defaults)
}
User space and kernel space
nginx runs in user space: it can compute and use its own memory, but it cannot touch the network card, another process or a socket's buffers. Everything else goes through a system call, which switches the CPU into the kernel and back (one call in full detail: A Linux System Call, Step by Step). The kernel owns the file descriptor table of each process (fd 8 is just an index into it), the sockets with their receive and send queues, the listening sockets' accept queues, the epoll instance, and the page cache. The TCP handshake, retransmissions and ACKs all happen in the kernel. nginx is not even woken up for them.
Sockets and their TCP states
Each card in the kernel half is one socket. An fd is only a number that points to a socket; the socket itself holds the TCP state, the two addresses, a receive queue (bytes that arrived and nobody has read yet) and a send queue (bytes written but not yet acknowledged by the peer), plus a list of waiters to wake when something changes: epoll's callback for nginx, the blocked thread for the app. A listening socket has no data queues. It has a SYN queue of half-open connections and an accept queue of finished ones waiting for accept().
Follow the states. The handshakes (SYN_RECV, SYN_SENT → ESTABLISHED) happen entirely in the kernel, before any program runs. The side that closes first goes FIN_WAIT_1 → FIN_WAIT_2 → TIME_WAIT and keeps the socket for 60 s after its fd is gone. Here that is the app for the backend connection and nginx for the client connection. The side that receives the FIN first sits in CLOSE_WAIT until it calls close(), then LAST_ACK, then the socket is freed. A server that shows thousands of sockets in CLOSE_WAIT has a bug: it never closes them.
Two threads, three states
The timeline under the kernel shows the life of nginx's worker thread (top row) and of the backend app's thread (bottom row), one column per step. nginx's thread is running nginx code (parsing, building buffers, deciding), inside a system call (the kernel working on its behalf), or asleep in epoll_wait. Notice where the grey is: waiting for the client's request, waiting for the backend to connect, waiting for the backend to compute. A thread-per-connection server would spend those same moments with a whole thread blocked in recv(). nginx's thread spends them in epoll_wait, where one sleep covers every connection it has. With thousands of connections the grey gaps fill up with other requests' work.
The system calls of one proxied request
| System call | Why |
|---|---|
epoll_wait | sleep until some fd is ready or a timer is due; called once per loop turn |
accept4(…, SOCK_NONBLOCK) | take a finished connection off the accept queue as a new, non-blocking fd |
epoll_ctl(ADD) | watch the new fd, edge-triggered |
recv | copy the request from the socket's receive queue |
socket, ioctl(FIONBIO), epoll_ctl, connect | open the backend connection without waiting (EINPROGRESS) |
getsockopt(SO_ERROR) | did the non-blocking connect succeed? |
writev | send header and body buffers in one call (to the backend, then to the client) |
recv / readv | read the backend's response into one or several buffers; 0 means the backend closed |
write | the access log line |
close | release the backend connection, then the client's |
About twenty system calls for a request, and none of them blocks. The only call that waits is epoll_wait, and it waits for all connections at once.
Memory: preallocated connections and pools
- Connections are not allocated per client. At startup each worker allocates an array of
worker_connectionsngx_connection_tstructures, with their read and write events, and keeps a free list. Accepting a connection takes one off the list, and a backend connection takes one too. - Pools. Everything a connection needs comes from
c->pool(connection_pool_size, 512 B on 64-bit), and everything a request needs fromr->pool(request_pool_size, 4 KB). Allocating from a pool is a pointer bump. Nothing is freed one object at a time: destroying the pool frees it all at once. That is fast, and nginx cannot leak memory across requests. - Buffers only when needed. The 1 KB header buffer is allocated when the first bytes arrive, not at accept, so an idle connection costs a few hundred bytes. The response uses
proxy_buffer_size(4 KB) for the header and takes up toproxy_buffers(8 × 4 KB) for the body only if it needs them. Compare the peak memory of the two response sizes. - Parsing without copying. The parsed method, URI and headers are pointers into the header buffer. The only copies inside user space are the ones that build something new: the backend request and the client's response header.
Copies between kernel and user memory
Every recv copies bytes from a socket's receive queue into nginx's memory, and every writev copies them back into a send queue. A proxied response therefore crosses the boundary twice inside nginx: in with recv/readv from the backend, out with writev to the client. The request crosses twice too, and the backend itself copies once more on each side. For a 20 KB response that is cheap. For large downloads through a proxy, these copies (and TLS encryption, when enabled) are where the CPU goes.
For static files nginx avoids them with sendfile(), which moves file pages from the page cache to a socket inside the kernel. A proxied response cannot use it, because its bytes come from a socket, not a file. With proxy_buffering on, a response larger than the buffers is written to a temporary file (proxy_max_temp_file_size) and can then be sent with sendfile.
Connection: close
With Connection: close, every request pays for two TCP handshakes (client and backend), two accept/socket + close pairs and fresh pools. Keep-alive to the client (the default in HTTP/1.1) and keepalive in an upstream block to the backend skip most of the kernel work shown here. The nginx page has a keep-alive demo.