The idea: a socket is a file descriptor for a kernel object

A program never touches the network itself. It asks the kernel for a socket and gets back a small number, a file descriptor (fd), just like opening a file. Behind that fd the kernel keeps an object with everything TCP needs: the connection's state, the two addresses, a send buffer and a receive buffer. The program only copies bytes in and out of those buffers. The kernel does all the TCP work: the handshake, sending, ACKs and closing.

curl's fd 3 points to kernel socket A with a send and a receive buffer; nginx's fd 3 points to a listening socket with an accept queue and its fd 4 to socket B; the request goes from A's send buffer to B's receive buffer, the response from B's send buffer to A's receive buffer
The program holds only a number; the kernel socket behind it has the state, the addresses and one buffer per direction.

On the canvas, curl on the left talks to nginx on the right. Press the buttons from left to right, or run Demo: the whole life of a connection. The buffers have 16 bytes so you can see them fill. Real ones are much bigger (net.ipv4.tcp_rmem starts at 128 KiB).

The server: listen() and accept()

A server uses two kinds of socket:

  • The listening socket. socket() creates it, bind() gives it port 80, and listen(fd, backlog) puts it in state LISTEN. It never carries data. It only collects connections that have finished their handshake in its accept queue, which holds at most backlog of them.
  • One connected socket per client. The kernel creates it during the handshake, without asking the program. accept() takes it off the queue and returns a new fd for it (fd 4 here). The listening fd 3 stays open for the next client.

If nginx calls accept() before any client has connected, it simply sleeps until one does.

The client: connect() and the handshake

connect() picks a free local port, sends SYN and puts the program to sleep. The two kernels exchange SYN, SYN-ACK and ACK. When the SYN-ACK arrives, connect() returns 0. If nothing listens on the port, the server's kernel answers with RST and connect() fails with ECONNREFUSED (Demo: connection refused). How an HTTP Connection Is Established shows the handshake and the segment headers in detail.

send() and recv() are copies

send() copies bytes from the program's buffer into the socket's send buffer and returns. At that moment nothing has gone over the network yet. The kernel then sends the bytes. It keeps them in the send buffer until the other side ACKs them, in case it has to send them again.

recv() copies bytes from the receive buffer into the program's buffer. If the buffer is empty, the program sleeps until data arrives.

Every ACK also says how much room is left in the receive buffer: the receive window. If the server stops reading, its receive buffer fills and the window drops to 0. The client's kernel must stop sending, its send buffer fills too, and the next send() blocks. This is flow control: a slow reader slows the writer down instead of losing data (Demo: full buffers make send() block).

nginx is idle, socket B's receive buffer is full, its ACK says window 0, so socket A sends no data, A's send buffer fills and curl's send() blocks
A full receive buffer advertises window 0, the sender's buffer fills behind it, and send() blocks instead of data being lost.
CallReturns at once whenBlocks when
accept()the accept queue has a connectionthe queue is empty
connect()(never at once: it waits for the handshake)until SYN-ACK or RST arrives
send()all bytes fit in the send bufferthe send buffer is full
recv()the receive buffer has data, or the peer sent FIN (returns 0)the buffer is empty

Data goes both ways over the same connection. Each socket has its own send buffer and its own receive buffer, so the request (blue, curl → nginx) and the response (orange, nginx → curl) never mix. Server: send copies HTTP/1.1 200 OK into socket B's send buffer, the server's kernel sends it, and it lands in socket A's receive buffer. Client: recv then copies it up into curl. If curl calls recv() before the response arrives, it sleeps until it does (Demo: curl waits in recv() for the response), which is what every HTTP client does after sending its request.

A non-blocking socket returns EAGAIN instead of sleeping. Servers like nginx use that together with epoll to handle thousands of sockets in one thread.

Closing: FIN, CLOSE_WAIT, TIME_WAIT

If a program closes a socket while unread data is still in its receive buffer, Linux sends RST instead of FIN, so the page asks curl to read the response first.

close() frees the fd at once, but the kernel keeps the socket until the connection has ended properly:

  1. The side that closes first (curl) sends FIN and goes to FIN_WAIT_1, then FIN_WAIT_2 after the ACK.
  2. The other side goes to CLOSE_WAIT, and its recv() returns 0, meaning end of file. It stays in CLOSE_WAIT until the program calls close(). Many sockets stuck in CLOSE_WAIT usually mean a program forgot to close them.
  3. nginx closes, sending its FIN (LAST_ACK). The last ACK frees socket B.
  4. The side that closed first waits in TIME_WAIT for 60 s on Linux, so a late segment can't be mistaken for a new connection.

Seeing it on a real machine

  • ss -tanp: every TCP socket with its state, addresses and owning process. For a LISTEN socket, Recv-Q is the number of connections waiting in the accept queue and Send-Q is the backlog. For a connected socket, they are the bytes waiting in the receive and send buffers.
  • ls -l /proc/<pid>/fd: a process's file descriptors; sockets show as socket:[inode].
  • strace -e trace=network curl http://example.com: the socket system calls one by one.

What the page leaves out

Packet loss and retransmission (see TCP vs UDP), congestion control (see TCP Congestion Control), the SYN queue and SYN cookies, Nagle's algorithm and delayed ACKs, buffers that grow on their own (autotuning), several clients at once, non-blocking sockets, SO_REUSEADDR, shutdown() for half-closing, and IPv6.