What happens between Enter and the first byte
You type http://example.com/index.html and press Enter. Before a single byte of the page can arrive, the browser has to find out where example.com is, set up a TCP connection to it, and send the request. Each of these steps needs at least one round trip (RTT): a packet goes out and the answer comes back. So a fresh request costs about three round trips before the first response byte arrives: DNS 1 RTT, the TCP handshake 1 RTT, the request and response 1 RTT. At 100 ms per round trip (a phone on the other side of an ocean) that is 300 ms of waiting in which no bandwidth helps: only fewer round trips do.
The canvas has three zones. On the left the client and on the right the server, each with its own stack from top to bottom: the application (the browser, nginx) and the system call it is in, the kernel's TCP state (socket state, sequence numbers, and on the server the SYN queue and accept queue), and the network card, under which a tcpdump-like log lists every packet that host sends (→) or receives (←). In the middle, the network: a sequence diagram in which time goes down and each packet is an arrow that slopes down by the one-way delay (RTT / 2), with the DNS resolver above it. Grey arrows are DNS (UDP), blue ones TCP control segments, green ones HTTP data, red dashed ones retransmissions; a lost packet stops half way with a ✕, one dropped by the receiver ends in a ✕. The time axis is to scale up to 1.5 RTT; longer waits (timeouts, idle time) are drawn as a short break with their length.
URL → name → IP address
The URL gives a scheme (http, so port 80 unless another is given), a host name and a path. TCP needs an IP address, so the browser first calls the resolver (getaddrinfo()). If the name is in the browser's or the operating system's cache, this costs nothing. Otherwise a small UDP query goes to the configured DNS resolver (usually the home router or the ISP), which answers from its cache, or first asks the root, .com and example.com's own name servers. The answer carries a TTL: how long it may be cached. UDP needs no handshake, which is why DNS uses it. What happens inside that one DNS round trip is shown on How a DNS Name Is Resolved. How each of these packets is wrapped in TCP, IP and Ethernet headers and carried across a switch and routers is shown on Internet Protocol Layers.
The TCP header
Every TCP segment carries the source and destination port (together with the two IP addresses they form the four-tuple that names the connection), a sequence number (the number of its first byte), an acknowledgement number (the next byte the sender expects to receive, so "I have everything before this"), a receive window (how much more it can take), and six flag bits. The page shows the ones that matter here:
- SYN: "synchronise", the first segment of each direction; it carries the sender's initial sequence number (ISN).
- ACK: the acknowledgement number is valid. Set on every segment after the first SYN.
- PSH: deliver to the application now; set on the last segment of a write.
- FIN: "I will send nothing more".
- RST: "this connection does not exist", abort.
The SYN also carries options that are only allowed there: the maximum segment size (MSS, 1 460 B on Ethernet: 1 500 minus 40 B of IP and TCP header), window scaling (to allow windows larger than 64 KB), SACK permitted and timestamps.
The three-way handshake
client server
SYN seq=1000 → "my bytes start after 1000"
← SYN-ACK seq=5000 ack=1001 "got it; mine start after 5000"
ACK seq=1001 ack=5001 → "got yours"
Each side picks a random ISN (the page uses 1000 and 5000 so the numbers stay readable). Random ISNs make it hard for an attacker who cannot see the traffic to inject segments, and keep segments of an old connection with the same four-tuple from being taken for new ones. SYN and FIN each consume one sequence number, although they carry no data, so that they can be acknowledged like data: ack=1001 means "I got your SYN, byte 1000".
Why three and not two? Both directions have to agree on a starting number, and each side must know the other received its ISN. The client's SYN is acknowledged by the SYN-ACK; the server's SYN is acknowledged only by the third segment. With two segments the server could never tell a real client from an old duplicate SYN still wandering the network, and it would start sending to someone who is not listening.
All of this is done by the kernel. The client's connect() just blocks (or, non-blocking, returns EINPROGRESS) until the SYN-ACK arrives; the server application is not involved at all.
SYN queue, accept queue and listen()'s backlog
A listening socket has two queues. When a SYN arrives, the kernel stores a small request sock in the SYN queue (the connection is half-open, SYN_RCVD) and answers with the SYN-ACK. When the final ACK arrives, it creates the full socket and moves it to the accept queue: the connection is ESTABLISHED although the application has not seen it yet. accept() only takes an already-finished connection off that queue and gives it a file descriptor. This is why the nginx, Tomcat and epoll pages start at "the kernel finished the handshake".
The backlog argument of listen(fd, backlog) limits the accept queue (capped by net.core.somaxconn, 4096 on current Linux; nginx asks for 511). If the application does not call accept() fast enough and the queue is full, Linux drops the handshake's final ACK (with the default tcp_abort_on_overflow = 0) and also new SYNs. The client meanwhile believes it is connected and sends its request, which is dropped too; the server retransmits its SYN-ACK after 1 s, and the client's answer finishes the handshake once there is room. Run Demo 6: a full backlog turns into seconds of delay rather than an error. (The SYN queue has its own limit; when it overflows during a SYN flood, Linux answers with SYN cookies instead of storing anything.)
If nothing listens on the port, the server's kernel answers the SYN with RST and connect() fails at once with ECONNREFUSED. A firewall that silently drops the SYN instead makes the client retry for more than a minute before ETIMEDOUT.
The HTTP/1.1 request and response
GET /index.html HTTP/1.1 HTTP/1.1 200 OK
Host: example.com Server: nginx
User-Agent: Mozilla/5.0 … Content-Type: text/html
Accept: text/html,*/* Content-Length: 1280
Accept-Encoding: gzip, br Connection: keep-alive
Connection: keep-alive
<!doctype html><html>…
HTTP/1.1 is plain text: a request line (method, path, version), header lines, and an empty line; the response has a status line, headers, an empty line and the body, whose end is given by Content-Length or by chunked encoding. Host is required, so one IP address can serve many sites. The request fits in one segment; a larger response is cut into MSS-sized segments, and the server may send up to its congestion window (10 segments at the start, on Linux) before it has to wait for ACKs. The server's acknowledgement of the request usually needs no packet of its own: the kernel delays a bare ACK for a few tens of milliseconds, nginx answers well within that time, and the ACK rides on the response. The time to first byte (TTFB) is the time from Enter until the first byte of the response arrives.
Loss and retransmission
TCP notices a lost segment in two ways. The retransmission timeout (RTO): if no ACK comes back in time, the segment is sent again and the timeout doubles. Before the first RTT is measured, Linux waits 1 s, so a lost SYN or SYN-ACK costs a whole second (then 2, 4, 8 s...): this is why a bad network makes page loads jump by whole seconds (Demo 4). Once the handshake has given an RTT sample, the RTO is about srtt + max(4 × rttvar, 200 ms). The second way is fast retransmit: when a segment is missing, the receiver ACKs every later segment with the same acknowledgement number (duplicate ACKs, with SACK blocks saying what did arrive); after the third duplicate ACK the sender resends the missing segment at once (Demo 7). With too few segments in flight to produce three duplicate ACKs, only the timeout is left (modern Linux adds tail loss probes and RACK to shorten this). What a loss does to the sending rate over many round trips is shown on How TCP Congestion Control Finds the Bandwidth.
Keep-alive and parallel connections
HTTP/1.1 keeps the connection open after the response (Connection: keep-alive is the default). The next request, for the page's CSS, script or images, skips DNS and the handshake: one round trip instead of two (Demo 3). A connection is also warmed up: its congestion window has grown and its RTT is known. But HTTP/1.1 can have only one request in progress per connection (pipelining exists but is not used), so browsers open up to six connections per host to fetch resources in parallel, each paying its own handshake. The server closes idle connections after a while (nginx: keepalive_timeout 65 in the stock config) to free their memory.
Closing: four segments, half-close and TIME_WAIT
Each direction is closed on its own. The side that closes first sends FIN (FIN_WAIT_1); the other side ACKs it and enters CLOSE_WAIT: the connection is now half-closed, the second side may still send data. When its application calls close() too, it sends its own FIN (LAST_ACK), and the first side ACKs it. So a close takes four segments (three when the ACK and the second FIN travel together).
The side that closed first then waits in TIME_WAIT for 2 × MSL (maximum segment lifetime; 60 s on Linux) before the socket disappears. Two reasons: if the last ACK is lost, the other side resends its FIN and it must still be answered; and old segments of this connection must die out before the same four-tuple can be used again. A busy server that closes many connections first collects many TIME_WAIT sockets; they cost little memory, but a client that opens many short connections to one server can run out of source ports. Run Demo 5 to see the server close first on its keep-alive timeout, or choose close: client first and press Close.
CLOSE_WAIT is the opposite case: the peer has closed, and the kernel is waiting for this side's application to call close(). No timer ends it, so a program that ignores end of file (a read() returning 0) and never closes the fd keeps the socket in CLOSE_WAIT for as long as the process lives, holding an fd and its buffers. Many CLOSE_WAIT sockets in ss -tan state close-wait almost always mean such a bug, and they end in "Too many open files". The other side is not stuck: its socket, already closed by its application, waits in FIN_WAIT_2 for at most tcp_fin_timeout (60 s) and is then dropped, so a FIN that arrives later only gets an RST. Run Demo 8 to see it.
Anyone on the path can read it
Look at the packet detail row while the request is on the wire: the request line, the Host, cookies if there were any, and the whole response are plain text. Every router, Wi-Fi access point or proxy on the path can read and even change them. HTTPS puts TLS between TCP and HTTP: after the TCP handshake, a TLS handshake agrees on keys and checks the server's certificate, and everything after it is encrypted. What that costs in round trips, and how TLS 1.3 brings it down to one, is shown on How an HTTPS Connection Is Established.
What HTTP/2 and HTTP/3 change
- HTTP/2 still runs on one TCP (and TLS) connection per host, but multiplexes many requests on it as interleaved streams of binary frames, with compressed headers. One connection replaces the six, so there is one handshake and one congestion window. Its weakness is TCP itself: one lost segment holds up every stream behind it (head-of-line blocking).
- HTTP/3 runs over QUIC, a transport on top of UDP that includes TLS 1.3. The transport and crypto handshakes are one: a new connection costs 1 RTT before the request instead of 2 (TCP + TLS 1.3), and a resumed one can send the request in the very first packet (0-RTT). Streams are independent, so a lost packet only delays its own stream, and a connection survives a change of IP address (Wi-Fi to mobile).
Neither is simulated here: this page is about establishing the connection, and the steps above (name lookup, a handshake counted in round trips, a request, loss recovery, closing) are still what every version does. See HTTP/1.1 vs HTTP/2 vs HTTP/3 for the same page loaded over all three side by side, and TCP vs UDP for what TCP's handshake, ACKs and retransmissions buy compared with plain datagrams.