The idea: two ways to use the same IP network
IP moves single packets from one host to another, with no promise: a packet can be lost,
arrive twice or arrive after a later one. The transport layer sits on top and decides what the
application gets. UDP adds almost nothing: port numbers and a checksum. Each
sendto() becomes one datagram, and the receiver gets whatever arrives.
TCP builds a connection: a reliable, ordered byte stream
on top of the same unreliable packets, paid for with a handshake, acknowledgements, timers and state
in both kernels.
The page sends the same messages over both, through a network that does the same thing to both (drops the same message, or delays it). The top band is TCP, the bottom band UDP. Each band shows the client application, the client kernel, the wire, the server kernel and the server application, and the table under them says when the server application actually got each message.
The clock ticks in 20 ms steps, and one tick is the one-way delay, so a round trip (RTT) is 40 ms. Messages are 100 bytes, one message per packet. To keep numbers small, TCP sequence numbers count segments here (SYN = 0, M1 = 1, …); real TCP counts bytes and starts from a random number.
Setting up and tearing down a connection
Before TCP can send a byte, the client kernel sends SYN, the server answers
SYN-ACK, and the client's ACK completes the three-way
handshake. That costs one RTT; connect() blocks for it, and the server's
accept() only returns once the handshake is done. The third segment may already carry data,
as it does here. UDP has nothing to set up: the first datagram leaves at t = 0 and is read at t = 1,
while TCP's first message is read at t = 3 (Demo: clean network).
For one question and one answer, like a DNS lookup, this is the whole difference: UDP answers in 1 RTT, TCP in 2 RTT, and then TCP still exchanges four more packets to close (Demo: one question, one answer).
Closing takes a FIN from each side. The side that closes first goes
FIN_WAIT_1 → FIN_WAIT_2 → TIME_WAIT; the other side sits in
CLOSE_WAIT until its application calls close(), then LAST_ACK →
CLOSED. TIME_WAIT lasts 2·MSL (60 s on Linux) so that a late duplicate
from this connection cannot be taken for data of the next connection with the same ports.
UDP has none of these states: the socket is just a port with a queue.
| state | who | means |
|---|---|---|
| LISTEN | server | waiting for SYNs on the port |
| SYN_SENT / SYN_RCVD | client / server | handshake half done |
| ESTABLISHED | both | data flows both ways |
| FIN_WAIT_1 / FIN_WAIT_2 | closes first | FIN sent / FIN acknowledged, waiting for the peer's FIN |
| CLOSE_WAIT | closes second | peer's FIN received; our application has not closed yet |
| LAST_ACK | closes second | our FIN sent, waiting for its ACK |
| TIME_WAIT | closes first | all done; waits 2·MSL before the port pair may be reused |
Reliability: ACKs, duplicate ACKs, retransmission
Every TCP segment has a sequence number, and the receiver answers with a
cumulative ACK: "ACK 3" means "I have everything before 3, send 3 next". The sender
keeps each segment in its send buffer until it is acknowledged
(snd_una is the oldest unacknowledged one, snd_nxt the next new one).
There are two ways to notice a loss. If later segments still arrive, each one makes the receiver
repeat the same ACK: a duplicate ACK. After the third one, the sender resends the
missing segment at once: fast retransmit (Demo: M3 lost). If nothing comes back
at all, for example because the lost segment was the last one, only the retransmission timer
(RTO) helps; here it is 6 ticks (120 ms) and doubles after every timeout. On Linux it is
srtt + 4·rttvar, at least 200 ms, and 1 s before the first RTT is measured
(Demo: last message lost).
UDP has no sequence numbers and no ACKs. A lost datagram is simply gone, and the sender never learns about it. If the application needs an answer, it must run its own timer and ask again, as a DNS resolver does (Demo: the question is lost).
Ordering and head-of-line blocking
A TCP receiver hands bytes to the application only in order. Segments that arrive after a hole wait in the out-of-order queue; when the retransmission fills the hole, they all move up together. In Demo: M3 lost, M4, M5 and M6 are in the server's kernel on time but the application gets them only with M3, several ticks later: head-of-line blocking. UDP delivers M4, M5 and M6 on time and never delivers M3.
Which is better depends on the data. For a file, a web page or a database query, a late byte is fine and a missing byte is not. For a voice frame or a game position, data that is 200 ms late is useless, and waiting for it delays the fresh data behind it. This is why real-time media and games use UDP, and why HTTP/3 runs QUIC over UDP: QUIC keeps reliability per stream, so one lost packet does not stall the other streams (see HTTP/1.1 vs HTTP/2 vs HTTP/3).
Reordering is the milder case: in Demo: reordering, TCP holds M5 until M4 arrives (one duplicate ACK, no retransmission); UDP hands M5 to the application before M4. A UDP application that cares must number its own messages.
Byte stream vs datagrams
TCP moves bytes, not messages. Six send() calls of 100 bytes do not mean six
read() calls of 100 bytes: after the loss in Demo: M3 lost, one
read() returns 400 bytes (M3 to M6), and the slow reader's read(fd, buf, 150)
returns one and a half messages. An application protocol over TCP must mark where each message ends:
a length prefix, a delimiter such as HTTP's blank line, or fixed-size records.
UDP keeps boundaries: one sendto() is one datagram and one recvfrom()
returns exactly one datagram, whole or not at all. The price is size: a datagram larger than the path
MTU (about 1 472 bytes of payload on Ethernet) is fragmented by IP, and losing any fragment loses the
whole datagram.
Flow control vs dropping
Each TCP ACK carries the receiver's free buffer space, the advertised window
(rwnd). The sender never has more unacknowledged bytes in flight than that. When the
application reads slowly, the window shrinks, possibly to zero, and the sender simply waits; the read
that frees space triggers a window update. Nothing is lost
(Demo: slow reader).
UDP has no window. Datagrams pile up in the socket's receive buffer (SO_RCVBUF), and when
it is full the receiving kernel drops the new ones and counts them
(RcvbufErrors in /proc/net/snmp, netstat -su). The network did
nothing wrong, and again the sender does not know.
Flow control protects the receiver. Protecting the network is a separate mechanism, congestion control, which this page leaves out; see How TCP Congestion Control Works.
Headers and state
| TCP | UDP | |
|---|---|---|
| header | 20 bytes (up to 60 with options) | 8 bytes |
| packets for the 6 messages here | 18 (handshake, data, ACKs, close) | 6 |
| first message read by the server | t = 3 (60 ms) | t = 1 (20 ms) |
| state per peer | a socket per connection: sequence numbers, buffers, timers, window | none; one socket can talk to any number of peers |
| delivery | every byte once, in order, or an error | each datagram at most once… or twice, or never, in any order |
| message boundaries | no (byte stream) | yes |
| flow / congestion control | yes / yes | no / no (the application's job) |
Because UDP keeps no per-peer state, a single UDP socket on a DNS server answers millions of clients, while a TCP server needs one accepted socket per client (see epoll for how one thread watches many of them).
When to use which
| use | transport | why |
|---|---|---|
| web pages, APIs, file transfer, email, databases, SSH | TCP | every byte must arrive, in order |
| DNS queries | UDP (TCP for large answers and zone transfers) | one small question, one answer; retrying is cheap |
| voice, video calls, live streaming, games | UDP (RTP, WebRTC) | late data is useless; no head-of-line blocking |
| HTTP/3 | QUIC over UDP | reliability per stream, 1-RTT (or 0-RTT) setup with TLS, runs in user space |
| DHCP, NTP, SNMP, syslog | UDP | tiny, stateless, often broadcast |
See also How an HTTP Connection Is Established (the handshake and close in detail), How TCP Congestion Control Works, Internet Protocol Layers (where TCP and UDP sit, IP and fragmentation), How a DNS Name Is Resolved and HTTP/1.1 vs HTTP/2 vs HTTP/3.
What the page leaves out
Congestion control (cwnd, slow start) is left out: the sender's window here is limited
only by the receiver. Real TCP counts bytes, starts from random sequence numbers, delays ACKs (one ACK per
two segments, or after up to 40 ms), uses selective ACKs (SACK) so several holes can be repaired in one
RTT, and may hold small writes back (Nagle's algorithm). The server's SYN-ACK retransmission, SYN floods and
SYN cookies, RST, keep-alive, TCP Fast Open, checksums and IP fragmentation are not drawn. UDP relatives
such as UDP-Lite, and SCTP (messages with reliability), are not covered.