The idea: let the device move the data

A packet has to get from the network card into RAM, where the kernel and the program can use it. There are two ways to do that.

  • Programmed I/O (PIO). The card keeps the packet in its own small buffer and raises an interrupt. The CPU then reads it out through an I/O port, one 16-bit word at a time (rep insw), and stores each word in RAM. Every byte passes through a CPU register. While it copies, the CPU does nothing else.
  • Direct memory access (DMA). The card is a bus master: it can send its own read and write requests over PCIe to RAM. The driver only tells it where to put packets. The card writes the data itself while the CPU runs other programs, and it raises one interrupt when there is something to look at.
Left, programmed I/O: after one interrupt the CPU reads the packet from the network card with rep insw and stores it into the skb buffer in RAM, busy copying and doing no other work. Right, DMA: the network card writes the packet into the skb buffer in RAM with its own PCIe writes and sends one interrupt, while the CPU runs P2
With programmed I/O every byte goes through the CPU; with DMA the card writes RAM itself and the CPU stays free for other programs.

On the canvas, P1 is a UDP server sleeping in recvfrom(), and P2 is a program that just wants the CPU. The P2 ticks counter shows how much CPU time the I/O leaves for real work. (Demo 1 and Demo 3.)

PIO cardDMA card
who moves the bytesthe CPU, word by wordthe card, over PCIe
CPU cost per packetgrows with the packet sizea small fixed cost (descriptor, skb, refill)
interruptsat least one per batch the CPU copiesone per batch, and NAPI polls the rest
buffer spacea few packets on the carda ring of buffers in RAM (256 to 4096 in real NICs)
new problemsnone: the CPU sees everythingaddresses, cache coherency, protection

Descriptor rings: head, tail and the DD bit

The driver and the card share a ring of descriptors in RAM. An RX descriptor holds the bus address of an empty buffer, and space for the card to report the length and a status. The two sides take turns:

  • The card owns the descriptors from the head (RDH, which only the card moves) up to just before the tail (RDT, which only the driver moves). On the canvas these are purple.
  • For each packet the card fetches the descriptor at the head, writes the data into that buffer, then writes the descriptor back with DD (descriptor done) set and moves the head on (yellow).
  • The driver walks the ring from where it stopped, takes every descriptor with DD, gives it a fresh buffer, and then moves the tail with one register write. That write is the doorbell.
  • If the head catches up with the tail, the card has no buffer left and must drop packets (rx_missed_errors, Demo 6). One slot always stays empty, so that "head == tail" can only mean "nothing for the card".
An RX ring of 8 descriptors: slots 0, 1 and 2 have DD set and wait for the driver; slots 3 to 6 hold empty buffers owned by the card, starting at the head RDH = 3; slot 7 at the tail RDT = 7 is the one slot kept free; after slot 7 the ring wraps to slot 0
The card fills buffers from the head and sets DD; the driver takes the DD slots, refills them and moves the tail to hand them back.

This pattern is everywhere: NVMe submission and completion queues, virtio virtqueues and io_uring all work the same way. A queue in shared memory has a producer index and a consumer index, and a doorbell says "there is new work".

Every step of one received packet

(Demo 2.) Purple boxes on the timeline are PCIe transactions made by the card itself.

#whowhat
1wire → NICThe packet lands in the card's small on-chip FIFO.
2NIC (DMA read)It fetches the descriptor at RDH and gets the buffer address.
3NIC (DMA write)It writes the packet into the buffer, one 64-byte line per tick here, through the IOMMU.
4NIC (DMA write)It writes back the descriptor: length, DD. RDH moves on and the RX cause is set in ICR.
5NIC → CPUMSI-X: a DMA write to address 0xfee00000 that the root complex turns into an interrupt.
6CPU, hard IRQThe handler masks the NIC interrupt and calls napi_schedule(). Nothing else.
7CPU, NAPI pollFor each descriptor with DD: dma_unmap_single(), build an skb around the page (no copy), queue it on the socket, wake P1.
8CPU, NAPI pollRefill: a new or recycled page, dma_map_single(), its address into the descriptor. Then one MMIO write to RDT.
9CPU, NAPI pollFewer packets than the budget: napi_complete_done() and the interrupt is unmasked.
10P1recvfrom() copies the 256 bytes to ubuf (copy_to_user) and returns.

Transmit is the mirror image (Demo 4). sendto() copies the data into an skb, the driver maps it (DMA_TO_DEVICE), writes a TX descriptor and rings TDT. The card reads the descriptor and the data, sends the packet, and writes back DD. In the next poll the driver unmaps the buffer and frees it.

Interrupts and NAPI

At a million packets per second, one interrupt per packet would use up the whole CPU. This is called receive livelock: the CPU spends all its time entering and leaving interrupt handlers, and the packets it accepts are never processed. Linux avoids it with NAPI. The first packet raises an interrupt. The handler masks further interrupts from the card and schedules a poll, which handles up to a budget of packets per round (64 by default; 4 on the canvas). While packets keep coming, the poll keeps running and there are no more interrupts. When a round finds less work than the budget, the driver unmasks the interrupt again. The card also has its own limit, interrupt moderation (ethtool -C eth0 rx-usecs 50): it waits at least that long between two interrupts (here ITR = 18 ticks). A PIO card has neither, and the copying itself is what overloads the CPU (Demo 3).

Devices don't use virtual addresses: the DMA API

A buffer that is contiguous for the program is a set of scattered physical pages, and a device knows nothing about page tables. So the driver must turn a buffer into a bus address, and in Linux this always goes through the DMA API:

  • dma_alloc_coherent(): memory shared with the device for a long time, set up so the CPU and the device always see the same data. The descriptor rings live here.
  • dma_map_single() / dma_map_sg(): streaming mappings for one transfer, with a direction (DMA_FROM_DEVICE, DMA_TO_DEVICE). This returns the address to put in the descriptor, and it does any cache work or bouncing that is needed. dma_unmap_* and dma_sync_single_for_cpu/for_device undo it or hand the buffer over again.

The rule is ownership. Between map and unmap, the buffer belongs to the device and the CPU must not touch it. After unmap (or sync_for_cpu) it belongs to the CPU again.

The IOMMU: translation and protection

The IOMMU (Intel VT-d, AMD-Vi, Arm SMMU) sits in the PCIe root complex. It does for devices what the MMU does for programs: the device uses IO virtual addresses (IOVAs), and the IOMMU's page table turns them into physical addresses. It also checks the permission (read-only for a TX buffer, write for an RX buffer). An access that is not mapped is blocked and logged as a DMAR fault.

Without an IOMMU, a device can read or write any byte of RAM. A driver bug (a descriptor that still holds the address of a freed page) or a malicious device then overwrites another program's memory without any error (Demo 7). With the IOMMU, the same bug is a dropped packet and a line in dmesg (Demo 8). The IOMMU is also what lets a VM own a real device safely (VFIO), and what stops DMA attacks through Thunderbolt or FireWire ports.

Left, IOMMU off: the card writes to a stale address, a freed page that now belongs to another program, which is silently overwritten with no error. Right, IOMMU on: the same write is looked up in the IOMMU page table, is not mapped, and is blocked with a DMAR fault; the packet is dropped and dmesg gets one line
Without an IOMMU a stale DMA address corrupts another program's memory; with one, the write is blocked and logged.

MMIO writes from the CPU to the card's registers (RDT, TDT, IMS) go the other way and do not pass through the IOMMU.

Cache coherency

The CPU reads and writes through its cache; the card reads and writes RAM. They can disagree in two ways:

  • Stale line: the CPU has an old copy of a line in its cache (S) and the card writes new data to RAM. The CPU keeps reading the old copy.
  • Dirty line: the CPU wrote new data that is still only in its cache (M). The card reads RAM and gets the old bytes.

On x86 the hardware solves this: every DMA access is snooped, so a device write invalidates cached copies and a device read gets dirty data from the cache. Intel DDIO even writes incoming packets straight into the last-level cache. On many ARM and other embedded SoCs, DMA is not coherent, and the DMA API does the work in software: dma_map(DMA_TO_DEVICE) cleans (writes back) the cache, and dma_map/unmap(DMA_FROM_DEVICE) invalidates it.

A driver that skips this, for example by recycling pages with DMA_ATTR_SKIP_CPU_SYNC and forgetting dma_sync_single_for_cpu(), works on x86 and breaks on ARM: P1 receives the previous packet that was in the same page (Demo 5). Switch the CPU to ARM with the driver correct to see the invalidates, or press P1: sendto with the bug to send old bytes.

DMA masks and bounce buffers

Some devices can only produce 32-bit addresses (dma_set_mask(dev, DMA_BIT_MASK(32))), so they cannot reach RAM above 4 GiB, where the buffer pages here live. Without an IOMMU, Linux uses swiotlb: it gives the device a bounce slot below 4 GiB, and on unmap the CPU copies the data to the real buffer (on transmit, before the device reads it). That copy is exactly the work DMA was supposed to save. With an IOMMU, the IOVA is simply chosen below 4 GiB and mapped to the high page, so there is no copy. Try DMA mask = 32-bit with the IOMMU off, then on. (Encrypted VMs, SEV and TDX, bounce all DMA through swiotlb for a different reason: the device may not read private guest memory.)

Zero copy

DMA removes the copy between the device and RAM. One copy is left: recvfrom() and sendto() copy between the kernel's skb and the program's buffer. sendfile() and splice() skip it for files, MSG_ZEROCOPY lets the card read user pages directly, and AF_XDP or DPDK let a program own the RX ring's buffers itself.

How to see it on a real system

ethtool -g eth0                      # RX/TX ring sizes
ethtool -S eth0 | grep -i -e miss -e drop -e err    # rx_missed_errors etc.
ethtool -c eth0                      # interrupt moderation (rx-usecs)
grep eth0 /proc/interrupts           # one MSI-X vector per queue, counts per CPU
dmesg | grep -i -e DMAR -e iommu -e swiotlb
ls /sys/kernel/iommu_groups/*/devices

What the page leaves out

  • One RX queue, one TX queue and one CPU. Real NICs have many queues, spread over CPUs by RSS, each with its own MSI-X vector.
  • Sizes are scaled down: 256-byte packets, a ring of 8, a budget of 4. Real PCIe writes carry 128 to 512 bytes, and real RX buffers are 2 KiB or a page.
  • Checksum offload, TSO/GRO, header split, PCIe ordering and flow-control rules, and the IOTLB cache and its invalidation are not drawn.
  • The CPU cache never evicts anything here, as if it were large. Real drivers usually use page_pool, which keeps recycled pages mapped.
  • The old ISA 8237 DMA controller, a third-party "DMA engine" that copied for a device, and peer-to-peer DMA between devices.

See also Linux Virtual Memory (the MMU side of address translation), CPU Cache (coherency and MESI), Linux Context Switch (what happens when P1 sleeps and wakes), nginx Proxy Request: Memory and System Calls (the copies between user and kernel space), Linux epoll and TCP vs UDP.