The idea: every memory is the same array

A memory array stores 2N words of M bits. Its depth is the number of words, its width the bits per word. The page uses the book's 4-word × 3-bit array: a 2-bit address A1 A0 and three data bits, Data2 to Data0.

The array is a grid of bit cells, one row per word and one column per bit. An N:2N decoder (the one on the Combinational Logic page) turns the address into exactly one HIGH wordline. That wordline connects its row of cells to the vertical bitlines. Reading senses the bitlines. Writing drives them. The kind of memory is decided only by what sits in each bit cell.

A 2:4 decoder turns address 01 into one HIGH wordline WL1; the three cells of word 01 (1, 1, 0) drive bitlines 2, 1, 0 into sense amplifiers, giving Data2..0 = 110; the other rows stay cut off
The decoder raises exactly one wordline, that row drives the bitlines, and the sense amplifiers read the word: address 01 gives 110.

The six tabs show four kinds of cells (DRAM, SRAM, ROM, register file) and two logic arrays that reuse the same grid to compute functions (PLA, FPGA lookup table).

Decoder, wordlines, bitlines

In Demo 1, address 01 makes the decoder raise WL1 only. The three cells of word 01 are connected to their bitlines; the other 9 cells are cut off. At the bottom of each column sits a sense amplifier, which turns a small bitline voltage into a clean 0 or 1, and a write driver, which forces the bitline to the new value.

Real arrays are much wider than one word per row: a row (a DRAM "page") holds thousands of bits, and a column mux picks the word. The page keeps one word per row, as the book does.

DRAM: one transistor and one capacitor

A DRAM cell is an access transistor and a capacitor. The wordline drives the transistor's gate. A charged capacitor is a 1, an empty one a 0. The cell is tiny, so its capacitance is much smaller than the bitline's (1 : 9 on this page). That causes everything else:

  • Precharge: before a read the bitlines are set to half of VDD, 0.50 V.
  • Charge sharing: when the wordline opens, the cell and the bitline share their charge. A full cell lifts the bitline only to 0.55 V, an empty one pulls it to 0.45 V.
  • Sensing: the sense amplifier compares the bitline with 0.50 V and drives it all the way to 1.00 V or 0 V.
  • Write-back: the read left the cell at about half charge. The read is destructive, so the amplified value is written back while the wordline is still open.

The charge also leaks away through the transistor. On the page a charged cell loses 10 % every 16 ms. A 1 written at t = 0 is at 60 % after 64 ms, which still reads as 1, but at 40 % after 96 ms, which reads as 0. Demo 2 waits 96 ms and then reads word 01: it returns 000, and the write-back makes the loss permanent. All six stored 1s are now lost.

The cure is refresh: read every row and write it back before any 1 falls below the threshold. In Demo 3 the page refreshes every 48 ms. The weakest 1 never drops below 70 %, and word 01 still reads 110 at t = 144 ms. The chart shows the saw-tooth. Real DRAM must refresh every row at least every 64 ms. The memory controller does it in short bursts between normal accesses, so a little bandwidth is lost.

Chart of a stored 1 losing 10 percent every 16 ms: without refresh it falls below the 50 percent threshold and reads 0 at 96 ms; refreshed every 48 ms it saw-tooths between 100 and 70 percent and still reads 1 at 144 ms
A DRAM 1 leaks toward the read threshold; refreshing it every 48 ms restores the charge before it is lost.
DemoWhat happens to word 01 (110)Lost bits
1: read at t = 0bitlines 0.55, 0.55, 0.45 V → 110, written back0
2: read at t = 96 ms, no refresh1s at 40 %: bitlines 0.49 V → 0006
3: refresh at 48 and 96 ms, read at 144 ms1s at 70 %: bitlines 0.52 V → 1100

SRAM: cross-coupled inverters

An SRAM cell is two cross-coupled inverters, the same loop as the bistable element on the Sequential Logic page, plus two access transistors. Q connects to the bitline BL and the opposite value Q̅ to bitline-bar BL̅. The inverters always drive the cell from VDD, so nothing leaks away and there is no refresh.

To read, both bitlines are precharged HIGH. When the wordline opens, the side that stores 0 pulls its bitline down a little, and a differential sense amplifier sees which of the pair is lower. The read does not disturb the cell. To write, strong drivers force BL = D and BL̅ = not D and overpower the inverters. Demo 4 writes 011 to word 10, reads it, waits 96 ms and reads 011 again.

Side by side: a DRAM cell is one access transistor and a capacitor on one bitline, whose charge leaks; an SRAM cell is two cross-coupled inverters with two access transistors on BL and BL-bar, which holds its value without refresh
DRAM stores a bit as charge in one tiny capacitor that leaks; SRAM stores it in a loop of two inverters that holds it as long as there is power.

Area, speed and cost per bit

The cell decides the cost. A big array needs long wordlines and bitlines, so it is also slower. That is why computers use a memory hierarchy: a few fast, costly bits near the processor and many cheap, slow bits behind them.

Bit cellTransistors per bitCostSpeedUsed for
flip-flop≈ 20very highfastestpipeline registers, state
register file (2 read + 1 write port)≈ 10highfastthe 32 MIPS registers
SRAM6mediummediumcaches
DRAM1 + a capacitorlowslowmain memory
ROM1 or nonelowestmediumfixed tables, boot code

Volatile and nonvolatile

DRAM and SRAM are volatile: they forget everything when the power goes. In Demo 5 the SRAM wakes up with arbitrary contents. Each cell falls to whichever inverter happens to be stronger. The page draws the values from a fixed pseudo-random sequence, so the demo always shows the same garbage. A ROM is nonvolatile. It has nothing to forget (Demo 7).

ROM and dot notation

A ROM stores a bit as the presence or absence of a transistor. In dot notation, a dot where a wordline crosses a bitline means a 1. The top row of the book's ROM has a single dot, under Data1, so address 11 holds 010.

On the page every dot is a transistor. The bitlines are weakly pulled HIGH. The raised wordline turns on the transistors of its row, which pull their bitlines LOW, and an inverter under each column turns LOW into 1. (The book's bit cell does it the other way round: a transistor stores a 0 and there is no inverter. Either way the dots show the data.) A mask ROM's transistors are placed when the chip is made, so a write does nothing (Demo 7). PROM, EPROM, EEPROM and flash can be programmed, slowly and a limited number of times.

Logic with a memory array

A 2N-word × M-bit ROM is a lookup table: the address is the input, the word is the output, and each column is one function of N inputs. Reading all four words (Demo 6) fills in the truth table of the book's ROM:

A1 A0Data2 = A1 ⊕ A0Data1 = A̅1 + A0Data0 = A̅1 · A̅0
00011
01110
10100
11010

Register files and ports

Each port of a memory is one address input with its own data lines. The MIPS register file is a 32-word × 32-bit array with three ports: read ports A1 → RD1 and A2 → RD2, and a write port A3, WD3 with write enable WE3. In one cycle an add rd, rs, rt reads rs and rt and writes rd (Demo 8: RD1 = 5, RD2 = 7, $t2 = 12). This is the register file of the single-cycle MIPS processor.

Each port has its own 5:32 decoder, and every cell gets one wordline and its own bitlines per port, so a register-file cell is larger than an SRAM cell. The reads are combinational. The write happens at the rising clock edge that ends the cycle. So add $t0, $t0, $t0 reads the old value 5 on both ports and writes 10 (Demo 9). Register $0 has no write wordline. It always reads 0, and writes to it are ignored (Demo 10).

Programmable logic arrays

A ROM used as logic spends one row on every one of the 2N input combinations: its decoder is a complete, fixed AND plane. A PLA makes the AND plane programmable too. It builds only the product terms it needs, then ORs them in an OR plane. The book's 3 × 3 × 2 PLA computes X = A̅B̅C + ABC̅ and Y = AB̅ with 3 product terms. The same functions as a ROM need 8 rows (Demos 11 and 12: both always agree; X = 1 only for 001 and 110, Y only for 100 and 101).

FPGAs: lookup tables and logic elements

An FPGA is an array of configurable logic elements (LEs) joined by programmable wiring. Each LE has a lookup table (LUT): 16 SRAM cells and a 16:1 mux whose select lines are the four inputs. In other words, a 16 × 1 memory. Any function of four inputs is just its truth table written into the cells. The LE also has a flip-flop and an output mux that chooses the combinational or the registered value.

The configuration is a bitstream. Loading it writes the LUT cells and the mux settings. In Demo 13 the LUT is programmed for X (1s in cells 2, 3, 12 and 13) and answers 1 for A B C D = 0010. In Demo 14 the same LUT is reprogrammed as 4-input parity, and the output for 1100 changes from 1 to 0 without any rewiring. In Demo 15 the output is registered: Y changes only at a rising clock edge.

ROMPLAFPGA
AND planefixed decoder, all 2N mintermsprogrammable, only the terms usednone: LUTs hold truth tables
Programmedat manufacture (mask) or onceat manufacture or onceat every power-up (SRAM), any number of times
Statenonoflip-flop in every LE

What the page leaves out

  • The voltages, the 1 : 9 capacitance ratio, the leak of 10 % per 16 ms and the 50 % threshold are illustrative. Real DRAM cells keep their charge far longer than 64 ms. The refresh interval is set by the weakest cells and by temperature.
  • One word per row. Real DRAM opens a whole row of thousands of bits into a row buffer, and a column mux picks the word (row hits and misses, banks, burst transfers, the DDR protocol and refresh commands are not shown).
  • Transistor-level detail: the access transistor's threshold drop, how a sense amplifier latches, the read-stability and write-margin sizing of the SRAM transistors, and leakage paths.
  • Only 8 of the 32 registers are drawn. Pipelined MIPS writes the register file in the first half of the cycle and reads it in the second (see the MIPS Pipeline page). Here the write is at the end of the cycle, as in the single-cycle processor.
  • An FPGA's programmable routing, carry chains, block RAMs, DSP blocks and I/O are not drawn. Only one LE is shown.
  • Flash, EEPROM, 3D NAND and new nonvolatile memories (MRAM, PCM) are only mentioned.

One level up, DRAM and SRAM become the main memory and caches of the CPU Cache page. On the Levels of Abstraction page, a printf ends up in DRAM cells like these.