The idea: a word is a bit pattern, the format gives it a value

Hardware stores only bits. The same 8 bits 1001 0110 are 150 as an unsigned number, −106 in two's complement, 9.375 as an unsigned fixed-point number with four fraction bits, and something else again inside a floating-point number. The format is a convention between the program and the hardware; the adder does not know it.

The bits 1001 0110 weighted three ways: unsigned weights 128 to 1 give 150; two's complement with top weight -128 gives -106; fixed point 4.4 with weights 8 to 1/16 gives 9.375
The same eight bits are 150, −106 or 9.375: only the weight given to each bit changes.

The page has one tab per format. Each tab starts with a sample, has its own demos, and lets you click any bit to flip it. The working panel lists every step; the narration under the canvas explains it.

Unsigned and two's complement integers

An N-bit unsigned number gives bit i the weight 2i, so 8 bits cover 0 … 255. Two's complement keeps every weight except the top one, which becomes −2N−1: 8 bits cover −128 … 127. Positive numbers look the same in both; a 1 in the top bit means "subtract 128".

Nunsignedtwo's complement
40 … 15−8 … 7
80 … 255−128 … 127
160 … 65,535−32,768 … 32,767
320 … 4,294,967,295−2,147,483,648 … 2,147,483,647

To negate a two's complement number, invert every bit and add 1. Inverting gives −x − 1, because x plus its inverse is all ones, which is −1. In Demo 1: −3 + 5, −3 is built from 3 = 0000 0011: inverted 1111 1100, plus 1 1111 1101. Zero has only one pattern, and −128 has no positive partner: negating it gives −128 again.

Sign extension widens a number without changing its value by copying the sign bit into the new bits (Demo 3: sign extension): −5 in 4 bits is 1011, in 8 bits 1111 1011. Filling with zeros instead (zero extension) gives 0000 1011 = 11, which is right only for unsigned numbers. MIPS sign-extends the 16-bit immediate of addi and lw, and the bytes of lb; lbu zero-extends.

Carry out and overflow

One adder serves both formats. A carry out of the top bit (c8) means the unsigned result does not fit. For two's complement a carry out is normal and is dropped: −3 + 5 gives 1 0000 0010, and the 8 bits left are 2, the right answer (Demo 1).

Signed overflow happens only when two numbers of the same sign give a result of the other sign. The hardware test is V = c7 XOR c8: the carry into the sign bit differs from the carry out of it. In Demo 2: overflow, 100 + 50 = 1001 0110: 150 unsigned (right, c8 = 0) but −106 in two's complement (wrong, V = 1 XOR 0 = 1). MIPS add traps on signed overflow; addu ignores it; C leaves signed overflow undefined. The Combinational Logic page shows the same V signal on a 4-bit ripple-carry adder, gate by gate.

A + B (8 bits)result bitsc8Vunsignedtwo's complement
−3 + 50000 001010253 + 5 → 2 ✗2 ✓
100 + 501001 011001150 ✓−106 ✗

Fixed-point numbers

A fixed-point number is an integer with an implied binary point. In the 4.4 format of the page, 4 bits come before the point and 4 after, worth 1/2, 1/4, 1/8 and 1/16; the value is the 8-bit integer divided by 16. H&H writes Ua.b for an unsigned format with a integer and b fraction bits and Qa.b for two's complement (some texts count the sign bit separately).

  • 6.75 = 4 + 2 + 1/2 + 1/4 = 0110.1100: the integer part as usual, the fraction by doubling: 0.75 × 2 = 1.5 → 1, 0.5 × 2 = 1.0 → 1, then 0, 0.
  • −2.375 in Q4.4: 2.375 = 0010.0110, inverted 1101.1001, plus one unit of the last place (1/16): 1101.1010. In sign/magnitude it would be 1010.0110.
  • Adding fixed-point numbers is plain integer addition (Demo 4: 6.75 − 2.375): 0110.1100 + 1101.1010 = 0100.0110 = 4.375, with a carry out that two's complement drops.

Only multiples of 1/16 exist. 0.1 doubles to 0.2, 0.4, 0.8, 1.6, … and never ends, so the page keeps 4 bits and cuts the rest off (truncation): 0.1 becomes 0000.0001 = 0.0625 and 0.2 becomes 0.1875. Their sum is 0.25, far from 0.3 (Demo 5: 0.1 + 0.2). Fixed point is fast and simple, which is why DSPs and audio codecs use it, but its range and precision are fixed when the format is chosen.

IEEE 754 single precision

Floating point is scientific notation in base 2: x = ±1.f × 2e. A 32-bit single-precision number stores three fields:

FieldBitsMeaning
sign s310 positive, 1 negative (sign/magnitude)
exponent E30 … 23 (8 bits)biased: e = E − 127, so E = 1 … 254 gives e = −126 … 127
fraction f22 … 0 (23 bits)the bits after the binary point of 1.f; the leading 1 is implicit (not stored)

value = (−1)s × 1.f × 2E − 127. Every nonzero binary number begins with 1, so storing it would waste a bit: the 23 stored bits give 24 bits of precision, about 7 decimal digits. The bias makes the exponent an unsigned field, so positive floats sort like integers.

The 32 bits of 0xC2690000 split into sign 1, exponent 10000100 = 132 and fraction 1101001 followed by zeros; decoded: negative, e = 5, mantissa 1.1101001, so -1.1101001 times 2 to the 5 = -58.25
A float is a sign bit, a biased exponent and the fraction after a hidden leading 1: 0xC2690000 decodes to −1.1101001₂ × 25 = −58.25.

Encoding, as in H&H (Demo 6: 228, −58.25):

Step228−58.25
signs = 0s = 1, go on with 58.25
binary11100100₂111010.01₂
normalize1.11001₂ × 271.1101001₂ × 25
biased exponent7 + 127 = 134 = 100001105 + 127 = 132 = 10000100
fraction1100100000000000000000011010010000000000000000
result0x436400000xC2690000

0.1 is the hard case (Demo 7: 0.1). Doubling 0.1 gives the bits 0.000110011001100…, repeating forever. Normalized: 1.10011001100110011001100 | 1100… × 2−4, so E = 123 = 01111011. The 23 fraction bits are followed by 1100…, more than half a unit in the last place, so the fraction is rounded up to …1101: 0x3DCCCCCD. That float is exactly 0.100000001490116119384765625, too large by about 1.49 × 10−9.

Special values and denormals

The smallest and the largest exponent fields are reserved (Demo 8: special values):

EfValueExample
00±00x00000000, 0x80000000 (−0 == +0, but 1/−0 = −∞)
0≠ 0denormal: ±0.f × 2−1260x00000001 = 2−149 ≈ 1.4 × 10−45
1 … 254anynormal: ±1.f × 2E − 1270x00800000 = 2−126, 0x7F7FFFFF ≈ 3.4028235 × 1038
2550±∞0x7F800000: overflow, 1/0
255≠ 0NaN0x7FC00000: 0/0, ∞ − ∞, √−1; NaN ≠ NaN

Denormals (subnormals) have no hidden 1 and keep the scale at 2−126. Without them the next number below 2−126 would be 0, and x − y could be 0 for x ≠ y. With them the gap to zero is filled with evenly spaced numbers (gradual underflow), at the cost of precision: 2−149 has a single significant bit.

Rounding

A result that needs more than 24 significant bits must be rounded to a neighbour. IEEE 754 defines four modes, all selectable on the page (rounding menu):

ModeRuleUsed for
round to nearest, ties to eventhe nearer neighbour; on an exact tie, the one whose last bit is 0the default everywhere
toward zerodrop the extra bits (truncate)float-to-int conversion in C
toward +∞ (up)the neighbour aboveinterval arithmetic (upper bound)
toward −∞ (down)the neighbour belowinterval arithmetic (lower bound)

Ties to even avoids a bias: always rounding halves up would push long sums upward. The hardware needs only three extra bits to decide: the guard bit (the first bit past the mantissa), the round bit (the next one), and the sticky bit, the OR of every bit after that. G = 0: round down. G = 1 and R or S = 1: round up. G = 1, R = S = 0: a tie, round to even.

Floating-point addition

H&H's algorithm for adding two floats (the FP addition tab, checklist on the left):

  1. Extract the exponent and fraction bits.
  2. Prepend the leading 1 to form the mantissa.
  3. Compare the exponents.
  4. Shift the smaller mantissa right by the difference.
  5. Add the mantissas.
  6. Normalize the mantissa and adjust the exponent if necessary.
  7. Round the result.
  8. Assemble the exponent and fraction back into a floating-point number.

Demo 9: 1.5 + 3.25 is the book's example, 0x3FC00000 + 0x40500000. The mantissas are 1.1₂ × 20 and 1.101₂ × 21. The first is shifted right by 1 to 0.11₂ × 21; the sum is 10.011₂ × 21; normalizing gives 1.0011₂ × 22; nothing needs rounding; E = 2 + 127 = 129, and the result is 0x40980000 = 4.75.

Adding 1.5 and 3.25: 1.1 times 2 to the 0 is shifted to 0.11 times 2 to the 1, added to 1.101 times 2 to the 1 giving 10.011 times 2 to the 1, normalized to 1.0011 times 2 to the 2 = 4.75
Floating-point addition lines up the binary points by shifting the smaller number, adds, then shifts the sum back to one digit before the point.

When the signs differ, the smaller magnitude is subtracted from the larger and the result takes the larger one's sign. Close numbers then cancel and the result is shifted left to normalize. Type 1.0000001 and -1 into A and B to see it. With a large exponent difference, the small number falls off the end into the sticky bit: 16777216.0 + 1 (224 + 1) is exactly a tie and rounds to even, back to 16777216. (Eight hex digits are read as a bit pattern; anything else as a decimal.)

The page's adder is bit-exact: for every pair it gives the same bits as a real FPU. Infinity and NaN operands skip the steps (∞ + x = ∞, ∞ − ∞ = NaN), and x + (−x) = +0.

Why 0.1 + 0.2 ≠ 0.3

None of the three numbers has a finite binary form, so each is rounded when it is stored, and the sum is rounded once more (Demo 10: 0.1 + 0.2). The exact values in single precision:

DecimalBitsExact stored value
0.10x3DCCCCCD0.100000001490116119384765625
0.20x3E4CCCCD0.20000000298023223876953125
0.1 + 0.2 (exact sum of the two)needs 26 bits0.300000004470348358154296875
0.1 + 0.2 (rounded)0x3E99999A0.300000011920928955078125
0.30x3E99999A0.300000011920928955078125
0.60x3F19999A0.60000002384185791015625
0.1 + 0.6 (rounded)0x3F3333340.7000000476837158203125
0.70x3F3333330.699999988079071044921875

In single precision 0.1 + 0.2 happens to round to the same float as 0.3, so 0.1f + 0.2f == 0.3f is true, although none of them is 0.3. But 0.1 + 0.6 lands one unit in the last place above 0.7. In double precision (JavaScript, Python, C double) it is the other way round: 0.1 + 0.2 = 0.30000000000000004 ≠ 0.3. The lesson is the same: compare floats with a tolerance, and use integers (cents) or decimal types for money.

What the page leaves out

  • Double precision (1 + 11 + 52 bits, bias 1023) works the same way with wider fields; half precision and bfloat16 (used in machine learning) too.
  • Multiplication and division: multiply the mantissas, add the exponents (subtracting one bias). H&H leaves them to exercises as well.
  • Signalling vs quiet NaNs, NaN payloads, the exception flags (inexact, overflow, underflow, invalid, divide by zero) and traps.
  • The MIPS floating-point coprocessor: registers $f0–$f31, add.s, lwc1, c.lt.s.
  • Fixed point is limited to 8-bit words and truncation; real fixed-point code often rounds and saturates instead of wrapping.
  • The integer adder is drawn as a column-by-column ripple; real adders use carry-lookahead or prefix adders.

Bits as a code for something else again: Base64 Encoding regroups bytes into 6-bit characters. Where the adder sits in the processor: The Single-Cycle MIPS Processor.