The idea: carry bytes through places that only accept text

Email bodies, JSON strings, URLs, HTTP headers and HTML attributes were made for text. Raw bytes break them: a zero byte, a line break or a quote ends the field early. Base64 writes any bytes using only 64 safe characters (A–Z a–z 0–9 + /), so they can pass through unchanged. The price is size: every 3 bytes become 4 characters.

3 bytes → 4 characters

64 = 26, so one character carries exactly 6 bits. The smallest number of bytes that splits evenly into 6-bit pieces is 3: 3 × 8 = 24 = 4 × 6. The encoder takes 3 bytes, writes their 24 bits in a row, reads them back 6 at a time, and looks each 6-bit number (0–63) up in the alphabet. Demo: Man → TWFu is the textbook example:

Other ways to read a group of bits as a number (unsigned, two's complement, fixed point, IEEE 754 floating point) are animated in Number Systems: Two's Complement, Fixed Point and IEEE 754.

textMan
bytes0x4D = 010011010x61 = 011000010x6E = 01101110
6-bit pieces010011 = 19010110 = 22000101 = 5101110 = 46
charactersTWFu
The bytes of Man, 0x4D 0x61 0x6E, written as 24 bits in three groups of 8, then the same bits regrouped into four groups of 6 with values 19, 22, 5 and 46, which become the characters T, W, F and u
Base64 regroups the same 24 bits from three 8-bit bytes into four 6-bit numbers, each one a character.

The alphabet is in index order: A–Z are 0–25, a–z are 26–51, 0–9 are 52–61, then + = 62 and / = 63.

Padding and the size formula

When the input length is not a multiple of 3, the last group has 1 or 2 bytes. The encoder adds zero bits up to the next 6-bit boundary and writes = for each character that carries no data (Demo: padding):

last groupbits+ zero bitscharactersexample
3 bytes2404Man → TWFu
2 bytes1623 + =Ma → TWE=
1 byte842 + ==M → TQ==
Ma is 16 bits; 2 zero bits are added to make three 6-bit groups T, W, E, then one = gives TWE=. M is 8 bits; 4 zero bits make two groups T, Q, then == gives TQ==
A short last group is filled with zero bits up to a 6-bit boundary, and = fills the missing characters.

With padding, n bytes always become 4·⌈n/3⌉ characters, about 33 % more. Without padding they become ⌈4n/3⌉. Padding keeps the length a multiple of 4. That lets a decoder read 4 characters at a time and lets you join two encoded strings safely. Since the decoder can work out the byte count from the length anyway, many formats leave the padding out.

Text is bytes first: UTF-8

Base64 encodes bytes. To encode text, you first have to pick a character encoding, and today that is almost always UTF-8. ASCII letters are one byte each, but รฉ is two (C3 A9) and ๐Ÿ‘‹ is four (F0 9F 91 8B). So Hi ๐Ÿ‘‹ is 7 bytes and encodes to SGkg8J+Riw== (Demo: UTF-8 first). JavaScript's btoa() accepts only characters below 256 and throws on ๐Ÿ‘‹. Encode with TextEncoder first, or use Uint8Array.prototype.toBase64().

Variants: standard, base64url, MIME

variant62, 63paddingline breaksused in
standard (RFC 4648 §4)+ /yesnodata: URIs, JSON, Basic auth
base64url (RFC 4648 §5)- _usually nonoJWT, URLs, file names, WebAuthn
MIME (RFC 2045)+ /yesCRLF every 76 charactersemail attachments
PEM (RFC 7468)+ /yesevery 64 characterscertificates and keys

In a URL, + can be read as a space and / as a path separator, so base64url swaps them for - and _. The two alphabets differ only in those two characters. A string that contains neither is valid in both. Demo: base64url encodes the same bytes both ways, then shows a base64url decoder rejecting a +.

Decoding, and the inputs it rejects

Decoding runs the steps backwards: look each character up, write 6 bits each, cut the bits into bytes, and drop the bits that belong to the padding (Demo: decode). Some inputs cannot come from a correct encoder (Demo: bad input):

inputproblemstrict (RFC 4648)lenient (atob, WHATWG)
TW@u@ is not in the alphabeterrorerror
TWElength not a multiple of 4 (missing =)errorMa
T6 bits cannot make a byteerrorerror
TWF=leftover bits 01 are not zeroerror (non-canonical)Ma
TQ=A= in the middleerrorerror

Rejecting non-zero leftover bits matters when Base64 strings are compared or signed. With a lenient decoder, TWE= and TWF= are two different strings that decode to the same bytes.

Encoding, not encryption

Base64 has no key. Anyone can reverse it. Authorization: Basic dXNlcjpwYXNz sends user:pass in readable form (Demo: not encryption), so Basic auth is only safe over HTTPS. A JWT's header and payload are base64url JSON that anyone can read. Only the signature protects them, and it stops changes, not reading (see OAuth). Base64 is not compression either: it makes data bigger. To make data smaller, use a code like Huffman coding.

encodingbits per charactersize for 3 bytesoverhead
Base16 (hex)46+100 %
Base3254.8 (8 per 5 bytes)+60 %
Base6464+33 %
Ascii85 / Z85~6.43.75 (5 per 4 bytes)+25 %

Where you meet Base64: email attachments, data: URIs for small images, PEM certificates, JSON fields that carry binary data, JWTs, cookie values and HTTP Basic auth.

What the page leaves out

The page handles at most 24 bytes (8 groups). It removes whitespace in one step and does not show MIME line wrapping. It does not cover streaming decoders that keep up to 3 characters between chunks, or the SIMD codecs that encode many groups at once. It also leaves out other alphabets, such as the one bcrypt and crypt(3) use.