LogicGates.org Open the simulatorSimulator

Base32 encode and decode

Turn text or bytes into Base32 or Base32 back into bytes, in the RFC 4648, base32hex or Crockford alphabet. Every group is drawn: five bytes, their 40 bits, the same bits cut into fives, and the character each five picks.

Alphabet
Padding
Input

Any text, including accents and emoji. It is turned into UTF-8 bytes first, and the bytes are encoded.

5 bytes, 8 characters

Step by step

Top to bottom: each byte, its 8 bits, the same 40 bits cut into fives, each five as a number from 0 to 31, and the character for that number. Colours follow each byte's bits into the fives; grey bits are zeros added to fill the last five.

48H65e6Cl6Cl6Fo010010000110010101101100011011000110111101001000011001010110110001101100011011119118222427315JBSWY3DP

Each group is 40 bits wide: scroll sideways to see all of it.

How Base32 works: 5 bytes become 8 characters

A Base32 character carries 5 bits, since 25 = 32, and a byte is 8. The smallest run of bits that is a whole number of both is 40: five bytes, or eight characters. So the encoder takes the bytes five at a time, writes out their 40 bits, cuts them into eight fives and looks each one up in the alphabet. Here is "Hello":

48H65e6Cl6Cl6Fo010010000110010101101100011011000110111101001000011001010110110001101100011011119118222427315JBSWY3DP

Each group is 40 bits wide: scroll sideways to see all of it.

H, e, l, l and o are the bytes 72, 101, 108, 108, 111. Their 40 bits cut into fives are 01001 00001 10010 10110 11000 11011 00011 01111, which are 9, 1, 18, 22, 24, 27, 3, 15, and in the RFC 4648 alphabet those are JBSWY3DP. Decoding runs the same steps upwards. Text goes through UTF-8 first, so an accented letter is two bytes before Base32 sees it.

The three Base32 alphabets

All three map the numbers 0 to 31 to characters; only the characters differ. Data encoded with one alphabet must be decoded with the same one.

Value Bits RFC 4648 base32hex Crockford
0 00000 A 0 0
1 00001 B 1 1
2 00010 C 2 2
3 00011 D 3 3
4 00100 E 4 4
5 00101 F 5 5
6 00110 G 6 6
7 00111 H 7 7
8 01000 I 8 8
9 01001 J 9 9
10 01010 K A A
11 01011 L B B
12 01100 M C C
13 01101 N D D
14 01110 O E E
15 01111 P F F
16 10000 Q G G
17 10001 R H H
18 10010 S I J
19 10011 T J K
20 10100 U K M
21 10101 V L N
22 10110 W M P
23 10111 X N Q
24 11000 Y O R
25 11001 Z P S
26 11010 2 Q T
27 11011 3 R V
28 11100 4 S W
29 11101 5 T X
30 11110 6 U Y
31 11111 7 V Z
  • RFC 4648 Base32 is the one meant by plain "Base32": capitals A–Z for 0 to 25, then 2–7. It is the alphabet of 2FA secret keys and of .onion addresses.
  • base32hex, from the same RFC, continues hexadecimal: 0–9, then A–V. Because the characters are in ASCII order, encoded strings sort the same way as the bytes they hold, as long as they are left unpadded (an = sorts after the digits). The bytes 0F, 80, FF are B4======, QA======, 74====== in RFC 4648 Base32, which sort as FF, 0F, 80: the alphabet puts 2–7 after Z, but ASCII puts digits before letters. In base32hex they are 1S======, G0======, VS======, which sort as 0F, 80, FF.
  • Crockford Base32, designed by Douglas Crockford for people to read and type, is 0–9 and the letters without I, L, O and U. A decoder ignores case and hyphens and reads I and L as 1 and O as 0, so a code read over the phone still decodes. It has no padding, and may end in a check symbol for the number modulo 37: one of the 32 characters or one of five extras, * ~ $ = U, which this decoder checks. His scheme encodes a number, with any spare bits at the front; this page, like many libraries, applies the alphabet to bytes the RFC 4648 way, with the spare bits at the end. The two agree whenever the length is a multiple of 8 characters. ULIDs, the sortable IDs the UUID and ULID decoder takes apart, are one 128-bit number in 26 characters, so 2 zero bits come first; paste one in Crockford decode mode and the page reads it as that number, with the bytes cut from the left a click away.

Padding: when the bytes do not fill a group

When fewer than five bytes are left, the last group is short. Its bits are filled out with zeros to a whole five, and = signs take the place of the characters with no bits at all, so the output is always a multiple of eight. These are the test vectors from RFC 4648:

Input Base32 base32hex = signs Zero bits addedFill bits Bytes Bits Characters
f MY====== CO====== 6 2 1 8 2
fo MZXQ==== CPNG==== 4 4 2 16 4
foo MZXW6=== CPNMU=== 3 1 3 24 5
foob MZXW6YQ= CPNMUOG= 1 3 4 32 7
fooba MZXW6YTB CPNMUOJ1 0 0 5 40 8

A decoder can tell how many bytes the last group holds from its number of characters alone, so the padding carries no information and is often left off, as in 2FA secrets. A group can never end after 1, 3 or 6 characters: no number of bytes produces that, which is how this decoder spots a missing or extra character.

How much bigger Base32 makes data

Every 5 bytes become 8 characters, so Base32 is 60% bigger than the data, against a third for Base64.

Bytes Base32 Without padding Base64
1 8 2 4
5 8 8 8
10 16 16 16
16 32 26 24
20 32 32 28
32 56 52 44
100 160 160 136
1,000 1,600 1,600 1,336

The rule: 8 × ⌈n ÷ 5⌉ characters for n bytes with padding, ⌈8n ÷ 5⌉ without. A 20-byte TOTP secret is exactly 32 characters, with no padding needed.

Base32 in two-factor authentication

An authenticator app and a website share a secret key, a run of random bytes. The QR code you scan when you set up 2FA holds an otpauth:// link with the key in Base32, and the "enter this key instead" text under it is the same Base32, often in blocks of four and in small letters. Base32 is used because the key may have to be typed by hand, and an alphabet with one case and no 0, 1 or 8 is hard to mistype. The example key JBSWY3DPEHPK3PXP is 10 bytes: "Hello!" followed by DE AD BE EF.

Common mistakes

  • Decoding with the wrong alphabet. A string can be valid in more than one alphabet, and then it decodes without an error into the wrong bytes. "abc" is MFRGG=== in RFC 4648 Base32; read as base32hex, the same characters give the bytes B3 F7 08, not 61 62 63.
  • Typing 0, 1 or 8 in RFC 4648 Base32. They are not in the alphabet (nor is 9). If a key seems to contain them, they are almost certainly the letters O, I and B.
  • Comparing Base32 strings as case-sensitive. Most decoders, this one included, read small letters as capitals, so jbswy3dp and JBSWY3DP are the same bytes. Some refuse small letters, such as Python's base64.b32decode unless it is given casefold=True, so write capitals when in doubt.
  • Treating it as encryption. Base32, like Base64, has no key; anyone can decode it.

Questions

What is Base32 used for?

Mostly for things people read, type or say, or that pass through systems that ignore case. The secret keys for two-factor authentication apps (TOTP) are Base32, such as JBSWY3DPEHPK3PXP, which decodes to the bytes 48 65 6C 6C 6F 21 DE AD BE EF. ULIDs use Crockford's Base32, and Tor's .onion addresses are Base32 in small letters.

Why does Base32 use 2 to 7 and no 0, 1, 8 or 9?

RFC 4648 Base32 needs 32 characters. The 26 capital letters give 26, and the other six are the digits 2 to 7. Leaving out 0, 1 and 8 means the result never has a digit that could be mistaken for the letters O, I or B, and once 2 to 7 fill the six places, 9 is not needed.

Why does Base32 end with ======?

Base32 turns every 5 bytes into 8 characters. When the data is not a multiple of 5 bytes, the last group is short and = fills it to 8 characters: one leftover byte gives 2 characters and six =, two give 4 and four =, three give 5 and three =, four give 7 and one =. Many systems, TOTP apps among them, leave the padding off.

What is the difference between Base32, base32hex and Crockford Base32?

RFC 4648 Base32 is A–Z then 2–7. base32hex is 0–9 then A–V, so encoded strings without padding sort in the same order as the bytes. Crockford's alphabet is 0–9 and the letters without I, L, O and U, and his scheme also has reading rules: a decoder reads I and L as 1 and O as 0, ignores hyphens and case, there is no padding, and an optional check symbol may end the string. "Hello" is JBSWY3DP, 91IMOR3F and 91JPRV3F in the three. Crockford's own scheme, and ULIDs, write a number rather than bytes, with the spare bits at the front instead of the end, so unless the length is a multiple of 8 characters the bits line up differently.

How much bigger does Base32 make data?

Eight characters for every 5 bytes, so 60% bigger, against a third bigger for Base64. The price buys an alphabet with no lower case and no symbols, which survives being read aloud, typed into a phone or used in a case-insensitive file name or DNS label.

Why did my Base32 fail to decode?

Usually a character from a different alphabet (a 0, 1, 8 or 9 in RFC 4648 Base32, or a W to Z in base32hex), a length that leaves 1, 3 or 6 characters after the groups of eight, which no number of bytes produces, or the wrong number of = signs. The error names which one and marks the character. Spaces, line breaks and small letters are fine.

For 6 bits per character, see Base64; for the alphabet that drops 0, O, I and l and works by division instead, Base58; and for base 36, the base 36 converter.