FP32
1 · 8 · 23, bias 127- Hex
- 3DCCCCCD
- Stored
- 0.100000001490116119384765625
- Error
- 1.49e−9 (1.49e−6% relative) rounded up
- Kind
- normal
- Neighbours
- below0.099999994 stored0.1 above0.10000001
Type a decimal to see it in every low-precision float used in machine learning at once: the bits, the value each format actually stores, and the rounding error. Or type a bit pattern to read it back.
Any decimal, such as 0.1, -2.5e-3 or 65504, with a point for the decimal, and also Infinity and NaN. It is rounded straight to each format, exactly, never through a JavaScript number.
FP16 of 0.1: hex 2E66, normal, 1.1001100110 × 2⁻⁴. Click a bit to flip it; the space after every fourth bit from the right marks a hex digit.
In binary the magnitude is 1.10011001100110011001… × 2⁻⁴. FP16 keeps 10 bits after the point.
Sticky is 1 if any bit after the round bit is 1 (10011001…). The guard bit is 0, so the rest is less than half a step: drop it, which rounds down (truncation). Result: 1.1001100110 × 2⁻⁴ = 0.0999755859375.
Every one of these is the same idea as the 32 and 64 bit floats on the IEEE 754 converter: a sign bit, an exponent stored with a bias, and the bits after the point of a binary number 1.mmm. They differ only in how many bits go to the exponent, which sets the range, and how many to the mantissa, which sets the precision. Every figure in this table is computed by the converter's engine.
sign exponent mantissa, drawn to scale. BF16's exponent is exactly as wide as FP32's.
| Format | Bits S·E·M, bias | Largest | Smallest normal | Smallest subnormal | Epsilon | Digits | Specials |
|---|---|---|---|---|---|---|---|
| FP32single precision, float | 1·8·23, 127 | 3.4028e38 | 2⁻¹²⁶≈ 1.1755e-38 | 2⁻¹⁴⁹≈ 1.4013e-45 | 2⁻²³ | 7.2 | ±∞16,777,214 NaN codes |
| FP16IEEE binary16, half precision | 1·5·10, 15 | 65504 | 2⁻¹⁴≈ 6.1035e-5 | 2⁻²⁴≈ 5.9605e-8 | 2⁻¹⁰ | 3.3 | ±∞2,046 NaN codes |
| BF16bfloat16, brain floating point | 1·8·7, 127 | 3.3895e38 | 2⁻¹²⁶≈ 1.1755e-38 | 2⁻¹³³≈ 9.1835e-41 | 2⁻⁷ | 2.4 | ±∞254 NaN codes |
| FP8 E4M3OCP E4M3FN | 1·4·3, 7 | 448 | 2⁻⁶≈ 0.015625 | 2⁻⁹≈ 0.0019531 | 2⁻³ | 1.2 | no ∞2 NaN codes |
| FP8 E5M2OCP E5M2 | 1·5·2, 15 | 57344 | 2⁻¹⁴≈ 6.1035e-5 | 2⁻¹⁶≈ 1.5259e-5 | 2⁻² | 0.9 | ±∞6 NaN codes |
| FP4 E2M1OCP MX FP4 | 1·2·1, 1 | 6 | 2⁰≈ 1 | 2⁻¹≈ 0.5 | 2⁻¹ | 0.6 | no ∞no NaN |
Epsilon is the gap between 1 and the next value up, 2 to the minus number of mantissa bits. Digits is the significand's bits, hidden 1 included, times log₁₀ 2: roughly how many significant decimal digits survive.
For a normal number, the value is (−1)sign × 1.mantissa × 2exponent − bias. The 1 before the point is not stored, because a normal number always has it. When the exponent field is all zeros the number is subnormal: the hidden bit becomes 0 and the exponent stays at 1 − bias, so the values fade evenly down to zero instead of stopping short of it.
The formats disagree about the all-ones exponent. The IEEE-style ones (FP32, FP16, BF16 and E5M2) reserve it: mantissa zero is infinity, anything else NaN. E4M3 keeps only one NaN, S.1111.111, and uses the rest of the top exponent for ordinary numbers, which is how it reaches 448 instead of 240. LLVM and PyTorch call it E4M3FN: F for finite, N for its non-IEEE NaN. FP4 E2M1 reserves nothing at all: every code is a number.
A decimal almost never lands on a value the format can store, so it is rounded to the nearest one. When it is exactly halfway between two, it goes to the one whose last mantissa bit is 0, the even one. Always rounding halves up would push sums upwards on average; ties to even goes up half the time and down half the time.
Hardware does this with three extra bits. Beyond the bits it keeps, it remembers the next bit (the guard), the one after that (the round bit) and whether anything at all is set further down (the sticky bit, an OR of the rest). Guard 0 means less than half a step: truncate. Guard 1 with round or sticky set means more than half: round up. Guard 1 with both clear is an exact tie, and the last kept bit decides.
kept 1.000 · G 1 · R 0 · S 0
→ 1.000 × 2⁰ =
1
Guard is 1 and round and sticky are 0: exactly halfway. The last kept bit is 0 (even), so keep it. Try it
kept 1.001 · G 1 · R 0 · S 0
→ 1.010 × 2⁰ =
1.25
Guard is 1 and round and sticky are 0: exactly halfway. The last kept bit is 1 (odd), so round to the even pattern, which is up. Try it
kept 1.000 · G 1 · R 0 · S 1
→ 1.001 × 2⁰ =
1.125
Guard is 1 and round is 0, which would be a tie, but sticky is 1: something further down is set, so the rest is more than half a step. Round up. Try it
Rounding works as if the exponent could keep growing, and only then checks whether the result fits. So the cut-off is not the largest value itself but halfway to the step above it. In FP16 that is 65520: exactly halfway, and the even neighbour is the pattern above 65504, infinity. In E4M3 it is 464, halfway between 448 and the 480 that S.1111.111 would have been if it were not NaN; here the even neighbour is 448, so 464 stays finite and 465 does not.
Formats with infinity overflow to it. E4M3 and FP4 have none. For E4M3 the OCP 8-bit floating point specification allows either saturating to ±448 or producing NaN, and it also lets conversions to E5M2 saturate instead of overflowing. This converter shows: infinity for FP32, FP16, BF16 and E5M2; for E4M3 whichever of the two you pick above (saturation unless you change it); and for FP4, which has neither infinity nor NaN, the largest value, ±6.
| Format | Largest | Twice the largest becomes | Infinity becomes |
|---|---|---|---|
| FP32 | 3.4028e38 | Infinity | Infinity |
| FP16 | 65504 | Infinity | Infinity |
| BF16 | 3.3895e38 | Infinity | Infinity |
| FP8 E4M3 | 448 | 448 or NaN | 448 or NaN |
| FP8 E5M2 | 57344 | Infinity | Infinity |
| FP4 E2M1 | 6 | 6 | 6 |
BF16 is the top 16 bits of an FP32: the same sign, the same 8 bit exponent with the same bias of 127, and the first 7 of FP32's 23 mantissa bits. So BF16 covers almost the same range as FP32, only less precisely: its largest value is 3.3895e38 against FP32's 3.4028e38, and its subnormals stop at 2⁻¹³³ rather than 2⁻¹⁴⁹. Converting is a matter of rounding off the low 16 bits. FP16 has more precision but a 5 bit exponent, so it runs out at 65504. The same values in all three:
| Value | FP32 hex | BF16 hex | BF16 value | FP16 value |
|---|---|---|---|---|
| 1 | 3F800000 | 3F80 | 1 | 1 |
| 0.1 | 3DCCCCCD | 3DCD | 0.1001 | 0.099976 |
| 3.14159 | 40490FD0 | 4049 | 3.140625 | 3.140625 |
| −0.001 | BA83126F | BA83 | −0.00099945 | −0.0010004 |
| 70000 | 4788B800 | 4789 | 70144 | Infinity |
| 1e38 | 7E967699 | 7E96 | 9.9692e37 | Infinity |
The BF16 pattern is the first four hex digits of the FP32 one, or one more when the low half is past halfway and rounds up, as it does for 0.1.
FP4's 7 positive values (8 with zero) are too few to store anything useful on their own. The OCP microscaling (MX) formats fix that by sharing a scale: a block of 32 values stores one 8-bit scale, an E8M0 number that is nothing but a power of two (an exponent with bias 127 and no mantissa), and each value is stored as an E2M1 after dividing by it. The scale exponent is the power of two of the block's largest magnitude minus 2, the largest exponent of E2M1, so the largest value lands between 4 and 8; anything above 6 is clamped.
Here is a block of 8 values (a real block has 32), worked by the engine. The largest magnitude sets the scale to 2⁻¹, stored as E8M0 code 126:
| Value | ÷ 2⁻¹ | E2M1 bits | E2M1 value | × 2⁻¹ again |
|---|---|---|---|---|
| 0.31 | 0.62 | 0 00 1 | 0.5 | 0.25 |
| −1.2 | −2.4 | 1 10 0 | −2 | −1 |
| 2.5 | 5 | 0 11 0 | 4 | 2 |
| 0.05 | 0.1 | 0 00 0 | 0 | 0 |
| −0.7 | −1.4 | 1 01 1 | −1.5 | −0.75 |
| 1.9 | 3.8 | 0 11 0 | 4 | 2 |
| 0.004 | 0.008 | 0 00 0 | 0 | 0 |
| −2.2 | −4.4 | 1 11 0 | −4 | −2 |
The block costs 4 bits per value plus 8 bits shared by all 32, 4.25 bits per value. Small values in a block with a large one lose the most: they fall into E2M1's coarse steps near zero.
| Bits S E M | Hex | Value | Kind |
|---|---|---|---|
| 0 00 0 | 0 | 0 | zero |
| 0 00 1 | 1 | 0.5 | subnormal |
| 0 01 0 | 2 | 1 | normal |
| 0 01 1 | 3 | 1.5 | normal |
| 0 10 0 | 4 | 2 | normal |
| 0 10 1 | 5 | 3 | normal |
| 0 11 0 | 6 | 4 | normal |
| 0 11 1 | 7 | 6 | normal |
| 1 00 0 | 8 | −0 | zero |
| 1 00 1 | 9 | −0.5 | subnormal |
| 1 01 0 | A | −1 | normal |
| 1 01 1 | B | −1.5 | normal |
| 1 10 0 | C | −2 | normal |
| 1 10 1 | D | −3 | normal |
| 1 11 0 | E | −4 | normal |
| 1 11 1 | F | −6 | normal |
Each row is one magnitude; the negative code is the same with the sign bit set, 0x80 higher.
| Bits S E M | Hex | Negative | Value | Kind |
|---|---|---|---|---|
| 0 0000 000 | 00 | 80 | 0 | zero |
| 0 0000 001 | 01 | 81 | 0.001953125 | subnormal |
| 0 0000 010 | 02 | 82 | 0.00390625 | subnormal |
| 0 0000 011 | 03 | 83 | 0.005859375 | subnormal |
| 0 0000 100 | 04 | 84 | 0.0078125 | subnormal |
| 0 0000 101 | 05 | 85 | 0.009765625 | subnormal |
| 0 0000 110 | 06 | 86 | 0.01171875 | subnormal |
| 0 0000 111 | 07 | 87 | 0.013671875 | subnormal |
| 0 0001 000 | 08 | 88 | 0.015625 | normal |
| 0 0001 001 | 09 | 89 | 0.017578125 | normal |
| 0 0001 010 | 0A | 8A | 0.01953125 | normal |
| 0 0001 011 | 0B | 8B | 0.021484375 | normal |
| 0 0001 100 | 0C | 8C | 0.0234375 | normal |
| 0 0001 101 | 0D | 8D | 0.025390625 | normal |
| 0 0001 110 | 0E | 8E | 0.02734375 | normal |
| 0 0001 111 | 0F | 8F | 0.029296875 | normal |
| 0 0010 000 | 10 | 90 | 0.03125 | normal |
| 0 0010 001 | 11 | 91 | 0.03515625 | normal |
| 0 0010 010 | 12 | 92 | 0.0390625 | normal |
| 0 0010 011 | 13 | 93 | 0.04296875 | normal |
| 0 0010 100 | 14 | 94 | 0.046875 | normal |
| 0 0010 101 | 15 | 95 | 0.05078125 | normal |
| 0 0010 110 | 16 | 96 | 0.0546875 | normal |
| 0 0010 111 | 17 | 97 | 0.05859375 | normal |
| 0 0011 000 | 18 | 98 | 0.0625 | normal |
| 0 0011 001 | 19 | 99 | 0.0703125 | normal |
| 0 0011 010 | 1A | 9A | 0.078125 | normal |
| 0 0011 011 | 1B | 9B | 0.0859375 | normal |
| 0 0011 100 | 1C | 9C | 0.09375 | normal |
| 0 0011 101 | 1D | 9D | 0.1015625 | normal |
| 0 0011 110 | 1E | 9E | 0.109375 | normal |
| 0 0011 111 | 1F | 9F | 0.1171875 | normal |
| 0 0100 000 | 20 | A0 | 0.125 | normal |
| 0 0100 001 | 21 | A1 | 0.140625 | normal |
| 0 0100 010 | 22 | A2 | 0.15625 | normal |
| 0 0100 011 | 23 | A3 | 0.171875 | normal |
| 0 0100 100 | 24 | A4 | 0.1875 | normal |
| 0 0100 101 | 25 | A5 | 0.203125 | normal |
| 0 0100 110 | 26 | A6 | 0.21875 | normal |
| 0 0100 111 | 27 | A7 | 0.234375 | normal |
| 0 0101 000 | 28 | A8 | 0.25 | normal |
| 0 0101 001 | 29 | A9 | 0.28125 | normal |
| 0 0101 010 | 2A | AA | 0.3125 | normal |
| 0 0101 011 | 2B | AB | 0.34375 | normal |
| 0 0101 100 | 2C | AC | 0.375 | normal |
| 0 0101 101 | 2D | AD | 0.40625 | normal |
| 0 0101 110 | 2E | AE | 0.4375 | normal |
| 0 0101 111 | 2F | AF | 0.46875 | normal |
| 0 0110 000 | 30 | B0 | 0.5 | normal |
| 0 0110 001 | 31 | B1 | 0.5625 | normal |
| 0 0110 010 | 32 | B2 | 0.625 | normal |
| 0 0110 011 | 33 | B3 | 0.6875 | normal |
| 0 0110 100 | 34 | B4 | 0.75 | normal |
| 0 0110 101 | 35 | B5 | 0.8125 | normal |
| 0 0110 110 | 36 | B6 | 0.875 | normal |
| 0 0110 111 | 37 | B7 | 0.9375 | normal |
| 0 0111 000 | 38 | B8 | 1 | normal |
| 0 0111 001 | 39 | B9 | 1.125 | normal |
| 0 0111 010 | 3A | BA | 1.25 | normal |
| 0 0111 011 | 3B | BB | 1.375 | normal |
| 0 0111 100 | 3C | BC | 1.5 | normal |
| 0 0111 101 | 3D | BD | 1.625 | normal |
| 0 0111 110 | 3E | BE | 1.75 | normal |
| 0 0111 111 | 3F | BF | 1.875 | normal |
| 0 1000 000 | 40 | C0 | 2 | normal |
| 0 1000 001 | 41 | C1 | 2.25 | normal |
| 0 1000 010 | 42 | C2 | 2.5 | normal |
| 0 1000 011 | 43 | C3 | 2.75 | normal |
| 0 1000 100 | 44 | C4 | 3 | normal |
| 0 1000 101 | 45 | C5 | 3.25 | normal |
| 0 1000 110 | 46 | C6 | 3.5 | normal |
| 0 1000 111 | 47 | C7 | 3.75 | normal |
| 0 1001 000 | 48 | C8 | 4 | normal |
| 0 1001 001 | 49 | C9 | 4.5 | normal |
| 0 1001 010 | 4A | CA | 5 | normal |
| 0 1001 011 | 4B | CB | 5.5 | normal |
| 0 1001 100 | 4C | CC | 6 | normal |
| 0 1001 101 | 4D | CD | 6.5 | normal |
| 0 1001 110 | 4E | CE | 7 | normal |
| 0 1001 111 | 4F | CF | 7.5 | normal |
| 0 1010 000 | 50 | D0 | 8 | normal |
| 0 1010 001 | 51 | D1 | 9 | normal |
| 0 1010 010 | 52 | D2 | 10 | normal |
| 0 1010 011 | 53 | D3 | 11 | normal |
| 0 1010 100 | 54 | D4 | 12 | normal |
| 0 1010 101 | 55 | D5 | 13 | normal |
| 0 1010 110 | 56 | D6 | 14 | normal |
| 0 1010 111 | 57 | D7 | 15 | normal |
| 0 1011 000 | 58 | D8 | 16 | normal |
| 0 1011 001 | 59 | D9 | 18 | normal |
| 0 1011 010 | 5A | DA | 20 | normal |
| 0 1011 011 | 5B | DB | 22 | normal |
| 0 1011 100 | 5C | DC | 24 | normal |
| 0 1011 101 | 5D | DD | 26 | normal |
| 0 1011 110 | 5E | DE | 28 | normal |
| 0 1011 111 | 5F | DF | 30 | normal |
| 0 1100 000 | 60 | E0 | 32 | normal |
| 0 1100 001 | 61 | E1 | 36 | normal |
| 0 1100 010 | 62 | E2 | 40 | normal |
| 0 1100 011 | 63 | E3 | 44 | normal |
| 0 1100 100 | 64 | E4 | 48 | normal |
| 0 1100 101 | 65 | E5 | 52 | normal |
| 0 1100 110 | 66 | E6 | 56 | normal |
| 0 1100 111 | 67 | E7 | 60 | normal |
| 0 1101 000 | 68 | E8 | 64 | normal |
| 0 1101 001 | 69 | E9 | 72 | normal |
| 0 1101 010 | 6A | EA | 80 | normal |
| 0 1101 011 | 6B | EB | 88 | normal |
| 0 1101 100 | 6C | EC | 96 | normal |
| 0 1101 101 | 6D | ED | 104 | normal |
| 0 1101 110 | 6E | EE | 112 | normal |
| 0 1101 111 | 6F | EF | 120 | normal |
| 0 1110 000 | 70 | F0 | 128 | normal |
| 0 1110 001 | 71 | F1 | 144 | normal |
| 0 1110 010 | 72 | F2 | 160 | normal |
| 0 1110 011 | 73 | F3 | 176 | normal |
| 0 1110 100 | 74 | F4 | 192 | normal |
| 0 1110 101 | 75 | F5 | 208 | normal |
| 0 1110 110 | 76 | F6 | 224 | normal |
| 0 1110 111 | 77 | F7 | 240 | normal |
| 0 1111 000 | 78 | F8 | 256 | normal |
| 0 1111 001 | 79 | F9 | 288 | normal |
| 0 1111 010 | 7A | FA | 320 | normal |
| 0 1111 011 | 7B | FB | 352 | normal |
| 0 1111 100 | 7C | FC | 384 | normal |
| 0 1111 101 | 7D | FD | 416 | normal |
| 0 1111 110 | 7E | FE | 448 | normal |
| 0 1111 111 | 7F | FF | NaN | NaN |
For the 32 and 64 bit formats in full detail, see the IEEE 754 converter. For patterns as plain integers, the hex to binary converter and the integer limits pages.
Both are 16 bits, but they split them differently. FP16 (IEEE half precision) has 5 exponent bits and 10 mantissa bits, so it is more precise, about 3.3 decimal digits, but its largest value is only 65504. BF16 has 8 exponent bits and 7 mantissa bits, the same exponent as FP32, so it reaches about 3.3895e38 but keeps only about 2.4 digits. Training tends to need range more than digits, which is why BF16 is popular there.
65504, stored as 7BFF: the largest exponent, 2¹⁵, times 1.1111111111 in binary. Anything from 65520 up rounds to infinity, because 65520 is exactly halfway to the next step and ties go to the even pattern, which is infinity. That is why 70000 overflows in FP16 but not in BF16.
E4M3 spends 4 bits on the exponent and 3 on the mantissa: finer steps but a largest value of 448. E5M2 spends 5 on the exponent and 2 on the mantissa: coarser steps but a largest value of 57344. E5M2 keeps IEEE-style infinities and NaNs; E4M3 gives up infinity so the top exponent can hold more numbers, and only S.1111.111 is NaN. The paper that proposed the pair suggests E4M3 for weights and activations and E5M2 for gradients.
In E5M2, as in FP16 and BF16, it becomes infinity. E4M3 has no infinity, so the OCP specification allows two behaviours: saturate to ±448, or produce NaN. This converter saturates by default, so everything above 448 becomes 448, and lets you switch to NaN. In NaN mode the threshold is 464, halfway to the step above 448; 464 itself rounds down to 448 because ties go to even.
Because it was designed that way: the sign and the 8 exponent bits are the same as FP32, and the 7 mantissa bits are the top 7 of FP32's 23. Converting is dropping the low 16 bits, after rounding, so BF16 covers almost the same range as FP32 (its largest value is 3.3895e38 against 3.4028e38) and converts to and from it cheaply. The cost is precision: about 2 to 3 significant decimal digits.
FP4 E2M1 has 1 sign bit, 2 exponent bits and 1 mantissa bit, with no infinity and no NaN, so all 16 codes are numbers: ±0, 0.5, 1, 1.5, 2, 3, 4, 6. On its own that is very little, so it is used with block scaling: in MXFP4, every 32 values share one 8-bit power-of-two scale.