Basic Uses
Everything on this page runs on a plain CPU with nothing installed beyond Double32s itself. Every output shown is checked by the documentation build.
Fourteen digits from Float32 words
A Double32 is the unevaluated sum of two Float32 words. Arithmetic keeps about 48 significant bits — ~14 decimal digits — where Float32 keeps ~7:
julia> using Double32s
julia> x = Double32(1.0f0) / 3
0.333333333333333
julia> Float64(x) # ~14 correct digits
0.33333333333333304
julia> Float64(1.0f0 / 3.0f0) # Float32 alone: ~7
0.3333333432674408The value displays as a decimal; Double32s.HILO shows the two words behind it, the high word being the nearest Float32 and the low word carrying the remainder.
julia> Float64(sqrt(Double32(2.0f0)))
1.4142135623730958
julia> sqrt(2.0) # Float64 reference
1.4142135623730951Accumulating Float32 data without loss
Summing Float32 data in Float32 loses digits to accumulated rounding. Summing it through Double32 needs no package-specific API — Base already spells it — and returns the sum of the stored values essentially exactly:
julia> A = fill(0.1f0, 10_000);
julia> sum(A) # Float32 accumulation: 4 digits survive
999.9998f0
julia> Float64(sum(Double32, A))
1000.0000149011612That …0149 is not an error — 0.1f0 is not exactly 1/10, and 1000.0000149011612 is the true sum of ten thousand copies of the value actually stored. Double32 reports the data as it is; Float32 accumulation buries it in rounding noise.
The same holds for ill-behaved series:
julia> v = Float32.(1 ./ (1:100_000));
julia> sum(v) # Float32: 6 digits
12.090148f0
julia> Float64(sum(Double32, v))
12.09014619539721The reference value to 17 digits is 12.090146195397210: every printed digit of the Double32 sum is correct.
Elementary functions
The whole elementary family — exp, log, trigonometric, hyperbolic, and their inverses — is implemented directly on the pair format and measured against a 400-bit reference (see the accuracy report):
julia> Float64(sin(Double32(1.0f0)))
0.8414709848078967
julia> sin(1.0) # Float64 reference
0.8414709848078965
julia> Float64(exp(Double32(1.0f0)))
2.718281828459041
julia> Float64(4 * atan(one(Double32)))
3.1415926535897967expm1 and log1p exist for the same reason they exist for Float64: near zero, exp(x) - 1 and log(1 + x) destroy exactly the digits being asked for.
julia> x = Double32(1.0f-7);
julia> Float64(expm1(x)) # correct: 1e-7 + 5e-15 - ...
1.0000000616860965e-7
julia> expm1(x) === exp(x) - one(Double32) # the naive spelling differs
falseInteger powers and exact results
Operations whose results are exactly representable remain exact, and stay within the type:
julia> Double32(2.0f0)^10
1024.0
julia> log2(Double32(1024.0f0))
10.0
julia> Double32(1.0f0) + 0.5f0 # Float32 promotes to Double32
1.5Double32 widens the significand, not the exponent: overflow happens at Float32's ~3.4e38, and exp saturates accordingly:
julia> exp(Double32(1000.0f0)), exp(Double32(-1000.0f0))
(Inf, 0.0)