Double32s Benchmark Report
Generated by docs/Reports/benchmarks.jl on Julia 1.12.6.
Double32 is a two-word Float32 format carrying ~48 significant bits (~14 decimal digits) with the Float32 exponent range, designed for GPUs where Float64 is slow or absent. These are host CPU timings — the GPU execution layer is roadmap phase 2+ — so the interesting columns are the ratio. Float64 is the CPU-native alternative that GPUs generally lack, so D32/F64 answers "what would this cost me if I had Float64 and used it instead".
This report covers the whole implemented surface — scalar arithmetic, elementary functions, trigonometry, hyperbolics, reductions, and dense linear algebra. GPU timings are not here: KernelAbstractions and CUDA are weak dependencies that the root project cannot load, so sum(Double32, A), dot(Double32, x, y), norm2, and matmul are measured by docs/Reports/gpu_report.jl, which has its own environment.
Timings are best-of-trials amortized averages on the machine that ran the script — indicative magnitudes, not precise measurements.
Scalar Arithmetic
| operation | Double32 | Float64 | D32/F64 |
|---|---|---|---|
+ | 3 ns | 0 ns | 9.1x |
- | 3 ns | 0 ns | 9.1x |
* | 2 ns | 0 ns | 5.4x |
/ | 7 ns | 1 ns | 9.3x |
sqrt | 12 ns | 1 ns | 11.7x |
inv | 7 ns | 1 ns | 9.5x |
abs | 1 ns | 0 ns | 2.3x |
fma | 6 ns | 0 ns | 18.5x |
hypot | 45 ns | 2 ns | 18.1x |
< | 1 ns | 0 ns | 1.7x |
^ (Int) | 10 ns | 3 ns | 3.6x |
^ | 89 ns | 14 ns | 6.4x |
/, inv, and sqrt include the operand scaling that keeps their residuals representable at exponent extremes, and * includes its overflow/underflow band guard — none of these are optional fast paths (see the design's section 7.5 and the guard docstrings).
Elementary Functions
Roadmap Phase 6, rebuilt 2026-07-29 around table range reductions, dominant-add Horner steps, and mixed-precision (Float32) series tails — design section 20.6 records the rebuild and its measurements (2–7x per function). log and atanh are now division-free; atan pays at most one division. Measured accuracy for every function is in accuracy_report.md.
| operation | Double32 | Float64 | D32/F64 |
|---|---|---|---|
cbrt | 63 ns | 3 ns | 18.4x |
rsqrt | 26 ns | 2 ns | 15.0x |
exp | 29 ns | 2 ns | 17.0x |
expm1 | 38 ns | 2 ns | 16.5x |
exp2 | 25 ns | 2 ns | 15.0x |
exp10 | 32 ns | 2 ns | 18.9x |
log | 38 ns | 3 ns | 12.1x |
log1p | 40 ns | 3 ns | 12.3x |
log2 | 31 ns | 3 ns | 9.5x |
log10 | 35 ns | 3 ns | 11.1x |
Trigonometric Functions
sin, cos, tan, and sincos return NaN beyond Double32s.TRIG_ARG_LIMIT = 2^20; these timings are inside that range. sincos computes both from one argument reduction, so it is far cheaper than sin plus cos.
| operation | Double32 | Float64 | D32/F64 |
|---|---|---|---|
sin | 31 ns | 3 ns | 10.5x |
cos | 31 ns | 3 ns | 10.8x |
sincos | 35 ns | 3 ns | 10.7x |
tan | 62 ns | 4 ns | 15.5x |
asin | 93 ns | 2 ns | 43.8x |
acos | 105 ns | 2 ns | 52.7x |
atan | 59 ns | 3 ns | 17.6x |
atan2 | 72 ns | 6 ns | 11.2x |
Hyperbolic Functions
tanh is listed twice: inside its series band (|x| <= 0.35, no division) and across the general range (one division). The gap is the cost of the division this package's / cannot avoid.
| operation | Double32 | Float64 | D32/F64 |
|---|---|---|---|
sinh | 37 ns | 3 ns | 14.1x |
cosh | 35 ns | 2 ns | 15.3x |
tanh | 62 ns | 3 ns | 21.0x |
tanh (series band) | 21 ns | 2 ns | 12.4x |
asinh | 70 ns | 7 ns | 9.4x |
acosh | 85 ns | 6 ns | 13.9x |
atanh | 75 ns | 5 ns | 14.6x |
Reductions (10000-element vectors)
| operation | time | baseline | ratio |
|---|---|---|---|
sum(Double32, A::Vector{Float32}) | 70.60 µs | sum(Float64, A) 917 ns | 77.0x |
sum(::Vector{Double32}) | 69.91 µs | sum(::Vector{Float64}) 741 ns | 94.4x |
compensated_dot(::Vector{Float32}, ...) | 71.38 µs | dot (BLAS, Float32) 639 ns | 111.7x |
compensated_dot(::Vector{Double32}, ...) | 71.89 µs | dot (BLAS, Float64) 985 ns | 73.0x |
compensated_sum(::Vector{Double32}) | 69.90 µs | sum(::Vector{Double32}) 69.91 µs | 1.0x |
compensated_sum(::Vector{Float32}) | 70.32 µs | sum(Float64, A) 917 ns | 76.6x |
compensated_norm2(::Vector{Double32}) | 108.26 µs | norm (BLAS, Float64) 1.86 µs | 58.2x |
compensated_mapreduce(abs2, ::Vector{Double32}) | 70.66 µs | sum(abs2, ::Vector{Float64}) 982 ns | 66.5x |
sum(Double32, A::Vector{Float32}) is the package's primary use: a Float32 data stream accumulated at ~48-bit precision through Base's own mapreduce machinery, requiring no package function.
Linear Algebra (generic algorithms)
Float64 timings use LAPACK/BLAS. Double32 A*B, A*x, A'B, AB', and dot use the package's split-channel compensated kernels (arrays/matmul.jl); lu and \ still run Julia's generic dense fallbacks. The device-side tiled kernel is measured in gpu_report.md.
| operation | Double32 | Float64 | D32/F64 |
|---|---|---|---|
A * x (n=32) | 616 ns | 102 ns | 6.0x |
A' * B (n=32) | 15.90 µs | 1.97 µs | 8.1x |
A * B' (n=32) | 15.63 µs | 1.91 µs | 8.2x |
dot(x,y) (n=10^4) | 14.07 µs | 680 ns | 20.7x |
lu(A) (n=32) | 60.51 µs | 3.71 µs | 16.3x |
A \ b (n=32) | 58.98 µs | 4.19 µs | 14.1x |
Matrix multiply
| operation | Double32 | Float64 | D32/F64 |
|---|---|---|---|
A * B (n=4) | 305 ns | 96 ns | 3.2x |
A * B (n=8) | 1.02 µs | 128 ns | 8.0x |
A * B (n=16) | 4.30 µs | 416 ns | 10.3x |
A * B (n=32) | 15.42 µs | 1.40 µs | 11.0x |
A * B (n=64) | 85.49 µs | 9.09 µs | 9.4x |
The Double32 column is Julia's generic *; Float64 is BLAS. The device-side tiled kernel Double32s.matmul with exact Float32 products is measured in gpu_report.md.
Notes
- Entries below ~10 ns are at the resolution of the timing loop; treat them (and their ratios) as "too fast to matter" rather than exact.
Double32is an immutable 8-byte bitstype and vectors store it inline, so none of these timings include allocation.- Scalar
+,*,/measured here at a few ns/op correspond to the session-recorded baselines (≈7 / ≈5 / ≈15 ns on the development machine). - Both columns time the same sample values, mapped into each function's own domain so nothing is measured on inputs it would reject.