Double32s Benchmark Report

Generated by docs/Reports/benchmarks.jl on Julia 1.12.6.

Double32 is a two-word Float32 format carrying ~48 significant bits (~14 decimal digits) with the Float32 exponent range, designed for GPUs where Float64 is slow or absent. These are host CPU timings — the GPU execution layer is roadmap phase 2+ — so the interesting columns are the ratio. Float64 is the CPU-native alternative that GPUs generally lack, so D32/F64 answers "what would this cost me if I had Float64 and used it instead".

This report covers the whole implemented surface — scalar arithmetic, elementary functions, trigonometry, hyperbolics, reductions, and dense linear algebra. GPU timings are not here: KernelAbstractions and CUDA are weak dependencies that the root project cannot load, so sum(Double32, A), dot(Double32, x, y), norm2, and matmul are measured by docs/Reports/gpu_report.jl, which has its own environment.

Timings are best-of-trials amortized averages on the machine that ran the script — indicative magnitudes, not precise measurements.

Scalar Arithmetic

operationDouble32Float64D32/F64
+3 ns0 ns9.1x
-3 ns0 ns9.1x
*2 ns0 ns5.4x
/7 ns1 ns9.3x
sqrt12 ns1 ns11.7x
inv7 ns1 ns9.5x
abs1 ns0 ns2.3x
fma6 ns0 ns18.5x
hypot45 ns2 ns18.1x
<1 ns0 ns1.7x
^ (Int)10 ns3 ns3.6x
^89 ns14 ns6.4x

/, inv, and sqrt include the operand scaling that keeps their residuals representable at exponent extremes, and * includes its overflow/underflow band guard — none of these are optional fast paths (see the design's section 7.5 and the guard docstrings).

Elementary Functions

Roadmap Phase 6, rebuilt 2026-07-29 around table range reductions, dominant-add Horner steps, and mixed-precision (Float32) series tails — design section 20.6 records the rebuild and its measurements (2–7x per function). log and atanh are now division-free; atan pays at most one division. Measured accuracy for every function is in accuracy_report.md.

operationDouble32Float64D32/F64
cbrt63 ns3 ns18.4x
rsqrt26 ns2 ns15.0x
exp29 ns2 ns17.0x
expm138 ns2 ns16.5x
exp225 ns2 ns15.0x
exp1032 ns2 ns18.9x
log38 ns3 ns12.1x
log1p40 ns3 ns12.3x
log231 ns3 ns9.5x
log1035 ns3 ns11.1x

Trigonometric Functions

sin, cos, tan, and sincos return NaN beyond Double32s.TRIG_ARG_LIMIT = 2^20; these timings are inside that range. sincos computes both from one argument reduction, so it is far cheaper than sin plus cos.

operationDouble32Float64D32/F64
sin31 ns3 ns10.5x
cos31 ns3 ns10.8x
sincos35 ns3 ns10.7x
tan62 ns4 ns15.5x
asin93 ns2 ns43.8x
acos105 ns2 ns52.7x
atan59 ns3 ns17.6x
atan272 ns6 ns11.2x

Hyperbolic Functions

tanh is listed twice: inside its series band (|x| <= 0.35, no division) and across the general range (one division). The gap is the cost of the division this package's / cannot avoid.

operationDouble32Float64D32/F64
sinh37 ns3 ns14.1x
cosh35 ns2 ns15.3x
tanh62 ns3 ns21.0x
tanh (series band)21 ns2 ns12.4x
asinh70 ns7 ns9.4x
acosh85 ns6 ns13.9x
atanh75 ns5 ns14.6x

Reductions (10000-element vectors)

operationtimebaselineratio
sum(Double32, A::Vector{Float32})70.60 µssum(Float64, A) 917 ns77.0x
sum(::Vector{Double32})69.91 µssum(::Vector{Float64}) 741 ns94.4x
compensated_dot(::Vector{Float32}, ...)71.38 µsdot (BLAS, Float32) 639 ns111.7x
compensated_dot(::Vector{Double32}, ...)71.89 µsdot (BLAS, Float64) 985 ns73.0x
compensated_sum(::Vector{Double32})69.90 µssum(::Vector{Double32}) 69.91 µs1.0x
compensated_sum(::Vector{Float32})70.32 µssum(Float64, A) 917 ns76.6x
compensated_norm2(::Vector{Double32})108.26 µsnorm (BLAS, Float64) 1.86 µs58.2x
compensated_mapreduce(abs2, ::Vector{Double32})70.66 µssum(abs2, ::Vector{Float64}) 982 ns66.5x

sum(Double32, A::Vector{Float32}) is the package's primary use: a Float32 data stream accumulated at ~48-bit precision through Base's own mapreduce machinery, requiring no package function.

Linear Algebra (generic algorithms)

Float64 timings use LAPACK/BLAS. Double32 A*B, A*x, A'B, AB', and dot use the package's split-channel compensated kernels (arrays/matmul.jl); lu and \ still run Julia's generic dense fallbacks. The device-side tiled kernel is measured in gpu_report.md.

operationDouble32Float64D32/F64
A * x (n=32)616 ns102 ns6.0x
A' * B (n=32)15.90 µs1.97 µs8.1x
A * B' (n=32)15.63 µs1.91 µs8.2x
dot(x,y) (n=10^4)14.07 µs680 ns20.7x
lu(A) (n=32)60.51 µs3.71 µs16.3x
A \ b (n=32)58.98 µs4.19 µs14.1x

Matrix multiply

operationDouble32Float64D32/F64
A * B (n=4)305 ns96 ns3.2x
A * B (n=8)1.02 µs128 ns8.0x
A * B (n=16)4.30 µs416 ns10.3x
A * B (n=32)15.42 µs1.40 µs11.0x
A * B (n=64)85.49 µs9.09 µs9.4x

The Double32 column is Julia's generic *; Float64 is BLAS. The device-side tiled kernel Double32s.matmul with exact Float32 products is measured in gpu_report.md.

Notes

  • Entries below ~10 ns are at the resolution of the timing loop; treat them (and their ratios) as "too fast to matter" rather than exact.
  • Double32 is an immutable 8-byte bitstype and vectors store it inline, so none of these timings include allocation.
  • Scalar +, *, / measured here at a few ns/op correspond to the session-recorded baselines (≈7 / ≈5 / ≈15 ns on the development machine).
  • Both columns time the same sample values, mapped into each function's own domain so nothing is measured on inputs it would reject.