I tried to verify whether algebraic_* in Rust 1.98 speeds up audio DSP

I tried to verify whether algebraic_* in Rust 1.98 speeds up audio DSP

The performance of gain and mixing was almost the same as normal arithmetic. Dot product was 2.75 to 6.94 times faster, and 64-tap FIR was approximately 3.5 times faster. On the other hand, in exchange for the speedup, differences from normal arithmetic occurred.
2026.09.10

This page has been translated by machine translation. View original

Introduction

In audio DSP, floating-point operations are repeated many times per sample within a short time window. It can be difficult to know how well ordinary Rust code is utilizing a CPU's SIMD capabilities, leaving you unsure whether to reach for CPU-specific intrinsics.

Rust 1.98 added algebraic_* methods to f32 and f64. This API relaxes constraints such as the order of floating-point operations, allowing the compiler to optimize more aggressively. While it offers the possibility of faster multiply-accumulate operations in standard Rust, the results are not guaranteed to match those of ordinary operations.

With that in mind, I compared behavior across gain, mixing, division, dot product, and a 64-tap FIR. The results showed that gain and mixing were nearly identical to ordinary operations, while dot product achieved 2.751–6.936× speedup and FIR achieved 3.466–3.513×. In exchange for the speedup, however, differences from ordinary operations emerged.

What are algebraic_*?

The methods that became stable in Rust 1.98 are the following five:

  • algebraic_add
  • algebraic_sub
  • algebraic_mul
  • algebraic_div
  • algebraic_rem

With ordinary floating-point addition, a + b + c + d is evaluated in the order ((a + b) + c) + d. Since each operation involves rounding, changing the order of addition can change the result.

When chaining algebraic_add, the compiler is free to compute partial sums concurrently, such as (a + b) + (c + d). The official documentation permits optimizations such as reassociating and reordering operations, converting division to reciprocal multiplication, and optimizations that do not distinguish negative zero. The exact set of optimizations is unspecified, and results may differ even for the same input. (However, this does not constitute undefined behavior.)

Test Environment

  • MacBook Pro, Apple M4, 10-core, 16 GB memory
  • macOS 26.6.2, arm64
  • rustc 1.98.1, LLVM 22.1.8
  • Rust 2024 Edition, opt-level = 3
  • codegen-units = 1, target-cpu=apple-m4
  • No external crates

Target Audience

  • Those implementing numerical computation or audio DSP in Rust
  • Those unsure whether to write CPU-specific intrinsics for SIMD
  • Those wanting to understand the performance of algebraic_* and how its results differ from ordinary operations

References

Methodology

I prepared five types of processing with different characteristics:

  • gain: multiply each sample by the same coefficient
  • mixing: add corresponding samples from two signals
  • division: divide each sample by the same value
  • dot product: multiply elements of two arrays and accumulate into a single value
  • 64-tap FIR: perform 64 multiply-accumulate operations per output sample

The ordinary and algebraic_* versions were given identical loop structures, differing only in the operations themselves. For division, I also added an implementation that computes the reciprocal of the divisor once and then performs ordinary multiplication.

For 64, 128, 256, and 512 frames, I determined the number of iterations needed for a single trial to exceed 50 ms, then measured each condition 15 times while varying execution order. The reported values are medians per single function call.

For numerical differences, I used finite values in [-1, 1] generated from a fixed seed, and compared results computed as the same multiply-accumulate with ordinary operations, algebraic_*, and f64. For dot product, I ran 10,000 samples per size; for FIR, 1,000 blocks per size.

Implementation

The essential part of dot product, where the difference is most apparent, is as follows:

fn dot_normal(left: &[f32], right: &[f32]) -> f32 {
    left.iter()
        .zip(right)
        .fold(0.0, |sum, (&left_sample, &right_sample)| {
            sum + left_sample * right_sample
        })
}

fn dot_algebraic(left: &[f32], right: &[f32]) -> f32 {
    left.iter()
        .zip(right)
        .fold(0.0, |sum, (&left_sample, &right_sample)| {
            sum.algebraic_add(left_sample.algebraic_mul(right_sample))
        })
}

The FIR also zips the input and coefficients and accumulates 64 elements using a fold of the same shape. The input range is sliced before the tap loop to avoid leaving bounds checks inside the inner loop.

Results

The speed ratio is normal / algebraic_*. A value greater than 1 indicates that algebraic_* is faster.

Operation 64 frames 128 frames 256 frames 512 frames
gain 1.007× 1.003× 0.997× 0.992×
mixing 1.002× 0.999× 1.003× 0.980×
division 0.981× 1.012× 1.047× 1.361×
dot product 2.751× 3.945× 5.408× 6.936×
64-tap FIR 3.466× 3.486× 3.507× 3.513×

For gain and mixing, the ratio was 0.980–1.007×, showing virtually no performance difference. Since neither has dependencies between elements, both the ordinary and algebraic_* versions compiled to the same kind of NEON instructions. These are operations that can be vectorized with ordinary operations as well.

For dot product, the speed ratio grew larger as the element count increased; at 512 elements, the ordinary version took 174.995 ns compared to 25.231 ns for the algebraic_* version. The 64-tap FIR was approximately 3.5× regardless of frame count, going from 5.024 µs to 1.430 µs at 512 frames.

For division, there was little difference at 64 elements, but at 512 elements the ratio reached 1.361×. Examining the assembly, the ordinary version performed vector division inside the loop, while the algebraic_div version computed the reciprocal first and then performed vector multiplication. The manual reciprocal-multiplication version also showed a ratio of 1.398× at 512 elements, exhibiting the same trend.

The dot product assembly showed the following difference. The ordinary version multiplies 4 elements together, then extracts each element to a scalar register and adds them sequentially.

fmul.4s  v1, v1, v5
mov      s18, v1[1]
fadd     s0, s0, s1
fadd     s0, s0, s18

The algebraic_* version updates 4 vector partial sums in parallel using fmla.4s, then performs a horizontal add at the end.

fmla.4s  v0, v16, v4
fmla.4s  v1, v17, v5
fmla.4s  v2, v18, v6
fmla.4s  v3, v19, v7
fadd.4s  v0, v1, v0
fadd.4s  v1, v3, v2
fadd.4s  v0, v1, v0
faddp.4s v0, v0, v0
faddp.2s s0, v0

In the LLVM IR, the algebraic_* side had reassoc nsz arcp contract attached. nnan and ninf were not present, so this is not an implementation that enables all fast-math flags.

In exchange for performance, numerical differences from ordinary operations emerged. The dot product results are as follows:

Elements Max absolute diff from normal Max absolute error of normal Max absolute error of algebraic_*
64 2.861023e-6 2.636478e-6 1.015914e-6
128 4.768372e-6 4.905520e-6 1.293832e-6
256 9.536743e-6 1.030360e-5 2.252584e-6
512 2.098083e-5 2.178359e-5 3.726489e-6

"Max absolute error" is the difference from a reference value computed by widening the input to f64. With the inputs used here, the algebraic_* version was closer to the reference value.

The maximum absolute difference from normal operations for the 64-tap FIR was 5.215406e-8–6.705523e-8 across all sizes. For 10,000 division samples, the difference between ordinary division and algebraic_div was at most 1 ULP, or 5.960464e-8 in absolute value.

Differences due to operation order and sign of zero were also observable.

Input and operation Normal algebraic_*
Add [1e10, -1e10, 1, 0] repeated 4 times 1.0 4.0
Add [MAX, MAX, -MAX, -MAX] Infinity NaN
Add 0.0 to -0.0 0.0 -0.0

Even with only finite values, reassociation caused a split between 1.0 and 4.0. For an addition where intermediate results overflow, ordinary operations and algebraic_* diverged into Infinity and NaN. Negative zero is equal to positive zero in numerical comparison, but differs in bit representation.

Discussion

For operations that process each element independently, like gain and mixing, I think it is sufficient to verify whether the ordinary loop is auto-vectorized rather than immediately switching to algebraic_*. If the same kinds of instructions are generated as in this case, there is no need to accept differences in computed results just to make the switch.

On the other hand, in dot product and FIR, because each addition uses the result of the previous one, ordinary operations must proceed sequentially. For operations where reordering is acceptable, computing multiple partial sums in parallel may be worth trying. However, bounds checks and loop structure can also impede vectorization. It is necessary to examine the instructions emitted by the compiler along with processing time for representative block sizes.

For division by a common value, converting to reciprocal multiplication reduced processing time for larger element counts. If the code can express that the divisor does not change within the loop, manually computing the reciprocal once is also worth considering as an alternative. In either case, it will likely be necessary to decide whether returning different results from per-element division versus multiplication by a pre-rounded reciprocal is acceptable.

For real-time audio processing, it would be best to identify multiply-accumulate operations where differences in operation order and rounding results are acceptable, then verify callback processing time and output signal on the target CPU before adopting them.

Summary

Using algebraic_* from Rust 1.98, I compared five types of processing modeled on audio DSP scenarios. Gain and mixing, which could already be vectorized with ordinary operations, showed nearly identical performance. Meanwhile, dot product achieved 2.751–6.936× and the 64-tap FIR achieved 3.466–3.513×.

algebraic_* is not an API that speeds up all floating-point processing. It seems best to narrow the scope to multiply-accumulate operations where differences in operation order and rounding results are acceptable, and to verify results for representative inputs, compiler-emitted instructions, processing time, and error all together.

Share this article