This file records the benchmark results for the RVV (RISC-V Vector 1.0) port of SBCL, measured on a riscv64 host on two different sets of cores:

  • default cores — VLEN=256 (VLMAX for e32 = 8 elements)

  • AI-accelerated cores — VLEN=1024 (VLMAX for e32 = 32 elements), reached with the $HOME/bin/ai launcher (which execs $HOME/bin/aix).

The Lisp benchmark is rvv-bench.lisp; the C calibration is cbench.c. Both assert or demonstrate correctness before timing, so every number below is from a verified run. Each Lisp cell is the best of three GC-quieted runs, reported in millions of elements per second (Melem/s, higher is better).

Environment

  • SBCL: built from this repository with :riscv-vector support (features contains :riscv-vector), and with the opt-in --with-riscv-zba flag (features contains :riscv-zba; see INSTALL).

  • gcc: (Bianbu 15.2.0-16ubuntu1bb5) 15.2.0

  • binutils: GNU Binutils for Bianbu 2.46

  • clang: Bianbu clang 21.1.8

  • The M1/chunk/blk/blk4 paths use SEW=e32/e64 and LMUL=1; the other benchmarks also exercise LMUL=2/4/8, whole-array reductions, mask compare/merge, fused multiply-add, same-width float <→ integer conversions, integer/float width changes (spec 07), and (spec 08) strided/indexed gather.

How to reproduce

# Lisp benchmark — default cores (VLEN=256), no aix:
./run-sbcl.sh --noinform --disable-debugger \
    --script rvv-results/rvv-bench.lisp

# Lisp benchmark — AI cores (VLEN=1024), with aix:
PATH="$HOME/bin:$PATH" \
  $HOME/bin/ai ./run-sbcl.sh --noinform --disable-debugger \
    --script rvv-results/rvv-bench.lisp

# C calibration — default cores:
gcc -O3 -march=rv64gcv -fno-tree-vectorize rvv-results/cbench.c -o cbench
./cbench

# C calibration — AI cores:
PATH="$HOME/bin:$PATH" $HOME/bin/ai ./cbench

What is measured

  • scalar — a plain Lisp loop over (simple-array single-float (*)) (or (unsigned-byte 32), double-float) compiled at (speed 3) (safety 0).

  • chunk — the %RVV-*-CHUNK VOPs called from a Lisp loop; each call does one vsetvli-bounded load/compute/store and returns how many elements it processed.

  • blk — %RVV-F32-ADD-BLOCK, the whole array handled by an assembly-internal strip-mine loop inside a single VOP.

  • blk4 — %RVV-F32-ADD-BLOCK4, the same but 4-way unrolled over all eight caller-saved vector registers (v24-v31).

  • chunk/blk m2, m4, m8 — the LMUL=2/4/8 register-group variants of the chunk and block loops; one instruction covers 2x/4x/8x the elements (chunk m8 uses the full v16-v31 scratch window).

  • blk8 — %RVV-F32-ADD-BLOCK8, an 8-way unrolled block loop (v16-v31).

  • reduce — %RVV-F32/F64-REDUCE-ADD and %RVV-U32-REDUCE-MAX, whole-array scalar-out reductions (a single VOP strip-mines the array and folds one vector into an accumulator).

  • mask — %RVV-U32-MAX-MASKED (compare + merge) and %RVV-U32-CMP-COUNT (compare + vcpop.m).

  • fma — %RVV-F32-FMA-CHUNK / %RVV-U32-MACC-CHUNK (dst = src1*src2 + src3).

  • fma overwrite — %RVV-F32-FMADD-CHUNK (vfmadd.vv), the overwrite-multiplicand form of the same dst = a*b + c contract.

  • convert — %RVV-U32→F32-CHUNK / %RVV-F32→U32-CHUNK (same-width float <→ integer conversion, SEW=e32).

  • width — %RVV-U8→U32-CHUNK (vzext.vf4), %RVV-U16→U32-CHUNK (vzext.vf2), %RVV-U32→U16-CHUNK (vnsrl.wi), %RVV-F32→F64-CHUNK (vfwcvt.f.f.v) and %RVV-F64→F32-CHUNK (vfncvt.f.f.w) — integer and float width changes (spec 07).

  • widen arith — %RVV-S32/U32-WIDEN-ADD-CHUNK (vwadd.vv/ vwaddu.vv) and %RVV-S32/U32-WIDEN-MUL-CHUNK (vwmul.vv/ vwmulu.vv), widening arithmetic from e32 to e64 (spec 07).

  • widen fma — %RVV-F32-WIDEN-FMA-CHUNK (vfwmacc.vv) and %RVV-U32-WIDEN-MACC-CHUNK (vwmaccu.vv), a widening multiply folded into a 64-bit accumulator (spec 06/07).

  • gather — %RVV-F32-STRIDED-LOAD (vlse32.v) and %RVV-F32-INDEXED-LOAD (vluxei32.v) gather a strided column or an indirectly-indexed set of elements into a compact array (spec 08).

  • segment — %RVV-U32-SEGMENT-LOAD-2…-8 (vlseg2e32.v… vlseg8e32.v) deinterleave an AoS u32 stream into a field-major SoA destination (spec 08).

Summary

  1. Net win — element-wise compute. RVV gives 2.1–6.3x over scalar Lisp for cache-resident f32/u32 (up to ~7.9x for f32 add with LMUL=4), and mask compare/merge, FMA, float↔int conversion, integer/float width change, and strided/indexed gather add another 1.2–6.8x. LMUL=2/4 closes the gap to gcc’s auto-vectorizer; every family is correctness-asserted before timing.

  2. Neutral to loss — memory- and latency-bound paths. Ordered float reductions sit at ~1x (order preservation forbids the parallel tree reduction), DRAM-bound f64 add and strided gather sit at ~0.6–1.0x, segment deinterleave at small nf is ~0.5–1.0x (unamortized vsetvli+vlseg+store overhead), and blk8 unrolling is neutral-to-loss.

  3. Caveat — weak scalar baseline. Lisp scalar is ~2.0–2.5x slower than C (per-element address recomputation, no indexed addressing, no loop unrolling); --with-riscv-zba recovers ~1.23x (gap → ~2.0x). So the RVV speedups are measured against a conservative scalar loop, not against C.

Findings

  1. RVV gives a real 2.1–6.3x element-wise speedup over scalar Lisp for cache-resident f32/u32 arrays, and ~1.0–1.4x for f64.

  2. LMUL=2/4 closes most of the previously identified gap: at VLEN=256 chunk m4 reaches ~3.9 Gelem/s at n=4093, in the same range as gcc’s auto-vectorizer (~3.4 Gelem/s), and the f32 add speedup goes from ~2.3x (LMUL=1) to ~7.9x (LMUL=4) in cache. At VLEN=1024 the LMUL win is larger still (~10–13x, see the aix tables below). LMUL=8 adds no further throughput at either VLEN: the memory-bound chunk is already bandwidth-saturated at M4, so the wider groups only add register pressure (see the LMUL key points).

  3. The bandwidth ceiling is the limiter only on the default cores: at n=1M (VLEN=256) all LMUL variants converge to ~0.72–0.76 Gelem/s and ~1.5–1.7x, matching the C DRAM-bound result (~1.1x); the AI cores' higher bandwidth keeps LMUL=4 at ~2 Gelem/s even at n=1M.

  4. Scalar-out reductions are a mixed bag: associative reductions with a branch (u32 max) win ~3.6–6.3x on the default cores (more on the slower AI-core scalar baseline), but the ordered float sum (vfredosum) is neutral against a tight scalar loop on the default cores (~1x) and only ~2x on the AI cores.

  5. Mask compare/merge (~2.1–4.1x), fused multiply-add (~2.4–3.8x), and float <→ integer conversion (~1.7–2.3x for u32→f32, ~4.5–6.8x for f32→u32) all deliver useful element-wise wins with no correctness issues.

  6. Integer/float width change (spec 07) is another solid win. The fractional-LMUL refinement (spec 07’s deferred half) raises LMUL for the widening VOPs so the narrow source fills a whole register: u8 → u32 widening goes from ~2.5x to ~6.0–6.5x, u16→u32 from ~2.3x to ~4.4–4.6x. u32→u16 narrowing stays ~4–6x and f32→f64 / f64→f32 width change ~1.3–2.3x, with no regressions.

  7. Unrolling does not scale past the memory-level parallelism the hardware can exploit: blk8 is neutral at VLEN=256 and a loss at VLEN=1024, so blk4/LMUL scaling is the better lever than wider unrolling for f32 add.

  8. Strided/indexed gather (spec 08) is a modest ~1.0–1.2x in cache: these ops are memory-bound (random-ish access) and their main value is enabling vectorized non-contiguous access rather than raw throughput.

  9. Unit-stride segment loads (spec 08’s last piece) deinterleave an AoS u32 stream into a field-major SoA destination in one vlseg<nf>e32.v plus nf stores. The win is small (~1.2–1.4x at nf=6..8, and a loss at nf=2) because both memory streams are already sequential; the value is the single-instruction AoS→SoA deinterleave rather than raw bandwidth.

  10. The C scalar column is ~2.4–2.5x faster than the Lisp scalar column for f32 add (994 vs 396 Melem/s at n=4093) purely from scalar code generation, not memory bandwidth. Lisp scalar is flat at ~396 Melem/s from n=251 (L1) to n=1M (DRAM) while C drops from ~960 to ~691, i.e. the Lisp loop is instruction/latency-bound. The cause is per-element array-address recomputation (RISC-V has no scaled/indexed addressing, so each aref costs slli+add+load) plus the absence of loop unrolling in SBCL. Manually unrolling the Lisp loop 4–8x lifts it to ~565–585 Melem/s. This is a scalar-baseline artifact: the RVV VOPs already match/beat gcc’s auto-vectorizer. See slow-scalar.adoc.

  11. Building with --with-riscv-zba replaces the per-aref slli+add address computation with the fused sh1add/sh2add/sh3add instructions, lifting the scalar f32 add baseline from ~396 to ~486–496 Melem/s (~1.23–1.28x). This narrows the gap to C’s ~994 Melem/s from ~2.5x to ~2.0x. The RVV Melem/s numbers are unchanged, so the scalar-vs-RVV speedup ratios are correspondingly lower on a Zba build (faster scalar, same vector throughput). Every table in this file was re-measured on the --with-riscv-zba build, so all scalar columns below reflect the Zba fused shift+add addressing.

  12. Widening arithmetic (spec 07’s last piece) and widening FMA (spec 06’s last piece) both deliver solid element-wise wins. vwadd.vv/ vwaddu.vv are ~2.6–3.0x and vwmul.vv/vwmulu.vv ~2.1–2.6x over scalar in cache, and vfwmacc.vv/vwmaccu.vv ~2.0–2.6x; all collapse to ~1.1–1.8x at n=1M where the 8-byte destination makes the loop DRAM-bound. The overwrite-multiplicand vfmadd.vv is within noise of the accumulate vfmacc.vv (~2.8x), as expected for the same instruction shape.

  13. A handful of VOPs are not faster than scalar, and that is expected rather than a bug. Three distinct causes:

    • Ordered reductions — the f32/f64 sum VOPs use vfredosum.vs, the ordered sum, to match Lisp reduce #'+ element order; ordering forbids the parallel tree reduction, so they sit at ~0.9–1.0x against a tight scalar loop.

    • Memory-bound, no compute to hide — f64 add at n=1M (~1.0x) and strided/indexed gather on the AI cores (~0.6–0.7x): both paths hit the same bandwidth ceiling, and the vector path’s extra instructions (plus, for gather, a defeated prefetcher) add cost without recovering it.

    • Unamortized per-call overhead — segment deinterleave at small nf/n (~0.5–1.0x): the vsetvli+vlseg+store round-trip costs more than the scalar two-load/store loop it replaces, and only wins once nf/n grow enough to amortize it. These last two families exist for correctness and expressiveness (single-instruction AoS→SoA deinterleave, vectorized non-contiguous access), not raw throughput.

Results without aix — default cores, VLEN=256 (VLMAX e32 = 8)

Table 1. SBCL scalar vs RVV (Melem/s, higher is better)
n scalar chunk blk blk4 speedup (blk4)

251

434

897

906

1498

3.5x

4093

495

1130

1522

3117

6.3x

65536

496

1254

1579

2159

4.4x

1048576

487

776

753

738

1.5x

Key points:

  • RVV is ~2.1–3.5x faster than scalar for the small/medium sizes where the arrays stay in cache, and ~6.3x at n=4093 for the 4-way-unrolled block loop.

  • blk and blk4 beat chunk substantially at n=4093 and n=65536, showing that the per-call Lisp loop overhead of the chunk API is real: moving the strip-mine loop into assembly recovers it.

  • At n=1048576 everything is DRAM-bound: vector throughput drops to ~0.74 Gelem/s and the advantage collapses to ~1.5x.

C calibration without aix (gcc -O3 -march=rv64gcv; scalar loop

compiled -fno-tree-vectorize, vector loop with tree-vectorize)

n C scalar C vector C speedup

251

960

3047

3.2x

4093

994

3384

3.4x

65536

752

2452

3.3x

1048576

691

759

1.1x

C shows the same shape (vector wins in cache, DRAM-bound collapse at n=1M), which confirms the Lisp numbers are measuring the memory hierarchy, not an SBCL artifact.

Why C scalar beats Lisp scalar

Note
this section analyzes the pre-Zba scalar baseline (a flat ~396 Melem/s); the following section shows how Zba narrows the gap.

The C scalar column (994 Melem/s at n=4093) is ~2.5x the Lisp scalar column (396). The gap is code generation, not memory bandwidth:

  • Lisp scalar is flat at ~396 Melem/s from n=251 (L1-resident) to n=1M (DRAM), whereas C scalar drops from ~960 to ~691. A memory-bound loop tracks the cache like C does; a flat line means the Lisp loop is instruction/latency-bound.

  • SBCL recomputes each array element’s byte address from the tagged fixnum index on every element (slli to untag/scale, then add to the base pointer), three times per element for the three arrays. RISC-V has no scaled/indexed addressing mode, so this cannot be folded into the load the way x86-64 (movss [base+idx*4+disp]) and arm64 (add with folded lsl) can. gcc strength-reduces the loop to three pointer-induction variables; SBCL does not.

  • SBCL does not auto-unroll scalar loops. Manual unrolling closes most of the gap:

variant (n=4093) Melem/s

base, 1 element/iter

396

unroll 4x

565

unroll 8x

585

  • The benchmark’s of-type index (SBCL’s index is (integer 0 array-dimension-limit), which this build reports as not a subtype of fixnum) also makes the loop-index arithmetic go through generic TWO-ARG-+/TWO-ARG→= calls. Switching the loop variable to fixnum produces a much tighter loop but does not move the measured number on this benchmark (the loop is already latency-bound), so it is a codegen smell rather than the bottleneck.

The scalar-vs-RVV speedups in this file are therefore measured against a conservative scalar baseline. The RVV VOPs already reach parity with gcc’s auto-vectorizer (chunk m4 ~3.7 vs ~3.4 Gelem/s at n=4093), so the RVV conclusions are unaffected. A fuller treatment, including what could change in the SBCL compiler, is in slow-scalar.adoc.

Zba (sh1add/sh2add/sh3add) scalar baseline

Building with --with-riscv-zba (opt-in, see INSTALL) lets the array VOPs use the RISC-V Zba fused shift+add instructions for scaled element addressing instead of slli+add. For a single-float aref the byte address becomes sh1add (the << 1 also untags the fixnum index), for double-float sh2add, and for (unsigned-byte 32)/complex sh2add/ sh3add, removing one instruction per array reference.

Table 2. Zba scalar f32 add (Melem/s, higher is better)
n scalar (no Zba) scalar (Zba) speedup

251

358

459.5

1.28x

4093

396

496.1

1.25x

65536

399

495.0

1.24x

1048576

396

485.8

1.23x

  • Zba removes one slli per aref: the scalar f32 add loop goes from a flat ~396 Melem/s to ~486–496 Melem/s (~1.23–1.28x), and the inner loop now disassembles to three sh1add and no slli.

  • This narrows but does not close the gap to C (~994 Melem/s at n=4093): the remaining ~2x is loop unrolling and pointer-induction strength reduction, compiler issues independent of Zba (see slow-scalar.adoc).

  • The scalar-vs-RVV tables in this file have all been re-measured on the Zba build, so their scalar columns already reflect the Zba speedup; the speedup ratios are correspondingly lower than on a non-Zba build (faster scalar, unchanged RVV throughput).

LMUL=1/2/4 chunk and block (f32 add, SEW=e32)

Table 3. LMUL scaling, f32 add chunk VOPs (Melem/s)
n scalar chunk m1 chunk m2 chunk m4 chunk m8

251

436

894

1433

1394

1238

4093

494

1132

2215

3883

3047

65536

497

1216

2335

2389

2310

1048576

457

738

727

786

730

Table 4. LMUL scaling and unroll, f32 add block VOPs (Melem/s)
n scalar blk m1 blk m2 blk m4 blk8

251

439

909

1227

1354

1170

4093

494

1528

2749

3944

3058

65536

495

1582

2353

2426

2186

1048576

492

751

762

765

758

Key points:

  • LMUL=2/4 gives a large, real speedup in cache: at n=4093 the chunk path goes from 2.3x (m1) to 4.5x (m2) to 7.9x (m4), and the block path reaches 8.4x (m4). This is exactly the instruction-count reduction predicted: one LMUL=4 vector instruction covers VLEN*4/32 = 32 elements.

  • At n=1048576 the memory system is the limit: m1/m2/m4/m8 all converge to ~1.5–1.6x and ~0.72–0.79 Gelem/s.

  • LMUL=8 is neutral to slightly worse than M4 at VLEN=256 (6.2x vs 7.9x at n=4093): with VLMAX=8 an M8 group already covers 64 elements per instruction, but the two 8-register groups (16 registers total) add register pressure without extra memory parallelism. See the VLEN=1024 tables below, where M8 covers 256 elements and is expected to help more.

  • blk8 (8-way unroll) is not a win over blk m4 at VLEN=256: at n=4093 blk8 is 6.2x vs blk m4’s ~8.0x, and at n=65536 4.5x vs 4.9x. With VLMAX=8 a 4-way unroll already keeps 32 elements in flight, so 8-way adds register pressure without more memory parallelism.

  • chunk m4 at n=4093 reaches ~3.9 Gelem/s, in the same neighborhood as the C auto-vectorizer (3.4 Gelem/s at n=4093).

Whole-array reductions (scalar out)

Table 5. Reduction VOPs (Melem/s; scalar vs RVV, whole array per call)
n f32 sum scalar f32 sum rvv u32 max scalar u32 max rvv u32 max speedup

251

575

545

437

1579

3.6x

4093

659

657

491

2996

6.1x

65536

664

664

498

3158

6.3x

1048576

662

663

497

2086

4.2x

Key points:

  • u32 max (vredmaxu.vs) is a clear win: ~3.6–6.3x over the scalar loop, because the integer max has a loop-carried branch that the scalar path cannot hide, while the vector reduction is parallel.

  • f32/f64 sum (vfredosum.vs) does not beat a tight scalar loop (~1.0x and ~0.9x). vfredosum is the ordered sum variant; it preserves element order (required to match Lisp reduce #'+) and therefore does not gain the tree-reduction parallelism of the unordered vfredusum. This is a documented, honest result rather than a bug.

Mask compare/merge (u32)

Table 6. Mask VOPs (Melem/s; scalar vs RVV)
n scalar max rvv max masked scalar cmp-count rvv cmp-count

251

428

917

429

760

4093

456

1284

467

1013

65536

320

1322

384

1037

1048576

268

743

350

855

Key points:

  • u32 max masked (vmsgtu.vv → vmerge.vvm) is ~2.1–4.1x over scalar max, comparable to (slightly below) the single-instruction vmaxu.vv chunk, as expected for a two-instruction sequence.

  • u32 cmp-count (vmsgtu.vv → vcpop.m) is ~1.8–2.7x over a scalar count loop.

Fused multiply-add (dst = a*b + c)

Table 7. FMA VOPs (Melem/s; scalar vs RVV)
n scalar f32 fma rvv f32 fma scalar u32 macc rvv u32 macc

251

363

864

323

764

4093

396

1141

337

1131

65536

367

1062

281

1067

1048576

353

473

279

476

Key points:

  • FMA/MACC fuses two operations into one instruction: ~2.4–2.9x for f32, ~2.4–3.8x for u32 in cache, collapsing to ~1.3–1.7x at n=1M where the arrays stop fitting in cache.

  • The overwrite-multiplicand form (vfmadd.vv, %RVV-F32-FMADD-CHUNK) performs within noise of the accumulate form (vfmacc.vv): ~2.8x at n=4093 and n=65536, ~2.2x at n=251, and ~1.3x at n=1M — the same instruction shape, so no extra cost.

Same-width float <→ integer conversions (SEW=e32)

Table 8. Conversion VOPs (Melem/s; scalar vs RVV)
n scalar u32→f32 rvv u32→f32 scalar f32→u32 rvv f32→u32

251

560

941

209

938

4093

657

1468

221

1458

65536

664

1506

222

1506

1048576

659

1350

221

1398

Key points:

  • f32→u32 is the standout: ~4.5–6.8x over scalar, because Lisp float to integer conversion (truncate) is expensive, while vfcvt.xu.f.v is a single instruction.

  • u32→f32 is ~1.7–2.3x over scalar coerce.

Integer/float width change (spec 07)

Table 9. Integer width-change VOPs (Melem/s; scalar vs RVV)
n scalar u8→u32 rvv u8→u32 scalar u16→u32 rvv u16→u32 scalar u32→u16 rvv u32→u16

4093

660

4093

660

2924

499

2924

65536

664

3998

664

3029

499

3029

1048576

664

3985

664

3066

497

1954

Table 10. Integer width-change speedups (higher is better)
n u8→u32 (vzext.vf4) u16→u32 (vzext.vf2) u32→u16 (vnsrl.wi)

4093

6.71x

4.33x

5.81x

65536

6.07x

4.53x

6.04x

1048576

5.97x

4.61x

3.92x

Table 11. Float width-change VOPs (Melem/s; scalar vs RVV)
n scalar f32→f64 rvv f32→f64 scalar f64→f32 rvv f64→f32

4093

660

1462

660

1462

65536

664

1515

664

1515

1048576

655

1286

655

878

Table 12. Float width-change speedups (higher is better)
n f32→f64 (vfwcvt.f.f.v) f64→f32 (vfncvt.f.f.w)

4093

2.16x

2.23x

65536

2.28x

2.28x

1048576

1.96x

1.34x

Key points:

  • Integer widening (u8/u16 → u32) is ~4.3–6.7x: the RVV path fuses load-extend-store, so it makes one memory pass where the scalar path zero-extends each narrow element into a 32-bit lane.

  • u32→u16 narrowing is ~3.9–6.0x: vnsrl.wi 0 truncates and halves the lane count in one instruction, while scalar must mask each element to 16 bits.

  • f32→f64 widening is ~2.0–2.3x and stays cache-resident through n=1M. f64→f32 narrowing drops to ~1.3x at n=1M: the 8-byte source element makes both paths DRAM-bound (the same effect seen in the f64 add and gather rows), so the instruction-count win no longer shows.

Widening arithmetic (spec 07, e32 → e64)

Note
measured on the --with-riscv-zba build; the scalar columns below already include the Zba fused shift+add addressing.
Table 13. Widening arithmetic VOPs (Melem/s; scalar vs RVV)
n scalar s32 wadd rvv s32 wadd scalar s32 wmul rvv s32 wmul

251

363.0

940.1

445.7

942.9

4093

396.4

1102.6

495.0

1104.8

65536

396.6

1177.0

497.9

1230.1

1048576

388.6

689.4

476.4

663.7

Table 14. Widening arithmetic (unsigned) VOPs (Melem/s; scalar vs RVV)
n scalar u32 wadd rvv u32 wadd scalar u32 wmul rvv u32 wmul

251

364.2

936.6

445.4

928.9

4093

396.7

1105.4

494.7

1130.4

65536

394.7

1183.8

496.1

1278.8

1048576

386.1

675.1

492.7

700.0

Key points:

  • Widening arithmetic turns two e32 lanes into one e64 result in a single instruction (vwadd.vv/vwaddu.vv, vwmul.vv/vwmulu.vv), so the RVV path is ~2.1–3.0x faster than the scalar loop in cache (add ~2.6–3.0x, mul ~2.1–2.6x) and ~1.4–1.8x at n=1M, where the 8-byte destination makes both paths DRAM-bound.

  • The scalar widening-multiply baseline (~495 Melem/s) is higher than widening-add (~396) because the scalar 32x32→64 multiply is more work, but the RVV numbers are the same ~930–1280 Melem/s for both, so the multiply’s speedup is slightly lower.

Widening FMA (spec 06/07, e32 sources + e64 accumulator)

Note
measured on the --with-riscv-zba build (same as the widening arithmetic table above).
Table 15. Widening FMA VOPs (Melem/s; scalar vs RVV, dst = a*b + c at 64 bits)
n scalar f32 wfma rvv f32 wfma scalar u32 wmacc rvv u32 wmacc

251

361.4

794.8

364.4

744.4

4093

385.5

1000.6

390.4

1026.6

65536

380.1

953.8

394.5

930.3

1048576

343.6

384.0

352.4

392.4

Key points:

  • vfwmacc.vv/vwmaccu.vv fuse a widening multiply with a 64-bit accumulate: ~2.0–2.6x in cache, collapsing to ~1.1x at n=1M where the 8-byte destination and 64-bit accumulator make the loop DRAM-bound (the same effect seen in the f64 rows).

Strided / indexed gather (spec 08, f32)

Table 16. Gather VOPs (scalar seconds vs RVV; higher speedup is better)
n scalar strided rvv strided scalar indexed rvv indexed indexed speedup

4093

0.014

0.012

0.013

0.012

1.09x

65536

0.182

0.172

0.152

0.148

1.03x

Key points:

  • Strided and indexed gather are ~1.0–1.2x over scalar at cache sizes. They are memory-bound (random-ish access) and the scalar baseline is a simple (aref src (* i k)) / (aref src (aref idx i)) loop, so there is little compute to hide; the real value is enabling vector code for non-contiguous access at all.

Unit-stride segment loads (spec 08, u32 AoS→SoA)

Table 17. Segment deinterleave (scalar seconds vs RVV; higher speedup is better)
nf n scalar rvv speedup

2

4093

0.084

0.169

0.49x

2

65536

1.073

0.996

1.08x

3

4093

0.108

0.156

0.69x

3

65536

1.310

1.172

1.12x

4

4093

0.132

0.163

0.81x

4

65536

1.606

1.424

1.13x

5

4093

0.159

0.207

0.77x

5

65536

1.981

1.912

1.04x

6

4093

0.188

0.192

0.98x

6

65536

2.380

2.474

0.96x

7

4093

0.217

0.225

0.96x

7

65536

3.073

2.895

1.06x

8

4093

0.247

0.260

0.95x

8

65536

4.288

3.084

1.39x

Key points:

  • AoS→SoA deinterleave is not a throughput win in the way the arithmetic VOPs are: both the scalar loop and the segment load walk both arrays sequentially, so there is no cache-behaviour gap to exploit, only the per-segment instruction-count saving.

  • The speedup grows with nf (more fields amortise the fixed vsetvli/vlseg/store-stream cost) and with n (the per-call VOP overhead amortises), reaching ~1.4x at nf=8, n=65536.

  • nf=2 at small n is a loss (~0.5x): the segment-load round-trip costs more than the scalar two-load/store loop it replaces. This is the expected, measurement-driven outcome for a memory-side op; its value is correctness + expressiveness (a single instruction deinterleaves AoS into SoA), recorded honestly rather than marketed as a speedup.

Results with aix — AI cores, VLEN=1024 (VLMAX e32 = 32)

Table 18. SBCL scalar vs RVV on VLEN=1024 (Melem/s)
n scalar chunk blk blk4 speedup (blk4)

251

205

863

903

1006

4.9x

4093

217

1104

1111

1019

4.7x

65536

221

1170

1170

1078

4.9x

1048576

209

1122

1127

998

4.8x

Table 19. C calibration with aix (VLEN=1024)
n C scalar C vector C speedup

251

255

2647

10.4x

4093

207

3324

16.1x

65536

253

5287

20.9x

1048576

229

2273

9.9x

Observations:

  • SBCL runs correctly at VLEN=1024 with no code changes: the VOPs query VLMAX at run time through vsetvli, so the same image works across VLEN=128/256/512/1024.

  • The higher speedups here (~4.9x vs ~3.5x) come from a slower scalar baseline (the AI cores run scalar code at ~210 Melem/s vs ~495 Melem/s on the default cores, both on the Zba build), while SBCL RVV throughput stays ~1.1 Gelem/s on both. The vector unit is bandwidth-saturated in both cases.

  • blk4 stops helping at VLEN=1024: with VLMAX=32 a single vector already covers 32 elements, so the 4-way unroll buys little beyond what blk gives. blk and chunk converge because the chunk loop overhead amortizes over 32 elements per call.

New-VOP benchmarks on VLEN=1024

Table 20. LMUL scaling, f32 add chunk VOPs (Melem/s)
n scalar chunk m1 chunk m2 chunk m4 chunk m8

251

206

891

850

1348

1466

4093

217

1101

1865

2457

1493

65536

222

1172

2036

2892

2034

1048576

209

1124

1871

2111

1384

Table 21. LMUL scaling and unroll, f32 add block VOPs (Melem/s)
n scalar blk m1 blk m2 blk m4 blk8

251

207

900

1081

1351

876

4093

218

1117

1728

2848

1195

65536

222

1170

2035

2869

862

1048576

209

1126

1855

2103

935

Table 22. Reduction VOPs (Melem/s)
n f32 sum scalar f32 sum rvv u32 max scalar u32 max rvv u32 max speedup

251

320

445

265

935

3.5x

4093

356

587

295

1843

6.2x

65536

358

597

295

1940

6.6x

1048576

357

596

255

1926

7.6x

Table 23. Mask VOPs (Melem/s)
n scalar max rvv max masked scalar cmp-count rvv cmp-count

251

203

827

231

828

4093

174

995

243

984

65536

172

1042

215

1020

1048576

162

1004

212

994

Table 24. FMA VOPs (Melem/s)
n scalar f32 fma rvv f32 fma scalar u32 macc rvv u32 macc

251

152

630

91

600

4093

160

368

85

556

65536

161

372

85

372

1048576

160

361

85

362

Table 25. Conversion VOPs (Melem/s)
n scalar u32→f32 rvv u32→f32 scalar f32→u32 rvv f32→u32

251

235

939

46

942

4093

250

1714

47

1732

65536

251

1806

47

1790

1048576

223

1674

47

1675

Table 26. Widening arithmetic VOPs (Melem/s; scalar vs RVV, e32 → e64)
n scalar s32 wadd rvv s32 wadd scalar s32 wmul rvv s32 wmul

251

167

810

168

759

4093

176

601

174

555

65536

178

610

176

566

1048576

176

579

175

550

Table 27. Widening arithmetic (unsigned) VOPs (Melem/s; scalar vs RVV)
n scalar u32 wadd rvv u32 wadd scalar u32 wmul rvv u32 wmul

251

166

820

142

757

4093

175

925

147

850

65536

177

609

148

565

1048576

176

590

147

553

Table 28. Widening FMA VOPs (Melem/s; scalar vs RVV, dst = a*b + c at 64 bits)
n scalar f32 wfma rvv f32 wfma scalar u32 wmacc rvv u32 wmacc

251

83

424

107

453

4093

84

350

110

574

65536

82

173

104

197

1048576

82

317

101

318

Key points at VLEN=1024:

  • LMUL=2/4 is an even bigger win than at VLEN=256: chunk m4 reaches ~13.0x at n=65536 and stays at ~10.1x at n=1048576 (the AI cores have enough bandwidth that the 1M case is not DRAM-bound the way it is on the default cores). With VLMAX=32, an LMUL=4 group covers 128 elements per instruction.

  • LMUL=8 is not a further win here either: it is ~6.6–9.2x, below M4’s ~10–13x. The load/compute/store chunk is memory-bound, so M4 already saturates bandwidth and the wider 8-register groups only add register pressure (16 registers for one binary op) without extra throughput. LMUL=8 is retained for completeness (spec 02) but is not faster.

  • blk8 is a loss at VLEN=1024 (~3.9–5.5x vs blk m4’s ~13x): with VLMAX=32, a single M1 vector already covers 32 elements, so 8-way unroll only adds per-iteration overhead and register pressure without more memory parallelism. This is the same "unroll saturates" effect seen with `blk4 at VLEN=256, just more pronounced.

  • f32/f64 sum (vfredosum) now shows a modest ~1.4–1.7x (vs ~1x at VLEN=256) because the scalar baseline on the AI cores is slower; u32 max stays a strong ~3.5–7.6x.

  • f32→u32 is the largest single speedup measured anywhere in this suite: ~20–38x at VLEN=1024, because the scalar truncate path runs at only ~47 Melem/s on these cores while vfcvt.xu.f.v is one instruction.