This file records the benchmark results for the RVV (RISC-V Vector 1.0) port of SBCL, measured on a riscv64 host on two different sets of cores:
-
default cores — VLEN=256 (VLMAX for e32 = 8 elements)
-
AI-accelerated cores — VLEN=1024 (VLMAX for e32 = 32 elements), reached with the
$HOME/bin/ailauncher (which execs$HOME/bin/aix).
The Lisp benchmark is rvv-bench.lisp; the C calibration is
cbench.c. Both assert or demonstrate correctness before timing, so
every number below is from a verified run. Each Lisp cell is the best
of three GC-quieted runs, reported in millions of elements per second
(Melem/s, higher is better).
Environment
-
SBCL: built from this repository with
:riscv-vectorsupport (featurescontains:riscv-vector), and with the opt-in--with-riscv-zbaflag (featurescontains:riscv-zba; see INSTALL). -
gcc: (Bianbu 15.2.0-16ubuntu1bb5) 15.2.0
-
binutils: GNU Binutils for Bianbu 2.46
-
clang: Bianbu clang 21.1.8
-
The M1/chunk/blk/blk4 paths use SEW=e32/e64 and LMUL=1; the other benchmarks also exercise LMUL=2/4/8, whole-array reductions, mask compare/merge, fused multiply-add, same-width float <→ integer conversions, integer/float width changes (spec 07), and (spec 08) strided/indexed gather.
How to reproduce
# Lisp benchmark — default cores (VLEN=256), no aix:
./run-sbcl.sh --noinform --disable-debugger \
--script rvv-results/rvv-bench.lisp
# Lisp benchmark — AI cores (VLEN=1024), with aix:
PATH="$HOME/bin:$PATH" \
$HOME/bin/ai ./run-sbcl.sh --noinform --disable-debugger \
--script rvv-results/rvv-bench.lisp
# C calibration — default cores:
gcc -O3 -march=rv64gcv -fno-tree-vectorize rvv-results/cbench.c -o cbench
./cbench
# C calibration — AI cores:
PATH="$HOME/bin:$PATH" $HOME/bin/ai ./cbench
What is measured
-
scalar — a plain Lisp
loopover(simple-array single-float (*))(or(unsigned-byte 32),double-float) compiled at(speed 3) (safety 0). -
chunk — the
%RVV-*-CHUNKVOPs called from a Lisp loop; each call does onevsetvli-bounded load/compute/store and returns how many elements it processed. -
blk —
%RVV-F32-ADD-BLOCK, the whole array handled by an assembly-internal strip-mine loop inside a single VOP. -
blk4 —
%RVV-F32-ADD-BLOCK4, the same but 4-way unrolled over all eight caller-saved vector registers (v24-v31). -
chunk/blk m2, m4, m8 — the LMUL=2/4/8 register-group variants of the chunk and block loops; one instruction covers 2x/4x/8x the elements (chunk m8 uses the full v16-v31 scratch window).
-
blk8 —
%RVV-F32-ADD-BLOCK8, an 8-way unrolled block loop (v16-v31). -
reduce —
%RVV-F32/F64-REDUCE-ADDand%RVV-U32-REDUCE-MAX, whole-array scalar-out reductions (a single VOP strip-mines the array and folds one vector into an accumulator). -
mask —
%RVV-U32-MAX-MASKED(compare + merge) and%RVV-U32-CMP-COUNT(compare +vcpop.m). -
fma —
%RVV-F32-FMA-CHUNK/%RVV-U32-MACC-CHUNK(dst = src1*src2 + src3). -
fma overwrite —
%RVV-F32-FMADD-CHUNK(vfmadd.vv), the overwrite-multiplicand form of the same dst = a*b + c contract. -
convert —
%RVV-U32→F32-CHUNK/%RVV-F32→U32-CHUNK(same-width float <→ integer conversion, SEW=e32). -
width —
%RVV-U8→U32-CHUNK(vzext.vf4),%RVV-U16→U32-CHUNK(vzext.vf2),%RVV-U32→U16-CHUNK(vnsrl.wi),%RVV-F32→F64-CHUNK(vfwcvt.f.f.v) and%RVV-F64→F32-CHUNK(vfncvt.f.f.w) — integer and float width changes (spec 07). -
widen arith —
%RVV-S32/U32-WIDEN-ADD-CHUNK(vwadd.vv/vwaddu.vv) and%RVV-S32/U32-WIDEN-MUL-CHUNK(vwmul.vv/vwmulu.vv), widening arithmetic from e32 to e64 (spec 07). -
widen fma —
%RVV-F32-WIDEN-FMA-CHUNK(vfwmacc.vv) and%RVV-U32-WIDEN-MACC-CHUNK(vwmaccu.vv), a widening multiply folded into a 64-bit accumulator (spec 06/07). -
gather —
%RVV-F32-STRIDED-LOAD(vlse32.v) and%RVV-F32-INDEXED-LOAD(vluxei32.v) gather a strided column or an indirectly-indexed set of elements into a compact array (spec 08). -
segment —
%RVV-U32-SEGMENT-LOAD-2…-8(vlseg2e32.v…vlseg8e32.v) deinterleave an AoS u32 stream into a field-major SoA destination (spec 08).
Summary
-
Net win — element-wise compute. RVV gives 2.1–6.3x over scalar Lisp for cache-resident f32/u32 (up to ~7.9x for f32 add with LMUL=4), and mask compare/merge, FMA, float↔int conversion, integer/float width change, and strided/indexed gather add another 1.2–6.8x. LMUL=2/4 closes the gap to gcc’s auto-vectorizer; every family is correctness-asserted before timing.
-
Neutral to loss — memory- and latency-bound paths. Ordered float reductions sit at ~1x (order preservation forbids the parallel tree reduction), DRAM-bound f64 add and strided gather sit at ~0.6–1.0x, segment deinterleave at small nf is ~0.5–1.0x (unamortized vsetvli+vlseg+store overhead), and blk8 unrolling is neutral-to-loss.
-
Caveat — weak scalar baseline. Lisp scalar is ~2.0–2.5x slower than C (per-element address recomputation, no indexed addressing, no loop unrolling);
--with-riscv-zbarecovers ~1.23x (gap → ~2.0x). So the RVV speedups are measured against a conservative scalar loop, not against C.
Findings
-
RVV gives a real 2.1–6.3x element-wise speedup over scalar Lisp for cache-resident f32/u32 arrays, and ~1.0–1.4x for f64.
-
LMUL=2/4 closes most of the previously identified gap: at VLEN=256 chunk m4 reaches ~3.9 Gelem/s at n=4093, in the same range as gcc’s auto-vectorizer (~3.4 Gelem/s), and the f32 add speedup goes from ~2.3x (LMUL=1) to ~7.9x (LMUL=4) in cache. At VLEN=1024 the LMUL win is larger still (~10–13x, see the aix tables below). LMUL=8 adds no further throughput at either VLEN: the memory-bound chunk is already bandwidth-saturated at M4, so the wider groups only add register pressure (see the LMUL key points).
-
The bandwidth ceiling is the limiter only on the default cores: at n=1M (VLEN=256) all LMUL variants converge to ~0.72–0.76 Gelem/s and ~1.5–1.7x, matching the C DRAM-bound result (~1.1x); the AI cores' higher bandwidth keeps LMUL=4 at ~2 Gelem/s even at n=1M.
-
Scalar-out reductions are a mixed bag: associative reductions with a branch (
u32 max) win ~3.6–6.3x on the default cores (more on the slower AI-core scalar baseline), but the ordered float sum (vfredosum) is neutral against a tight scalar loop on the default cores (~1x) and only ~2x on the AI cores. -
Mask compare/merge (~2.1–4.1x), fused multiply-add (~2.4–3.8x), and float <→ integer conversion (~1.7–2.3x for u32→f32, ~4.5–6.8x for f32→u32) all deliver useful element-wise wins with no correctness issues.
-
Integer/float width change (spec 07) is another solid win. The fractional-LMUL refinement (spec 07’s deferred half) raises LMUL for the widening VOPs so the narrow source fills a whole register: u8 → u32 widening goes from ~2.5x to ~6.0–6.5x, u16→u32 from ~2.3x to ~4.4–4.6x. u32→u16 narrowing stays ~4–6x and f32→f64 / f64→f32 width change ~1.3–2.3x, with no regressions.
-
Unrolling does not scale past the memory-level parallelism the hardware can exploit:
blk8is neutral at VLEN=256 and a loss at VLEN=1024, soblk4/LMUL scaling is the better lever than wider unrolling for f32 add. -
Strided/indexed gather (spec 08) is a modest ~1.0–1.2x in cache: these ops are memory-bound (random-ish access) and their main value is enabling vectorized non-contiguous access rather than raw throughput.
-
Unit-stride segment loads (spec 08’s last piece) deinterleave an AoS u32 stream into a field-major SoA destination in one
vlseg<nf>e32.vplusnfstores. The win is small (~1.2–1.4x atnf=6..8, and a loss atnf=2) because both memory streams are already sequential; the value is the single-instruction AoS→SoA deinterleave rather than raw bandwidth. -
The C scalar column is ~2.4–2.5x faster than the Lisp scalar column for f32 add (994 vs 396 Melem/s at n=4093) purely from scalar code generation, not memory bandwidth. Lisp scalar is flat at ~396 Melem/s from n=251 (L1) to n=1M (DRAM) while C drops from ~960 to ~691, i.e. the Lisp loop is instruction/latency-bound. The cause is per-element array-address recomputation (RISC-V has no scaled/indexed addressing, so each
arefcostsslli+add+load) plus the absence of loop unrolling in SBCL. Manually unrolling the Lisp loop 4–8x lifts it to ~565–585 Melem/s. This is a scalar-baseline artifact: the RVV VOPs already match/beat gcc’s auto-vectorizer. See slow-scalar.adoc. -
Building with
--with-riscv-zbareplaces the per-arefslli+addaddress computation with the fusedsh1add/sh2add/sh3addinstructions, lifting the scalar f32 add baseline from ~396 to ~486–496 Melem/s (~1.23–1.28x). This narrows the gap to C’s ~994 Melem/s from ~2.5x to ~2.0x. The RVV Melem/s numbers are unchanged, so the scalar-vs-RVV speedup ratios are correspondingly lower on a Zba build (faster scalar, same vector throughput). Every table in this file was re-measured on the--with-riscv-zbabuild, so all scalar columns below reflect the Zba fused shift+add addressing. -
Widening arithmetic (spec 07’s last piece) and widening FMA (spec 06’s last piece) both deliver solid element-wise wins.
vwadd.vv/vwaddu.vvare ~2.6–3.0x andvwmul.vv/vwmulu.vv~2.1–2.6x over scalar in cache, andvfwmacc.vv/vwmaccu.vv~2.0–2.6x; all collapse to ~1.1–1.8x at n=1M where the 8-byte destination makes the loop DRAM-bound. The overwrite-multiplicandvfmadd.vvis within noise of the accumulatevfmacc.vv(~2.8x), as expected for the same instruction shape. -
A handful of VOPs are not faster than scalar, and that is expected rather than a bug. Three distinct causes:
-
Ordered reductions — the
f32/f64 sumVOPs usevfredosum.vs, the ordered sum, to match Lispreduce #'+element order; ordering forbids the parallel tree reduction, so they sit at ~0.9–1.0x against a tight scalar loop. -
Memory-bound, no compute to hide —
f64 addat n=1M (~1.0x) and strided/indexed gather on the AI cores (~0.6–0.7x): both paths hit the same bandwidth ceiling, and the vector path’s extra instructions (plus, for gather, a defeated prefetcher) add cost without recovering it. -
Unamortized per-call overhead — segment deinterleave at small
nf/n(~0.5–1.0x): thevsetvli+vlseg+store round-trip costs more than the scalar two-load/store loop it replaces, and only wins oncenf/ngrow enough to amortize it. These last two families exist for correctness and expressiveness (single-instruction AoS→SoA deinterleave, vectorized non-contiguous access), not raw throughput.
-
Results without aix — default cores, VLEN=256 (VLMAX e32 = 8)
| n | scalar | chunk | blk | blk4 | speedup (blk4) |
|---|---|---|---|---|---|
251 |
434 |
897 |
906 |
1498 |
3.5x |
4093 |
495 |
1130 |
1522 |
3117 |
6.3x |
65536 |
496 |
1254 |
1579 |
2159 |
4.4x |
1048576 |
487 |
776 |
753 |
738 |
1.5x |
Key points:
-
RVV is ~2.1–3.5x faster than scalar for the small/medium sizes where the arrays stay in cache, and ~6.3x at n=4093 for the 4-way-unrolled block loop.
-
blkandblk4beatchunksubstantially at n=4093 and n=65536, showing that the per-call Lisp loop overhead of the chunk API is real: moving the strip-mine loop into assembly recovers it. -
At n=1048576 everything is DRAM-bound: vector throughput drops to ~0.74 Gelem/s and the advantage collapses to ~1.5x.
compiled -fno-tree-vectorize, vector loop with tree-vectorize)
| n | C scalar | C vector | C speedup |
|---|---|---|---|
251 |
960 |
3047 |
3.2x |
4093 |
994 |
3384 |
3.4x |
65536 |
752 |
2452 |
3.3x |
1048576 |
691 |
759 |
1.1x |
C shows the same shape (vector wins in cache, DRAM-bound collapse at n=1M), which confirms the Lisp numbers are measuring the memory hierarchy, not an SBCL artifact.
Why C scalar beats Lisp scalar
|
Note
|
this section analyzes the pre-Zba scalar baseline (a flat ~396 Melem/s); the following section shows how Zba narrows the gap. |
The C scalar column (994 Melem/s at n=4093) is ~2.5x the Lisp scalar column (396). The gap is code generation, not memory bandwidth:
-
Lisp scalar is flat at ~396 Melem/s from n=251 (L1-resident) to n=1M (DRAM), whereas C scalar drops from ~960 to ~691. A memory-bound loop tracks the cache like C does; a flat line means the Lisp loop is instruction/latency-bound.
-
SBCL recomputes each array element’s byte address from the tagged fixnum index on every element (
sllito untag/scale, thenaddto the base pointer), three times per element for the three arrays. RISC-V has no scaled/indexed addressing mode, so this cannot be folded into the load the way x86-64 (movss [base+idx*4+disp]) and arm64 (addwith foldedlsl) can. gcc strength-reduces the loop to three pointer-induction variables; SBCL does not. -
SBCL does not auto-unroll scalar loops. Manual unrolling closes most of the gap:
| variant (n=4093) | Melem/s |
|---|---|
base, 1 element/iter |
396 |
unroll 4x |
565 |
unroll 8x |
585 |
-
The benchmark’s
of-type index(SBCL’sindexis(integer 0 array-dimension-limit), which this build reports as not a subtype offixnum) also makes the loop-index arithmetic go through genericTWO-ARG-+/TWO-ARG→=calls. Switching the loop variable tofixnumproduces a much tighter loop but does not move the measured number on this benchmark (the loop is already latency-bound), so it is a codegen smell rather than the bottleneck.
The scalar-vs-RVV speedups in this file are therefore measured against a conservative scalar baseline. The RVV VOPs already reach parity with gcc’s auto-vectorizer (chunk m4 ~3.7 vs ~3.4 Gelem/s at n=4093), so the RVV conclusions are unaffected. A fuller treatment, including what could change in the SBCL compiler, is in slow-scalar.adoc.
Zba (sh1add/sh2add/sh3add) scalar baseline
Building with --with-riscv-zba (opt-in, see INSTALL) lets the array
VOPs use the RISC-V Zba fused shift+add instructions for scaled element
addressing instead of slli+add. For a single-float aref the byte
address becomes sh1add (the << 1 also untags the fixnum index), for
double-float sh2add, and for (unsigned-byte 32)/complex sh2add/
sh3add, removing one instruction per array reference.
| n | scalar (no Zba) | scalar (Zba) | speedup |
|---|---|---|---|
251 |
358 |
459.5 |
1.28x |
4093 |
396 |
496.1 |
1.25x |
65536 |
399 |
495.0 |
1.24x |
1048576 |
396 |
485.8 |
1.23x |
-
Zba removes one
slliperaref: the scalar f32 add loop goes from a flat ~396 Melem/s to ~486–496 Melem/s (~1.23–1.28x), and the inner loop now disassembles to threesh1addand noslli. -
This narrows but does not close the gap to C (~994 Melem/s at n=4093): the remaining ~2x is loop unrolling and pointer-induction strength reduction, compiler issues independent of Zba (see slow-scalar.adoc).
-
The scalar-vs-RVV tables in this file have all been re-measured on the Zba build, so their scalar columns already reflect the Zba speedup; the speedup ratios are correspondingly lower than on a non-Zba build (faster scalar, unchanged RVV throughput).
LMUL=1/2/4 chunk and block (f32 add, SEW=e32)
| n | scalar | chunk m1 | chunk m2 | chunk m4 | chunk m8 |
|---|---|---|---|---|---|
251 |
436 |
894 |
1433 |
1394 |
1238 |
4093 |
494 |
1132 |
2215 |
3883 |
3047 |
65536 |
497 |
1216 |
2335 |
2389 |
2310 |
1048576 |
457 |
738 |
727 |
786 |
730 |
| n | scalar | blk m1 | blk m2 | blk m4 | blk8 |
|---|---|---|---|---|---|
251 |
439 |
909 |
1227 |
1354 |
1170 |
4093 |
494 |
1528 |
2749 |
3944 |
3058 |
65536 |
495 |
1582 |
2353 |
2426 |
2186 |
1048576 |
492 |
751 |
762 |
765 |
758 |
Key points:
-
LMUL=2/4 gives a large, real speedup in cache: at n=4093 the chunk path goes from 2.3x (m1) to 4.5x (m2) to 7.9x (m4), and the block path reaches 8.4x (m4). This is exactly the instruction-count reduction predicted: one LMUL=4 vector instruction covers VLEN*4/32 = 32 elements.
-
At n=1048576 the memory system is the limit: m1/m2/m4/m8 all converge to ~1.5–1.6x and ~0.72–0.79 Gelem/s.
-
LMUL=8 is neutral to slightly worse than M4 at VLEN=256 (6.2x vs 7.9x at n=4093): with VLMAX=8 an M8 group already covers 64 elements per instruction, but the two 8-register groups (16 registers total) add register pressure without extra memory parallelism. See the VLEN=1024 tables below, where M8 covers 256 elements and is expected to help more.
-
blk8(8-way unroll) is not a win overblk m4at VLEN=256: at n=4093 blk8 is 6.2x vs blk m4’s ~8.0x, and at n=65536 4.5x vs 4.9x. With VLMAX=8 a 4-way unroll already keeps 32 elements in flight, so 8-way adds register pressure without more memory parallelism. -
chunk m4 at n=4093 reaches ~3.9 Gelem/s, in the same neighborhood as the C auto-vectorizer (3.4 Gelem/s at n=4093).
Whole-array reductions (scalar out)
| n | f32 sum scalar | f32 sum rvv | u32 max scalar | u32 max rvv | u32 max speedup |
|---|---|---|---|---|---|
251 |
575 |
545 |
437 |
1579 |
3.6x |
4093 |
659 |
657 |
491 |
2996 |
6.1x |
65536 |
664 |
664 |
498 |
3158 |
6.3x |
1048576 |
662 |
663 |
497 |
2086 |
4.2x |
Key points:
-
u32 max(vredmaxu.vs) is a clear win: ~3.6–6.3x over the scalar loop, because the integer max has a loop-carried branch that the scalar path cannot hide, while the vector reduction is parallel. -
f32/f64 sum(vfredosum.vs) does not beat a tight scalar loop (~1.0x and ~0.9x).vfredosumis the ordered sum variant; it preserves element order (required to match Lispreduce #'+) and therefore does not gain the tree-reduction parallelism of the unorderedvfredusum. This is a documented, honest result rather than a bug.
Mask compare/merge (u32)
| n | scalar max | rvv max masked | scalar cmp-count | rvv cmp-count |
|---|---|---|---|---|
251 |
428 |
917 |
429 |
760 |
4093 |
456 |
1284 |
467 |
1013 |
65536 |
320 |
1322 |
384 |
1037 |
1048576 |
268 |
743 |
350 |
855 |
Key points:
-
u32 max masked(vmsgtu.vv → vmerge.vvm) is ~2.1–4.1x over scalar max, comparable to (slightly below) the single-instructionvmaxu.vvchunk, as expected for a two-instruction sequence. -
u32 cmp-count(vmsgtu.vv → vcpop.m) is ~1.8–2.7x over a scalar count loop.
Fused multiply-add (dst = a*b + c)
| n | scalar f32 fma | rvv f32 fma | scalar u32 macc | rvv u32 macc |
|---|---|---|---|---|
251 |
363 |
864 |
323 |
764 |
4093 |
396 |
1141 |
337 |
1131 |
65536 |
367 |
1062 |
281 |
1067 |
1048576 |
353 |
473 |
279 |
476 |
Key points:
-
FMA/MACC fuses two operations into one instruction: ~2.4–2.9x for f32, ~2.4–3.8x for u32 in cache, collapsing to ~1.3–1.7x at n=1M where the arrays stop fitting in cache.
-
The overwrite-multiplicand form (
vfmadd.vv,%RVV-F32-FMADD-CHUNK) performs within noise of the accumulate form (vfmacc.vv): ~2.8x at n=4093 and n=65536, ~2.2x at n=251, and ~1.3x at n=1M — the same instruction shape, so no extra cost.
Same-width float <→ integer conversions (SEW=e32)
| n | scalar u32→f32 | rvv u32→f32 | scalar f32→u32 | rvv f32→u32 |
|---|---|---|---|---|
251 |
560 |
941 |
209 |
938 |
4093 |
657 |
1468 |
221 |
1458 |
65536 |
664 |
1506 |
222 |
1506 |
1048576 |
659 |
1350 |
221 |
1398 |
Key points:
-
f32→u32is the standout: ~4.5–6.8x over scalar, because Lisp float to integer conversion (truncate) is expensive, whilevfcvt.xu.f.vis a single instruction. -
u32→f32is ~1.7–2.3x over scalarcoerce.
Integer/float width change (spec 07)
| n | scalar u8→u32 | rvv u8→u32 | scalar u16→u32 | rvv u16→u32 | scalar u32→u16 | rvv u32→u16 |
|---|---|---|---|---|---|---|
4093 |
660 |
4093 |
660 |
2924 |
499 |
2924 |
65536 |
664 |
3998 |
664 |
3029 |
499 |
3029 |
1048576 |
664 |
3985 |
664 |
3066 |
497 |
1954 |
| n | u8→u32 (vzext.vf4) |
u16→u32 (vzext.vf2) |
u32→u16 (vnsrl.wi) |
|---|---|---|---|
4093 |
6.71x |
4.33x |
5.81x |
65536 |
6.07x |
4.53x |
6.04x |
1048576 |
5.97x |
4.61x |
3.92x |
| n | scalar f32→f64 | rvv f32→f64 | scalar f64→f32 | rvv f64→f32 |
|---|---|---|---|---|
4093 |
660 |
1462 |
660 |
1462 |
65536 |
664 |
1515 |
664 |
1515 |
1048576 |
655 |
1286 |
655 |
878 |
| n | f32→f64 (vfwcvt.f.f.v) |
f64→f32 (vfncvt.f.f.w) |
|---|---|---|
4093 |
2.16x |
2.23x |
65536 |
2.28x |
2.28x |
1048576 |
1.96x |
1.34x |
Key points:
-
Integer widening (u8/u16 → u32) is ~4.3–6.7x: the RVV path fuses load-extend-store, so it makes one memory pass where the scalar path zero-extends each narrow element into a 32-bit lane.
-
u32→u16 narrowing is ~3.9–6.0x:
vnsrl.wi 0truncates and halves the lane count in one instruction, while scalar must mask each element to 16 bits. -
f32→f64 widening is ~2.0–2.3x and stays cache-resident through n=1M. f64→f32 narrowing drops to ~1.3x at n=1M: the 8-byte source element makes both paths DRAM-bound (the same effect seen in the f64 add and gather rows), so the instruction-count win no longer shows.
Widening arithmetic (spec 07, e32 → e64)
|
Note
|
measured on the --with-riscv-zba build; the scalar columns
below already include the Zba fused shift+add addressing.
|
| n | scalar s32 wadd | rvv s32 wadd | scalar s32 wmul | rvv s32 wmul |
|---|---|---|---|---|
251 |
363.0 |
940.1 |
445.7 |
942.9 |
4093 |
396.4 |
1102.6 |
495.0 |
1104.8 |
65536 |
396.6 |
1177.0 |
497.9 |
1230.1 |
1048576 |
388.6 |
689.4 |
476.4 |
663.7 |
| n | scalar u32 wadd | rvv u32 wadd | scalar u32 wmul | rvv u32 wmul |
|---|---|---|---|---|
251 |
364.2 |
936.6 |
445.4 |
928.9 |
4093 |
396.7 |
1105.4 |
494.7 |
1130.4 |
65536 |
394.7 |
1183.8 |
496.1 |
1278.8 |
1048576 |
386.1 |
675.1 |
492.7 |
700.0 |
Key points:
-
Widening arithmetic turns two e32 lanes into one e64 result in a single instruction (
vwadd.vv/vwaddu.vv,vwmul.vv/vwmulu.vv), so the RVV path is ~2.1–3.0x faster than the scalar loop in cache (add ~2.6–3.0x, mul ~2.1–2.6x) and ~1.4–1.8x at n=1M, where the 8-byte destination makes both paths DRAM-bound. -
The scalar widening-multiply baseline (~495 Melem/s) is higher than widening-add (~396) because the scalar 32x32→64 multiply is more work, but the RVV numbers are the same ~930–1280 Melem/s for both, so the multiply’s speedup is slightly lower.
Widening FMA (spec 06/07, e32 sources + e64 accumulator)
|
Note
|
measured on the --with-riscv-zba build (same as the widening
arithmetic table above).
|
| n | scalar f32 wfma | rvv f32 wfma | scalar u32 wmacc | rvv u32 wmacc |
|---|---|---|---|---|
251 |
361.4 |
794.8 |
364.4 |
744.4 |
4093 |
385.5 |
1000.6 |
390.4 |
1026.6 |
65536 |
380.1 |
953.8 |
394.5 |
930.3 |
1048576 |
343.6 |
384.0 |
352.4 |
392.4 |
Key points:
-
vfwmacc.vv/vwmaccu.vvfuse a widening multiply with a 64-bit accumulate: ~2.0–2.6x in cache, collapsing to ~1.1x at n=1M where the 8-byte destination and 64-bit accumulator make the loop DRAM-bound (the same effect seen in the f64 rows).
Strided / indexed gather (spec 08, f32)
| n | scalar strided | rvv strided | scalar indexed | rvv indexed | indexed speedup |
|---|---|---|---|---|---|
4093 |
0.014 |
0.012 |
0.013 |
0.012 |
1.09x |
65536 |
0.182 |
0.172 |
0.152 |
0.148 |
1.03x |
Key points:
-
Strided and indexed gather are ~1.0–1.2x over scalar at cache sizes. They are memory-bound (random-ish access) and the scalar baseline is a simple
(aref src (* i k))/(aref src (aref idx i))loop, so there is little compute to hide; the real value is enabling vector code for non-contiguous access at all.
Unit-stride segment loads (spec 08, u32 AoS→SoA)
| nf | n | scalar | rvv | speedup |
|---|---|---|---|---|
2 |
4093 |
0.084 |
0.169 |
0.49x |
2 |
65536 |
1.073 |
0.996 |
1.08x |
3 |
4093 |
0.108 |
0.156 |
0.69x |
3 |
65536 |
1.310 |
1.172 |
1.12x |
4 |
4093 |
0.132 |
0.163 |
0.81x |
4 |
65536 |
1.606 |
1.424 |
1.13x |
5 |
4093 |
0.159 |
0.207 |
0.77x |
5 |
65536 |
1.981 |
1.912 |
1.04x |
6 |
4093 |
0.188 |
0.192 |
0.98x |
6 |
65536 |
2.380 |
2.474 |
0.96x |
7 |
4093 |
0.217 |
0.225 |
0.96x |
7 |
65536 |
3.073 |
2.895 |
1.06x |
8 |
4093 |
0.247 |
0.260 |
0.95x |
8 |
65536 |
4.288 |
3.084 |
1.39x |
Key points:
-
AoS→SoA deinterleave is not a throughput win in the way the arithmetic VOPs are: both the scalar loop and the segment load walk both arrays sequentially, so there is no cache-behaviour gap to exploit, only the per-segment instruction-count saving.
-
The speedup grows with
nf(more fields amortise the fixedvsetvli/vlseg/store-stream cost) and withn(the per-call VOP overhead amortises), reaching ~1.4x atnf=8, n=65536. -
nf=2at smallnis a loss (~0.5x): the segment-load round-trip costs more than the scalar two-load/store loop it replaces. This is the expected, measurement-driven outcome for a memory-side op; its value is correctness + expressiveness (a single instruction deinterleaves AoS into SoA), recorded honestly rather than marketed as a speedup.
Results with aix — AI cores, VLEN=1024 (VLMAX e32 = 32)
| n | scalar | chunk | blk | blk4 | speedup (blk4) |
|---|---|---|---|---|---|
251 |
205 |
863 |
903 |
1006 |
4.9x |
4093 |
217 |
1104 |
1111 |
1019 |
4.7x |
65536 |
221 |
1170 |
1170 |
1078 |
4.9x |
1048576 |
209 |
1122 |
1127 |
998 |
4.8x |
| n | C scalar | C vector | C speedup |
|---|---|---|---|
251 |
255 |
2647 |
10.4x |
4093 |
207 |
3324 |
16.1x |
65536 |
253 |
5287 |
20.9x |
1048576 |
229 |
2273 |
9.9x |
Observations:
-
SBCL runs correctly at VLEN=1024 with no code changes: the VOPs query VLMAX at run time through
vsetvli, so the same image works across VLEN=128/256/512/1024. -
The higher speedups here (~4.9x vs ~3.5x) come from a slower scalar baseline (the AI cores run scalar code at ~210 Melem/s vs ~495 Melem/s on the default cores, both on the Zba build), while SBCL RVV throughput stays ~1.1 Gelem/s on both. The vector unit is bandwidth-saturated in both cases.
-
blk4stops helping at VLEN=1024: with VLMAX=32 a single vector already covers 32 elements, so the 4-way unroll buys little beyond whatblkgives.blkandchunkconverge because the chunk loop overhead amortizes over 32 elements per call.
New-VOP benchmarks on VLEN=1024
| n | scalar | chunk m1 | chunk m2 | chunk m4 | chunk m8 |
|---|---|---|---|---|---|
251 |
206 |
891 |
850 |
1348 |
1466 |
4093 |
217 |
1101 |
1865 |
2457 |
1493 |
65536 |
222 |
1172 |
2036 |
2892 |
2034 |
1048576 |
209 |
1124 |
1871 |
2111 |
1384 |
| n | scalar | blk m1 | blk m2 | blk m4 | blk8 |
|---|---|---|---|---|---|
251 |
207 |
900 |
1081 |
1351 |
876 |
4093 |
218 |
1117 |
1728 |
2848 |
1195 |
65536 |
222 |
1170 |
2035 |
2869 |
862 |
1048576 |
209 |
1126 |
1855 |
2103 |
935 |
| n | f32 sum scalar | f32 sum rvv | u32 max scalar | u32 max rvv | u32 max speedup |
|---|---|---|---|---|---|
251 |
320 |
445 |
265 |
935 |
3.5x |
4093 |
356 |
587 |
295 |
1843 |
6.2x |
65536 |
358 |
597 |
295 |
1940 |
6.6x |
1048576 |
357 |
596 |
255 |
1926 |
7.6x |
| n | scalar max | rvv max masked | scalar cmp-count | rvv cmp-count |
|---|---|---|---|---|
251 |
203 |
827 |
231 |
828 |
4093 |
174 |
995 |
243 |
984 |
65536 |
172 |
1042 |
215 |
1020 |
1048576 |
162 |
1004 |
212 |
994 |
| n | scalar f32 fma | rvv f32 fma | scalar u32 macc | rvv u32 macc |
|---|---|---|---|---|
251 |
152 |
630 |
91 |
600 |
4093 |
160 |
368 |
85 |
556 |
65536 |
161 |
372 |
85 |
372 |
1048576 |
160 |
361 |
85 |
362 |
| n | scalar u32→f32 | rvv u32→f32 | scalar f32→u32 | rvv f32→u32 |
|---|---|---|---|---|
251 |
235 |
939 |
46 |
942 |
4093 |
250 |
1714 |
47 |
1732 |
65536 |
251 |
1806 |
47 |
1790 |
1048576 |
223 |
1674 |
47 |
1675 |
| n | scalar s32 wadd | rvv s32 wadd | scalar s32 wmul | rvv s32 wmul |
|---|---|---|---|---|
251 |
167 |
810 |
168 |
759 |
4093 |
176 |
601 |
174 |
555 |
65536 |
178 |
610 |
176 |
566 |
1048576 |
176 |
579 |
175 |
550 |
| n | scalar u32 wadd | rvv u32 wadd | scalar u32 wmul | rvv u32 wmul |
|---|---|---|---|---|
251 |
166 |
820 |
142 |
757 |
4093 |
175 |
925 |
147 |
850 |
65536 |
177 |
609 |
148 |
565 |
1048576 |
176 |
590 |
147 |
553 |
| n | scalar f32 wfma | rvv f32 wfma | scalar u32 wmacc | rvv u32 wmacc |
|---|---|---|---|---|
251 |
83 |
424 |
107 |
453 |
4093 |
84 |
350 |
110 |
574 |
65536 |
82 |
173 |
104 |
197 |
1048576 |
82 |
317 |
101 |
318 |
Key points at VLEN=1024:
-
LMUL=2/4 is an even bigger win than at VLEN=256: chunk m4 reaches ~13.0x at n=65536 and stays at ~10.1x at n=1048576 (the AI cores have enough bandwidth that the 1M case is not DRAM-bound the way it is on the default cores). With VLMAX=32, an LMUL=4 group covers 128 elements per instruction.
-
LMUL=8 is not a further win here either: it is ~6.6–9.2x, below M4’s ~10–13x. The load/compute/store chunk is memory-bound, so M4 already saturates bandwidth and the wider 8-register groups only add register pressure (16 registers for one binary op) without extra throughput. LMUL=8 is retained for completeness (spec 02) but is not faster.
-
blk8is a loss at VLEN=1024 (~3.9–5.5x vsblk m4’s ~13x): with VLMAX=32, a single M1 vector already covers 32 elements, so 8-way unroll only adds per-iteration overhead and register pressure without more memory parallelism. This is the same "unroll saturates" effect seen with `blk4at VLEN=256, just more pronounced. -
f32/f64 sum(vfredosum) now shows a modest ~1.4–1.7x (vs ~1x at VLEN=256) because the scalar baseline on the AI cores is slower;u32 maxstays a strong ~3.5–7.6x. -
f32→u32is the largest single speedup measured anywhere in this suite: ~20–38x at VLEN=1024, because the scalartruncatepath runs at only ~47 Melem/s on these cores whilevfcvt.xu.f.vis one instruction.