Elapsed times are hh:mm:ss since task start (wall clock 03:33
PKT, 2026-09-06). Items marked ~ are reconstructed from build
logs, file mtimes and turn timestamps (minute-level accuracy);
everything else is anchored to recorded times. One line per item.
Timeline
-
00:00:00 (03:33) Task received: add RVV 1.0 (vector) support to SBCL’s riscv64gc backend; free rein on branch
rvv. -
00:01:00 ~ Survey repo tree: src/compiler/riscv (24 files), src/assembly/riscv, src/code/riscv-vm.lisp, src/runtime, src/cold.
-
00:03:00 ~ Read src/compiler/riscv/insts.lisp: formats, printers, bitfield emitters, arg types, inst/inst*.
-
00:05:00 ~ Read vm.lisp (storage bases/classes), c-call.lisp (C ABI), assem-rtns.lisp (save/restore-c-registers).
-
00:07:00 ~ Study x86-64 SIMD precedent: avx2-insts.lisp, simd-pack*.lisp, cpu feature detection; arm64-simd.lisp.
-
00:08:30 ~ Study disassembler internals: printer compilation, :cond/:unless, :constant test, instruction-flavors.
-
00:10:00 ~ Probe host: /proc/cpuinfo shows v/zvbb; clang 21 + binutils 2.46 accept -march=rv64gcv; native SBCL 2.6.6 installed.
-
00:12:00 ~ First gas experiments: store syntax
vse32.v (a0), v1rejected by gas; LLVM integrated assembler accepts vd-first form. -
00:14:00 ~ More syntax probes: whole-register ops need group-aligned registers; masked (v0.t) acceptability probed per family.
-
00:16:00 ~ Write /tmp/rvvgen.py: mnemonic table → .s → clang-assemble → objdump → verified TSV of 32-bit encodings.
-
00:18:00 ~ Iterate generator against assembler errors: .wv/.wi/.wx suffixes, vwmacc (not vwsmacc), unary .m forms, vfsqrt.v, vwsll is Zvbb, vmacc operand order.
-
00:21:00 ~ Add missing families: vaadd/vaaddu/vasub/vasubu, vmin/minu/max/maxu, vrgather.vi/vx; final table 585+ rows.
-
00:23:00 ~ Decode fields of every word; discover pseudo-ops: vmsgt.vv/vmsgtu.vv (swapped vmslt), vmclr/vmset/vmnot, vmerge/vmv vm-bit roles.
-
00:25:00 ~ Verify vtype layout from 0x0d0372d7: lmul m1=000..mf2=110; sew 5:3; ta/ma 6/7; vset* distinguished by bits 31:30.
-
00:27:00 ~ Write /tmp/rvvspec.py: classify each mnemonic into operand kinds (vv/vi/vx/vf/acc/un/wmv/mem/…).
-
00:29:00 ~ Write /tmp/rvvgen2.py: emit src/compiler/riscv/rvv-insts.lisp (formats + emitters + ~560 define-instructions).
-
00:32:00 ~ Emit tests/rvv-assembler.pure.lisp (572 encode-vs-gas cases).
-
00:33:30 ~ Install: vector printers in target-insts.lisp; build-order entry
#+riscv {arch}/rvv-insts. -
00:35:00 ~ Write src/compiler/riscv/vector.lisp v1: chunk VOPs + defknowns + convenience wrappers.
-
00:36:30 ~ Start build #1; it dies at rvv-insts (vsetvli bitfield overlap, vsetvl arg count) then hangs in the SBCL debugger.
-
00:38:00 (04:11) User advises non-interactive builds; build #1 killed.
-
00:39:00 ~ Fix vsetvli/vsetivli emitters byte 10 20 and vsetvl arg count; discover cached output/build-config xc-host gotcha.
-
00:41:00 (04:20) Restart build #2 — silently SIGTTIN-stopped reading the terminal.
-
00:44:00 ~ Diagnose via /proc wchan (do_signal_stop); kill tree (one pkill self-kill accident); rm build-config; restart detached.
-
00:47:00 ~ Build #3 (--disable-debugger, stdin=/dev/null): rvv-insts.lisp compiles; aborts on STYLE-WARNING "undefined function EMIT-RVV-MEM" (warnings are fatal in make-host-1).
-
00:50:00 ~ Add emit-rvv-mem wrapper; restart build #4.
-
00:53:00 (04:13) Build #4 fails: "Comma not inside a backquote" reader error (line 265); user says continue.
-
00:54:00 ~ Debug reader error: hexdump + paren/string scanner + sbcl read-script pinpoint a missing closing quote in a generator printer template (line 267).
-
00:59:00 ~ Repair generator quoting (three misfired attempts, incl. duplicated patch code); regenerate; reader-clean confirmed.
-
01:02:00 (04:35) Restart build #5: clean host compile; make-host-2 cold-init fails: pd-error "unknown argument RD" (xd/fd/vs/vf printers on the rvv-un format).
-
01:06:00 ~ Add four dedicated formats (rvv-xd/fd/vs/vf); add vector.lisp to build-order; restart build #6.
-
01:10:00 (04:18) Build #6 fails in vector.lisp: "operand type T" — VOPs need explicit :arg-types/:result-types; user says retry.
-
01:12:00 ~ Add :arg-types/:result-types to the chunk VOP macro.
-
01:15:00 (04:58) Build #7 fails at load: defknowns invisible from the VOP file — move them to generic/vm-fndb.lisp; drop vector-sap wrappers (cross-compiler type clash).
-
01:17:00 (05:05) Build #8 fails at vector.fasl load: "%RVV-F32-ADD-CHUNK is not a known function"; user notes builds take ~600 s.
-
01:19:00 ~ Diagnose package mismatch: defknown names interned in SB-C vs VOP :translate in SB-VM; find in-tree precedent sb-vm::%vector-cas-pair.
-
01:23:00 (05:13) Apply sb-vm:: prefix to the six defknowns; restart build #9.
-
01:25:00 (05:24) Build #9 still "not a known function"; user says retry.
-
01:27:00 ~ Confirm the names are in vm-fndb.fasl (strings) but registered under SB-C; verify prefix edit landed; restart build #10.
-
01:45:00 ~ Build #10 runs full length; make-host-1/2 and runtime build complete.
-
02:29:00 (06:02) BUILD EXIT: 0 — first fully successful build.
-
02:31:00 ~ First test run: pure fails (dotted test-entry shape); impure: chunk functions not callable at runtime; roundtrip already shows working disassembly and exposes the vmsgt.vv/vmslt.vv printer ambiguity.
-
02:34:00 ~ Add self-calling defun stubs in src/code/riscv-vm.lisp; fix test shapes, test-util:with-test, paren balance, sub-test expectation.
-
02:37:00 ~ Convert vmsgt.vv/vmsgtu.vv to alias emitters (swapped operands, no printer); write doc/internals-notes/rvv-support.txt while the next build runs.
-
02:40:00 ~ Restart build #11.
-
02:56:00 ~ Tests: f32/chunk tests PASS on hardware; pure down to one mismatch: VMSBC.VV funct6; roundtrip: vmsleu.vi/vmsltu.vi collision.
-
02:59:00 ~ Fix vmsbc/vsbc funct6 templates (hardcoded vmadc/vadc values); run a systematic constraint-collision scan over the generated file.
-
03:02:00 ~ Scan shows only the two vi aliases remain: convert vmslt.vi/vmsltu.vi to imm-1 alias emitters.
-
03:04:00 ~ Pure test finds vid.v/vmclr.m/vmset.m arg-count and operand-derived-field bugs; fix kinds (vid 1-arg with printer; clrset/vmnot alias emitters).
-
03:07:00 ~ Restart build #12.
-
03:23:00 ~ Tests: pure fails on vsbc.vvm (same funct6 class); fix acv/aci/acx templates; make roundtrip assertions register-name agnostic.
-
03:26:00 ~ Restart build #13.
-
04:24:00 (07:57) ALL TESTS SUCCEED: 572 encodings, functional RVV math on hardware, chunk VOPs at odd lengths, disassembler round-trip.
-
04:25:00 (07:58) Regression subset: assembler.pure, disassem.pure-cload.lisp (after first trying a nonexistent disassem.pure.lisp), disassem.impure, interface.pure — all clean.
-
04:26:00 (07:59) compiler.pure.lisp passes (one pre-existing expected-failure marker).
-
04:28:00 ~ NEWS entry; restore xperfecthash63.lisp-expr build churn; git add.
-
04:30:00 ~ First commit attempt fails (no git identity); set local identity.
-
04:32:00 ~ Commit 1bbaf3c26 "riscv: assembler, disassembler and VOP support for RVV 1.0" (10 files, +5074).
-
04:32:00 (07:59) Session goes idle.
Round 2-4 — benchmark, spec review, deferred-spec implementation
Round 1 was the assembler/disassembler/VOP bring-up above. Rounds 2-4
are recorded here too (they had previously been left only in
rvv-sample-tests.adoc and rvv-todo.adoc). Timestamps below are
absolute wall-clock (PKT, 2026-09-07), reconstructed from commit
author dates and file mtimes.
-
02:58 — Rewrite the RVV instruction source: generate
src/compiler/riscv/rvv-insts.lispfrom a Lisp table instead of the Python generator; regenerate the perfect-hash table; rename docs to.adoc; writetimeline.adoc,work.adoc,buildnotes.adoc,changes.adoc,rvv-implementation.adoc; add the scalar-vs-RVV timing benchmark and gap analysis (commits 23bb5d229, 32c8c25d4, 3aa78d5a2, b91e2d0da, f70b11544, a0a62be06). -
04:17 — Add runtime RVV CPU feature detection (
getauxval(AT_HWCAP)/SB-VM:RVV-SUPPORTED-P); spec 00. -
04:39 — Build-time RVV toolchain autodetection in
make-config.shplus--without-riscv-vector; spec 01. -
05:10 — Parameterize chunk/block VOPs by LMUL (1/2/4); spec 02.
-
05:22 — Generalize block unroll to
:unroll N; spec 03. -
05:57 — Scalar-out reduction VOPs; spec 04.
-
06:19 — Mask-producing compare and mask-consuming merge VOPs; spec 05.
-
06:29 — Fused multiply-add (accumulate) chunk VOPs; spec 06.
-
06:40 — Same-width float↔integer conversion VOPs; spec 07.
-
06:42 — Record remaining specs (08-10) status in the implementation log.
-
12:12 — Benchmark the new VOPs and refresh
rvv-results/results.adoc. -
12:31 — Retire the interpreter/evaluator caveat in
doc/internals-notes/rvv-support.txt(fixed in SBCL 2.6.8). -
12:36 — Mark specs 00-07 done and 08-10 deferred in the spec files and index (commit d004a6be5).
-
13:41 — Spec 08: strided load/store and indexed gather VOPs plus the whole-register move primitive (
vl1r.v/vs1r.v); segment loads deferred (commit edc72da5e). -
16:58 — Spec 09: fixed-128-bit first-class vector values (the SIMD-PACK analogue) — move/box/unbox/construct/extract VOPs,
:sb-simd-packgating, primitive types, register/stack SCs,simd-pack-dispatch; cross-call/GC liveness deferred to spec 10 (commit 0ff5316cc). -
17:08 — Re-run the full benchmark suite on both core types (VLEN=256 and VLEN=1024) plus the spec-08 gather benchmark and the C calibration; refresh
rvv-results/results.adocand prune the superseded status text acrossrvv-todo.adoc,rvv-specs-implementation.adoc,rvv-implementation.adoc,buildnotes.adoc,changes.adoc,work.adoc,rvv-progress.adocandrvv-sample-tests.adoc. -
18:24 — Spec 10 (partial): C-runtime vector-state accessor (
os_context_vstate/os_context_vector_register_addr) reaching the kernel’s magic-tail__riscv_v_ext_stateblock, plus the VLENB probe (riscv_vector_vlenb/SB-VM:RVV-VLENB) andtests/rvv-state.impure.lisp. Callee-saved save/restore (v1-v7 + v24-v31, psABI vector-CC variant) and the interrupt-context CSR round-trip are deferred to the general-VLEN value. -
18:49 — Spec 10 decision revised: the callee-saved set is now v1-v7
v24-v31 (psABI Standard Vector Calling Convention Variant, riscv-elf-psabi-doc PR #389), replacing the v8-v23 draft. The FFI boundary is unaffected (alien/callback calls follow the standard all-caller-saved C convention) and there is no runtime performance benefit to diverging, so we follow gcc/clang. -
19:43 — Dedicated vector register storage base: add a separate
vector-registersstorage base (size 32) insrc/compiler/riscv/vm.lispand move the vector SCs (vector-reg/int-vector-reg/double-vector-reg/single-vector-reg) offfloat-registers, so FP f0-f31 and vector v0-v31 are independent allocator slots (RISC-V V has no F/V aliasing). The chunk/block/unroll/reduce/strided/indexed VOPs insrc/compiler/riscv/vector.lispnow use thevector-regscratch SC instead of the FPdouble-regSC for their wired temporaries;rvv-insts.lispandtarget-insts.lispencode/disassemble vector operands fromvector-registers. Full rebuild green; rvv, rvv-simd, rvv-state and rvv-assembler tests pass. -
20:46 — Integer width-change VOPs (spec 07, the deferred half):
%rvv-u8→u32-chunk(vzext.vf4),%rvv-u16→u32-chunk(vzext.vf2) and%rvv-u32→u16-chunk(vnsrl.wi 0) with defknowns, runtime stubs, correctness tests and benchmark rows. Key fix: a narrowing instruction’s vtype SEW is the destination width, so the VOP programsSEW=e16and loads the e32 source into a 2-register group (v24-v25); the first attempt usedSEW=e32, which makes the source e64 and puts the destination in the high half of the source group — a reserved overlap encoding (SIGILL on the Spacemit X100). Also allowlisted the dead x86-64/arm64%simd-pack-int-to-*references indebug-int.lispso a--disable-debuggercold build completes. Widening benchmarks ≈2.3-2.6x, narrowing ≈4-6x, no regressions in the existing rows. -
21:17 — Float width-change VOPs (spec 07):
%rvv-f32→f64-chunk(vfwcvt.f.f.v) and%rvv-f64→f32-chunk(vfncvt.f.f.w). Widening float conversions follow the widening-arithmetic convention (source EEW=SEW, destination EEW=2*SEW), so the VOP uses SEW=e32 with a 2-register e64 destination group; the first cut used SEW=e64 (a reserved overlap encoding, SIGILL). f32→f64 ≈1.9-2.3x, f64→f32 ≈1.3-2.3x (the large-n f64→f32 row is DRAM-bound), no regressions. -
21:47 — Float FMA subtract/negate accumulate forms (spec 06):
%rvv-f32-fmsac-chunk(vfmsac.vv),%rvv-f32-fnmacc-chunk(vfnmacc.vv) and%rvv-f32-fnmsac-chunk(vfnmsac.vv). RVV FMA naming is non-obvious:vfnmaccnegates the product and subtracts the addend (-(a*b)-c), whilevfnmsacnegates the product and adds (-(a*b)+c); the first test swapped the two. Existing benchmark rows unchanged. -
22:20 — Signed integer width-change VOPs (spec 07):
%rvv-s8→s32-chunk(vsext.vf4) and%rvv-s16→s32-chunk(vsext.vf2), completing the signed-extension half of spec 07 by reusing the widen-VOP shape (vle8.v/vle16.vload the narrow source into the low fraction of a scratch register,vsextsign-extends it to a full e32 lane,vse32.vstores). Adds defknowns, runtime stubs and signed round-trip tests. -
22:27 — Predicated store + find-first (spec 05):
%rvv-u32-store-gt(vmsgtu.vvbuilds the comparison mask in v0, then the,v0.tmaskedvse32.voverwrites only the active lanes, leaving the rest undisturbed) and%rvv-u32-first-gt(compare +vfirst.m, which returns -1 for an all-zero mask and is detected with a signed>= 0branch to yield the first matching lane index, or COUNT when none). Full rebuild green; rvv, rvv-simd, rvv-state and rvv-assembler all pass. -
22:55 — Predicated (masked) load (spec 05, completing the spec):
%rvv-u32-masked-loadimplements the,v0.tmasked-load form and selects thetu/mu(tail/mask undisturbed) policy via(inst vsetvli … nil nil): it preloads the destination fromdst, materialises the u32 mask into v0 withvmsne.vi v0, vmask, 0, thenvle32.v vd, src, v0.t(vm=0) overwrites only the active lanes andvse32.vstores the merged result. This closes the last deferred piece of spec 05. Full rebuild green; all RVV suites pass including the new:rvv-masked-load-vop. -
23:15 — Fractional-LMUL refinement (spec 07, completing the spec): the widening extension VOPs now program LMUL=4 (
vzext.vf4/vsext.vf4) or LMUL=2 (vzext.vf2/vsext.vf2) instead of LMUL=1. Because the vtype SEW for a widening extension is the destination EEW (e32), the narrow source EMUL is LMUL/4 or LMUL/2; raising LMUL makes the source a whole register and grows VLMAX from VLEN/32 to VLEN/8 (vf4) or VLEN/16 (vf2), with a 4- or 2-register e32 destination group. u8→u32 speedup went from ~2.5x to ~6.0-6.5x and u16→u32 from ~2.3x to ~4.4-4.6x. Full rebuild green; all RVV suites pass. -
23:48 — Unit-stride segment loads (spec 08, completing the spec):
%rvv-u32-segment-load-2…-8(vlseg2e32.v…vlseg8e32.v) deinterleave an AoS u32 stream into a field-major SoA destination in onevlseg<nf>e32.vplusnfstores. The VOP keeps the strided-load shape (SAP SAP GPR GPR) for everynfby taking a single destination base plus a per-field byte stride, instead ofnfdestination SAPs (which would not fit the register file atnf=8). The destination group isnfconsecutive scratch registers at v24…v24+nf-1; at LMUL=1 there is no per-field alignment constraint, so oddnf(3/5/7) are valid. Full rebuild green; all RVV suites pass including the new:rvv-segment-load-vopstest; benchmark shows ~1.2-1.4x atnf=6..8(AoS→SoA is a small win because both memory streams are already sequential).
Step Summary (Round 1)
Durations by step type, aggregated from the Round 1 timeline above. Active work spans 03:33–08:05; note that some steps overlap (documentation and design notes were written while builds ran), so the column totals slightly exceed the wall-clock span. Full builds measured ≈16 min; make-host-1-stage failures ≈5–8 min; make-host-2 (cold-init) failures ≈15 min.
| Step type | Total | Notes |
|---|---|---|
Explore/survey |
~10 min |
Repo layout, backend file inventory. |
Study/reading |
~20 min |
insts.lisp, vm.lisp, x86-64/arm64 SIMD, disassem/assem internals. |
Experiment/probe (assembler oracle) |
~25 min |
gas vs clang syntax, register alignment, mask probes, mnemonic/operand-order corrections, vtype decode. |
Write generators (Python) |
~25 min |
rvvgen.py, rvvspec.py, rvvgen2.py, including re-runs after assembler feedback. |
Generate/install Lisp + tests |
~10 min |
rvv-insts.lisp (3999 lines), 572-case test, printers, build-order, vector.lisp. |
Build SBCL (13 builds) |
~2 h 10 m |
4 full-length successes ≈16 min each; 6 early fails ≈6 min; 3 cold-init fails ≈15 min; the dominant cost of the session. |
Debug/diagnose |
~1 h |
Reader errors, SIGTTIN, style-warning fatality, cold-init pd-error, VOP/defknown plumbing, alias/ambiguity hunts, funct6 mismatches. |
Patch/fix (source, generator, tests) |
~35 min |
Emitters, formats, aliases, arg kinds, test shapes, expectations. |
Test runs |
~15 min |
≈8 rvv runs of 1–2 min each + the regression subset. |
Git (identity, 2 commits) |
~5 min |
1bbaf3c26 (code). |
Key economics:
-
Conceptual work that produced every encoding and the design (study + probes + generators) cost ≈80 min — about the length of two build cycles.
-
Each fix validated by a full rebuild cost ≈16 min; the three late encoding-class fixes (vmsgt alias; vmsbc/vsbc funct6; the vi aliases plus vid/vmclr kinds) each consumed one build, and the vsbc fix needed one more.
-
A faster inner loop — reloading just rvv-insts.fasl into an existing core, or the (eventually written) constraint-collision pre-check run before building — would have cut the session to roughly half.
-
Environment friction (debugger hangs, SIGTTIN stops, cached build-config, self-pkill) cost ≈15 min once and nothing after the
--disable-debugger+setsid </dev/nullinvocation became standard.