Elapsed times are hh:mm:ss since task start (wall clock 03:33 PKT, 2026-09-06). Items marked ~ are reconstructed from build logs, file mtimes and turn timestamps (minute-level accuracy); everything else is anchored to recorded times. One line per item.

Timeline

  • 00:00:00 (03:33) Task received: add RVV 1.0 (vector) support to SBCL’s riscv64gc backend; free rein on branch rvv.

  • 00:01:00 ~ Survey repo tree: src/compiler/riscv (24 files), src/assembly/riscv, src/code/riscv-vm.lisp, src/runtime, src/cold.

  • 00:03:00 ~ Read src/compiler/riscv/insts.lisp: formats, printers, bitfield emitters, arg types, inst/inst*.

  • 00:05:00 ~ Read vm.lisp (storage bases/classes), c-call.lisp (C ABI), assem-rtns.lisp (save/restore-c-registers).

  • 00:07:00 ~ Study x86-64 SIMD precedent: avx2-insts.lisp, simd-pack*.lisp, cpu feature detection; arm64-simd.lisp.

  • 00:08:30 ~ Study disassembler internals: printer compilation, :cond/:unless, :constant test, instruction-flavors.

  • 00:10:00 ~ Probe host: /proc/cpuinfo shows v/zvbb; clang 21 + binutils 2.46 accept -march=rv64gcv; native SBCL 2.6.6 installed.

  • 00:12:00 ~ First gas experiments: store syntax vse32.v (a0), v1 rejected by gas; LLVM integrated assembler accepts vd-first form.

  • 00:14:00 ~ More syntax probes: whole-register ops need group-aligned registers; masked (v0.t) acceptability probed per family.

  • 00:16:00 ~ Write /tmp/rvvgen.py: mnemonic table → .s → clang-assemble → objdump → verified TSV of 32-bit encodings.

  • 00:18:00 ~ Iterate generator against assembler errors: .wv/.wi/.wx suffixes, vwmacc (not vwsmacc), unary .m forms, vfsqrt.v, vwsll is Zvbb, vmacc operand order.

  • 00:21:00 ~ Add missing families: vaadd/vaaddu/vasub/vasubu, vmin/minu/max/maxu, vrgather.vi/vx; final table 585+ rows.

  • 00:23:00 ~ Decode fields of every word; discover pseudo-ops: vmsgt.vv/vmsgtu.vv (swapped vmslt), vmclr/vmset/vmnot, vmerge/vmv vm-bit roles.

  • 00:25:00 ~ Verify vtype layout from 0x0d0372d7: lmul m1=000..mf2=110; sew 5:3; ta/ma 6/7; vset* distinguished by bits 31:30.

  • 00:27:00 ~ Write /tmp/rvvspec.py: classify each mnemonic into operand kinds (vv/vi/vx/vf/acc/un/wmv/mem/…​).

  • 00:29:00 ~ Write /tmp/rvvgen2.py: emit src/compiler/riscv/rvv-insts.lisp (formats + emitters + ~560 define-instructions).

  • 00:32:00 ~ Emit tests/rvv-assembler.pure.lisp (572 encode-vs-gas cases).

  • 00:33:30 ~ Install: vector printers in target-insts.lisp; build-order entry #+riscv {arch}/rvv-insts.

  • 00:35:00 ~ Write src/compiler/riscv/vector.lisp v1: chunk VOPs + defknowns + convenience wrappers.

  • 00:36:30 ~ Start build #1; it dies at rvv-insts (vsetvli bitfield overlap, vsetvl arg count) then hangs in the SBCL debugger.

  • 00:38:00 (04:11) User advises non-interactive builds; build #1 killed.

  • 00:39:00 ~ Fix vsetvli/vsetivli emitters byte 10 20 and vsetvl arg count; discover cached output/build-config xc-host gotcha.

  • 00:41:00 (04:20) Restart build #2 — silently SIGTTIN-stopped reading the terminal.

  • 00:44:00 ~ Diagnose via /proc wchan (do_signal_stop); kill tree (one pkill self-kill accident); rm build-config; restart detached.

  • 00:47:00 ~ Build #3 (--disable-debugger, stdin=/dev/null): rvv-insts.lisp compiles; aborts on STYLE-WARNING "undefined function EMIT-RVV-MEM" (warnings are fatal in make-host-1).

  • 00:50:00 ~ Add emit-rvv-mem wrapper; restart build #4.

  • 00:53:00 (04:13) Build #4 fails: "Comma not inside a backquote" reader error (line 265); user says continue.

  • 00:54:00 ~ Debug reader error: hexdump + paren/string scanner + sbcl read-script pinpoint a missing closing quote in a generator printer template (line 267).

  • 00:59:00 ~ Repair generator quoting (three misfired attempts, incl. duplicated patch code); regenerate; reader-clean confirmed.

  • 01:02:00 (04:35) Restart build #5: clean host compile; make-host-2 cold-init fails: pd-error "unknown argument RD" (xd/fd/vs/vf printers on the rvv-un format).

  • 01:06:00 ~ Add four dedicated formats (rvv-xd/fd/vs/vf); add vector.lisp to build-order; restart build #6.

  • 01:10:00 (04:18) Build #6 fails in vector.lisp: "operand type T" — VOPs need explicit :arg-types/:result-types; user says retry.

  • 01:12:00 ~ Add :arg-types/:result-types to the chunk VOP macro.

  • 01:15:00 (04:58) Build #7 fails at load: defknowns invisible from the VOP file — move them to generic/vm-fndb.lisp; drop vector-sap wrappers (cross-compiler type clash).

  • 01:17:00 (05:05) Build #8 fails at vector.fasl load: "%RVV-F32-ADD-CHUNK is not a known function"; user notes builds take ~600 s.

  • 01:19:00 ~ Diagnose package mismatch: defknown names interned in SB-C vs VOP :translate in SB-VM; find in-tree precedent sb-vm::%vector-cas-pair.

  • 01:23:00 (05:13) Apply sb-vm:: prefix to the six defknowns; restart build #9.

  • 01:25:00 (05:24) Build #9 still "not a known function"; user says retry.

  • 01:27:00 ~ Confirm the names are in vm-fndb.fasl (strings) but registered under SB-C; verify prefix edit landed; restart build #10.

  • 01:45:00 ~ Build #10 runs full length; make-host-1/2 and runtime build complete.

  • 02:29:00 (06:02) BUILD EXIT: 0 — first fully successful build.

  • 02:31:00 ~ First test run: pure fails (dotted test-entry shape); impure: chunk functions not callable at runtime; roundtrip already shows working disassembly and exposes the vmsgt.vv/vmslt.vv printer ambiguity.

  • 02:34:00 ~ Add self-calling defun stubs in src/code/riscv-vm.lisp; fix test shapes, test-util:with-test, paren balance, sub-test expectation.

  • 02:37:00 ~ Convert vmsgt.vv/vmsgtu.vv to alias emitters (swapped operands, no printer); write doc/internals-notes/rvv-support.txt while the next build runs.

  • 02:40:00 ~ Restart build #11.

  • 02:56:00 ~ Tests: f32/chunk tests PASS on hardware; pure down to one mismatch: VMSBC.VV funct6; roundtrip: vmsleu.vi/vmsltu.vi collision.

  • 02:59:00 ~ Fix vmsbc/vsbc funct6 templates (hardcoded vmadc/vadc values); run a systematic constraint-collision scan over the generated file.

  • 03:02:00 ~ Scan shows only the two vi aliases remain: convert vmslt.vi/vmsltu.vi to imm-1 alias emitters.

  • 03:04:00 ~ Pure test finds vid.v/vmclr.m/vmset.m arg-count and operand-derived-field bugs; fix kinds (vid 1-arg with printer; clrset/vmnot alias emitters).

  • 03:07:00 ~ Restart build #12.

  • 03:23:00 ~ Tests: pure fails on vsbc.vvm (same funct6 class); fix acv/aci/acx templates; make roundtrip assertions register-name agnostic.

  • 03:26:00 ~ Restart build #13.

  • 04:24:00 (07:57) ALL TESTS SUCCEED: 572 encodings, functional RVV math on hardware, chunk VOPs at odd lengths, disassembler round-trip.

  • 04:25:00 (07:58) Regression subset: assembler.pure, disassem.pure-cload.lisp (after first trying a nonexistent disassem.pure.lisp), disassem.impure, interface.pure — all clean.

  • 04:26:00 (07:59) compiler.pure.lisp passes (one pre-existing expected-failure marker).

  • 04:28:00 ~ NEWS entry; restore xperfecthash63.lisp-expr build churn; git add.

  • 04:30:00 ~ First commit attempt fails (no git identity); set local identity.

  • 04:32:00 ~ Commit 1bbaf3c26 "riscv: assembler, disassembler and VOP support for RVV 1.0" (10 files, +5074).

  • 04:32:00 (07:59) Session goes idle.

Round 2-4 — benchmark, spec review, deferred-spec implementation

Round 1 was the assembler/disassembler/VOP bring-up above. Rounds 2-4 are recorded here too (they had previously been left only in rvv-sample-tests.adoc and rvv-todo.adoc). Timestamps below are absolute wall-clock (PKT, 2026-09-07), reconstructed from commit author dates and file mtimes.

  • 02:58 — Rewrite the RVV instruction source: generate src/compiler/riscv/rvv-insts.lisp from a Lisp table instead of the Python generator; regenerate the perfect-hash table; rename docs to .adoc; write timeline.adoc, work.adoc, buildnotes.adoc, changes.adoc, rvv-implementation.adoc; add the scalar-vs-RVV timing benchmark and gap analysis (commits 23bb5d229, 32c8c25d4, 3aa78d5a2, b91e2d0da, f70b11544, a0a62be06).

  • 04:17 — Add runtime RVV CPU feature detection (getauxval(AT_HWCAP) / SB-VM:RVV-SUPPORTED-P); spec 00.

  • 04:39 — Build-time RVV toolchain autodetection in make-config.sh plus --without-riscv-vector; spec 01.

  • 05:10 — Parameterize chunk/block VOPs by LMUL (1/2/4); spec 02.

  • 05:22 — Generalize block unroll to :unroll N; spec 03.

  • 05:57 — Scalar-out reduction VOPs; spec 04.

  • 06:19 — Mask-producing compare and mask-consuming merge VOPs; spec 05.

  • 06:29 — Fused multiply-add (accumulate) chunk VOPs; spec 06.

  • 06:40 — Same-width float↔integer conversion VOPs; spec 07.

  • 06:42 — Record remaining specs (08-10) status in the implementation log.

  • 12:12 — Benchmark the new VOPs and refresh rvv-results/results.adoc.

  • 12:31 — Retire the interpreter/evaluator caveat in doc/internals-notes/rvv-support.txt (fixed in SBCL 2.6.8).

  • 12:36 — Mark specs 00-07 done and 08-10 deferred in the spec files and index (commit d004a6be5).

  • 13:41 — Spec 08: strided load/store and indexed gather VOPs plus the whole-register move primitive (vl1r.v/vs1r.v); segment loads deferred (commit edc72da5e).

  • 16:58 — Spec 09: fixed-128-bit first-class vector values (the SIMD-PACK analogue) — move/box/unbox/construct/extract VOPs, :sb-simd-pack gating, primitive types, register/stack SCs, simd-pack-dispatch; cross-call/GC liveness deferred to spec 10 (commit 0ff5316cc).

  • 17:08 — Re-run the full benchmark suite on both core types (VLEN=256 and VLEN=1024) plus the spec-08 gather benchmark and the C calibration; refresh rvv-results/results.adoc and prune the superseded status text across rvv-todo.adoc, rvv-specs-implementation.adoc, rvv-implementation.adoc, buildnotes.adoc, changes.adoc, work.adoc, rvv-progress.adoc and rvv-sample-tests.adoc.

  • 18:24 — Spec 10 (partial): C-runtime vector-state accessor (os_context_vstate / os_context_vector_register_addr) reaching the kernel’s magic-tail __riscv_v_ext_state block, plus the VLENB probe (riscv_vector_vlenb / SB-VM:RVV-VLENB) and tests/rvv-state.impure.lisp. Callee-saved save/restore (v1-v7 + v24-v31, psABI vector-CC variant) and the interrupt-context CSR round-trip are deferred to the general-VLEN value.

  • 18:49 — Spec 10 decision revised: the callee-saved set is now v1-v7
    v24-v31 (psABI Standard Vector Calling Convention Variant, riscv-elf-psabi-doc PR #389), replacing the v8-v23 draft. The FFI boundary is unaffected (alien/callback calls follow the standard all-caller-saved C convention) and there is no runtime performance benefit to diverging, so we follow gcc/clang.

  • 19:43 — Dedicated vector register storage base: add a separate vector-registers storage base (size 32) in src/compiler/riscv/vm.lisp and move the vector SCs (vector-reg/int-vector-reg/double-vector-reg/ single-vector-reg) off float-registers, so FP f0-f31 and vector v0-v31 are independent allocator slots (RISC-V V has no F/V aliasing). The chunk/block/unroll/reduce/strided/indexed VOPs in src/compiler/riscv/vector.lisp now use the vector-reg scratch SC instead of the FP double-reg SC for their wired temporaries; rvv-insts.lisp and target-insts.lisp encode/disassemble vector operands from vector-registers. Full rebuild green; rvv, rvv-simd, rvv-state and rvv-assembler tests pass.

  • 20:46 — Integer width-change VOPs (spec 07, the deferred half): %rvv-u8→u32-chunk (vzext.vf4), %rvv-u16→u32-chunk (vzext.vf2) and %rvv-u32→u16-chunk (vnsrl.wi 0) with defknowns, runtime stubs, correctness tests and benchmark rows. Key fix: a narrowing instruction’s vtype SEW is the destination width, so the VOP programs SEW=e16 and loads the e32 source into a 2-register group (v24-v25); the first attempt used SEW=e32, which makes the source e64 and puts the destination in the high half of the source group — a reserved overlap encoding (SIGILL on the Spacemit X100). Also allowlisted the dead x86-64/arm64 %simd-pack-int-to-* references in debug-int.lisp so a --disable-debugger cold build completes. Widening benchmarks ≈2.3-2.6x, narrowing ≈4-6x, no regressions in the existing rows.

  • 21:17 — Float width-change VOPs (spec 07): %rvv-f32→f64-chunk (vfwcvt.f.f.v) and %rvv-f64→f32-chunk (vfncvt.f.f.w). Widening float conversions follow the widening-arithmetic convention (source EEW=SEW, destination EEW=2*SEW), so the VOP uses SEW=e32 with a 2-register e64 destination group; the first cut used SEW=e64 (a reserved overlap encoding, SIGILL). f32→f64 ≈1.9-2.3x, f64→f32 ≈1.3-2.3x (the large-n f64→f32 row is DRAM-bound), no regressions.

  • 21:47 — Float FMA subtract/negate accumulate forms (spec 06): %rvv-f32-fmsac-chunk (vfmsac.vv), %rvv-f32-fnmacc-chunk (vfnmacc.vv) and %rvv-f32-fnmsac-chunk (vfnmsac.vv). RVV FMA naming is non-obvious: vfnmacc negates the product and subtracts the addend (-(a*b)-c), while vfnmsac negates the product and adds (-(a*b)+c); the first test swapped the two. Existing benchmark rows unchanged.

  • 22:20 — Signed integer width-change VOPs (spec 07): %rvv-s8→s32-chunk (vsext.vf4) and %rvv-s16→s32-chunk (vsext.vf2), completing the signed-extension half of spec 07 by reusing the widen-VOP shape (vle8.v/vle16.v load the narrow source into the low fraction of a scratch register, vsext sign-extends it to a full e32 lane, vse32.v stores). Adds defknowns, runtime stubs and signed round-trip tests.

  • 22:27 — Predicated store + find-first (spec 05): %rvv-u32-store-gt (vmsgtu.vv builds the comparison mask in v0, then the ,v0.t masked vse32.v overwrites only the active lanes, leaving the rest undisturbed) and %rvv-u32-first-gt (compare + vfirst.m, which returns -1 for an all-zero mask and is detected with a signed >= 0 branch to yield the first matching lane index, or COUNT when none). Full rebuild green; rvv, rvv-simd, rvv-state and rvv-assembler all pass.

  • 22:55 — Predicated (masked) load (spec 05, completing the spec): %rvv-u32-masked-load implements the ,v0.t masked-load form and selects the tu/mu (tail/mask undisturbed) policy via (inst vsetvli …​ nil nil): it preloads the destination from dst, materialises the u32 mask into v0 with vmsne.vi v0, vmask, 0, then vle32.v vd, src, v0.t (vm=0) overwrites only the active lanes and vse32.v stores the merged result. This closes the last deferred piece of spec 05. Full rebuild green; all RVV suites pass including the new :rvv-masked-load-vop.

  • 23:15 — Fractional-LMUL refinement (spec 07, completing the spec): the widening extension VOPs now program LMUL=4 (vzext.vf4/vsext.vf4) or LMUL=2 (vzext.vf2/vsext.vf2) instead of LMUL=1. Because the vtype SEW for a widening extension is the destination EEW (e32), the narrow source EMUL is LMUL/4 or LMUL/2; raising LMUL makes the source a whole register and grows VLMAX from VLEN/32 to VLEN/8 (vf4) or VLEN/16 (vf2), with a 4- or 2-register e32 destination group. u8→u32 speedup went from ~2.5x to ~6.0-6.5x and u16→u32 from ~2.3x to ~4.4-4.6x. Full rebuild green; all RVV suites pass.

  • 23:48 — Unit-stride segment loads (spec 08, completing the spec): %rvv-u32-segment-load-2…-8 (vlseg2e32.v…vlseg8e32.v) deinterleave an AoS u32 stream into a field-major SoA destination in one vlseg<nf>e32.v plus nf stores. The VOP keeps the strided-load shape (SAP SAP GPR GPR) for every nf by taking a single destination base plus a per-field byte stride, instead of nf destination SAPs (which would not fit the register file at nf=8). The destination group is nf consecutive scratch registers at v24…v24+nf-1; at LMUL=1 there is no per-field alignment constraint, so odd nf (3/5/7) are valid. Full rebuild green; all RVV suites pass including the new :rvv-segment-load-vops test; benchmark shows ~1.2-1.4x at nf=6..8 (AoS→SoA is a small win because both memory streams are already sequential).

Step Summary (Round 1)

Durations by step type, aggregated from the Round 1 timeline above. Active work spans 03:33–08:05; note that some steps overlap (documentation and design notes were written while builds ran), so the column totals slightly exceed the wall-clock span. Full builds measured ≈16 min; make-host-1-stage failures ≈5–8 min; make-host-2 (cold-init) failures ≈15 min.

Step type Total Notes

Explore/survey

~10 min

Repo layout, backend file inventory.

Study/reading

~20 min

insts.lisp, vm.lisp, x86-64/arm64 SIMD, disassem/assem internals.

Experiment/probe (assembler oracle)

~25 min

gas vs clang syntax, register alignment, mask probes, mnemonic/operand-order corrections, vtype decode.

Write generators (Python)

~25 min

rvvgen.py, rvvspec.py, rvvgen2.py, including re-runs after assembler feedback.

Generate/install Lisp + tests

~10 min

rvv-insts.lisp (3999 lines), 572-case test, printers, build-order, vector.lisp.

Build SBCL (13 builds)

~2 h 10 m

4 full-length successes ≈16 min each; 6 early fails ≈6 min; 3 cold-init fails ≈15 min; the dominant cost of the session.

Debug/diagnose

~1 h

Reader errors, SIGTTIN, style-warning fatality, cold-init pd-error, VOP/defknown plumbing, alias/ambiguity hunts, funct6 mismatches.

Patch/fix (source, generator, tests)

~35 min

Emitters, formats, aliases, arg kinds, test shapes, expectations.

Test runs

~15 min

≈8 rvv runs of 1–2 min each + the regression subset.

Git (identity, 2 commits)

~5 min

1bbaf3c26 (code).

Key economics:

  • Conceptual work that produced every encoding and the design (study + probes + generators) cost ≈80 min — about the length of two build cycles.

  • Each fix validated by a full rebuild cost ≈16 min; the three late encoding-class fixes (vmsgt alias; vmsbc/vsbc funct6; the vi aliases plus vid/vmclr kinds) each consumed one build, and the vsbc fix needed one more.

  • A faster inner loop — reloading just rvv-insts.fasl into an existing core, or the (eventually written) constraint-collision pre-check run before building — would have cut the session to roughly half.

  • Environment friction (debugger hangs, SIGTTIN stops, cached build-config, self-pkill) cost ≈15 min once and nothing after the --disable-debugger + setsid </dev/null invocation became standard.