Companion to
2026-04-19-hdt-fstar-status.md,
which audits what's in F* today. This note describes what it takes to
actually ship the ballyhoo binary-format stack — specifically the
F*-verified Parquet parser and the COTTAS path that rides on top of
it — across native + js_of_ocaml + wasm_of_ocaml.
./build-ocaml.sh on claude/main produces a
factoidal / w3c_runner that can --data-cottas file.parquet
(or equivalent) and run SPARQL queries over the resulting quads.docs/fstar-extracted/factoidal.js can load a Parquet
file (via Fetch / HTTP Range) and query it in the browser.build-ocaml.sh wasm path grows the
same capability, with a Zstd primitive available to the wasm binary.hdtSearch is explicitly not shipped to browser/wasm
targets (it spawns subprocesses). Browser story is COTTAS-via-Parquet.
Native can keep HDT optional-if-hdtSearch-is-on-PATH.Out of scope for this plan:
Parquet.Footer.fst.What exists on claude/main (commit 8d4fa67):
Parquet.Footer.fst (1453 lines), five
Parser.Ballyhoo*.fst, SPARQL11.Store.fst.experimental_ocaml_glue/{ballyhoo_hdt_runtime.sh, cottas_runtime.sh, parquet_footer_runtime.sh, parquet_zstd_stubs.c}..ml in ocaml-output/ (mtime predates the
.fst sources; won't match a fresh extract).What's missing on claude/main but exists on
origin/codex/ballyhoo-baseline:
build-ocaml.sh extract-list entries for the new .fst files.build-ocaml.sh COMMON_MODULES entries for the new .ml files.parquet_zstd_stubs.c compile step.So the cherry-pick did source but not wiring. Plan is to finish the job.
These estimates assume coding-agent-paced work (read-many-files-in- parallel, cherry-pick hunks mechanically, run builds in background, iterate on failures while other things continue). Human-paced would multiply by ~3×. Also: the mechanical cherry-picks are fast and reliable; the Zstd integration paths have genuine unknowns that could eat a day if the obvious option turns out not to work.
Goal: ./build-ocaml.sh on claude/main produces a factoidal
that can open a COTTAS Parquet file and return quads.
Steps:
build-ocaml.sh hunks from
origin/codex/ballyhoo-baseline:
for fst in …: add Parquet.Footer.fst,
Parser.Ballyhoo.fst, Parser.BallyhooBloom.fst,
Parser.BallyhooHDT.fst, Parser.BallyhooHDTQ.fst,
Parser.BallyhooCOTTAS.fst, SPARQL11.Store.fst. Order
matters — RDF.Graph.Executable.fst first, then
Parquet.Footer.fst, then the Ballyhoo modules (HDT before
COTTAS — COTTAS runtime glue calls Parquet_Footer), then
SPARQL11.Algebra.fst, then SPARQL11.Store.fst.COMMON_MODULES: mirror the ordering.ocamlfind ocamlopt call: add -package zstd (or
-cclib -lzstd + the C stub .o).experimental_ocaml_glue/parquet_zstd_stubs.c into a
.o inside build-ocaml.sh compile step, link alongside the
.cmx files. ocamlfind ocamlopt with -cc cc -cclib -lzstd
and including parquet_zstd_stubs.o as a plain object.ocaml-patches.sh runs the three ballyhoo glue scripts
in the right order: ballyhoo_hdt_runtime.sh →
cottas_runtime.sh → parquet_footer_runtime.sh. COTTAS glue
calls into Parquet_Footer, so the Parquet patch should land
first; re-read both scripts to confirm the direction.zstd (the zstd opam package is the bindings to
libzstd; also needs the libzstd-dev system package on
Debian/macOS-brew — already present on most CI setups).examples/berlin_hdt_*.sh script that uses a
COTTAS artifact, runs end-to-end against a small fixture,
returns the right row count.bin/darwin-arm64/ per CLAUDE.md rule 9.Estimate: 30 min–2 hours with an agent. The only thing that can go wrong is zstd linking (opam version / system library mismatch). Fallback: declare the C stub compilation optional and gate Parquet support on it, so if libzstd is missing the build still succeeds without Parquet rather than failing.
Exit criterion: a w3c_runner / factoidal that can answer BGPs over a small COTTAS/Parquet file, with the same SPARQL pass rate as the pre-ballyhoo build plus a new test or two exercising the new path.
Goal: docs/fstar-extracted/factoidal.js (and the npm package
built from it) can open a Parquet file and run SPARQL over it.
Steps:
fzstd (BSD, ~20 KB,
active), zstddec (LGPL, might be heavier), numcodecs (big,
overkill). Default pick: fzstd.caml_parquet_zstd_decompress_hex that mirrors the C stub's
contract: take hex string + expected size, return decompressed
hex or None. The primitive lives as a //Provides: block in
a .js runtime file referenced from the build-ocaml.sh js
step (pattern established by existing fstar_int_stubs.js).
Implementation: hex→Uint8Array, call fzstd.decompress,
Uint8Array→hex.parquet_read_tail_hex / _range_hex
currently do open_in_bin + seek_in. In js_of_ocaml this
works for the virtual FS (files mounted via
Sys_js.update_file), so local-file factoidal-js use cases
work as-is. For remote Parquet URLs, add a separate
parquet_fetch_range_hex primitive using browser Fetch with
Range: headers. Wire it as an alternate path in the
parquet_footer_runtime.sh-equivalent for js targets.fzstd into the factoidal-js build. Either inline its
source (10–20 KB) or load it via ES module import, depending
on how the existing js build packages third-party JS.docs/fstar-extracted/index.html is the demo) plus a
new code path that fetches a small Parquet file and runs a
SPARQL query over it.factoidal.js. Expected: +30–60 KB
for fzstd + glue. If unacceptable, lazy-load fzstd only when
a Parquet path is actually used.Estimate: 2–6 hours with an agent. The bottleneck is step 2/3
if fzstd's hex↔bytes conversion is awkward or if the Bytes vs
String distinction in js_of_ocaml surfaces a snag. Step 4
(bundling) can be slow if the existing build pipeline is picky
about new dependencies.
Known potential snag: Parquet.Footer.fst is nat-heavy in
its varint / zigzag loops (same issue as Turtle parser — see
2026-04-19-turtle-parser-speed.md). In JS this extracts to
zarith_stubs_js bignum ops. Functionally correct but slow —
a few-MB Parquet file may take seconds to parse in the browser.
Not a blocker for Phase 2, but track it as a follow-on; the same
pos_t migration that helps Turtle would help here.
Exit criterion: a browser demo that loads a small Parquet file (say, 100 KB) and runs a SELECT over it without timing out.
Goal: ./build-ocaml.sh wasm produces a wasm bundle that
includes Parquet/COTTAS capability.
Per CLAUDE.md, the wasm path is already partially working — SPARQL suites like bind/bindings/aggregates pass identically to native. The gaps are SHA/MD5 stubs (separate issue) and now Zstd.
Two options for Zstd in wasm:
wasm_of_ocaml
supports external primitives via .wat + JS shim pairs (see
CLAUDE.md mention of janestreet/zarith_stubs_js's
(wasm_of_ocaml (wasm_files runtime_wasm.js runtime.wat))).
The caml_parquet_zstd_decompress_hex primitive gets a
.wat stub that tail-calls into a JS function that runs
fzstd. Cheap, identical semantics to Phase 2.Steps (assuming Option A):
wasm_runtime/parquet_zstd_stubs.wat that exports a
function matching the caml_parquet_zstd_decompress_hex
contract, forwarding to JS.wasm_runtime/parquet_zstd_stubs.js providing the JS
side (same fzstd call as Phase 2; shared if layered right).build-ocaml.sh wasm's linker flags,
next to the existing Zarith runtime wiring.Parser_BallyhooHDT.ml patched with Unix.open_process_full
— that runtime dependency can't be satisfied in a browser.
Two approaches:
COMMON_MODULES split: COMMON_MODULES_NATIVE (with HDT)
vs COMMON_MODULES_BROWSER (without).hdt_search that returns
empty results and logs a warning. Uglier but keeps one
module list.
Recommend (a); the code split follows a clean boundary.Estimate: 4 hours–1 day with an agent, assuming Option A works. Option B adds a day for emscripten integration.
Known snag: the hex-string memory ceiling. A 10 MB
compressed Parquet page becomes a 20 MB hex input to the Zstd
stub, and the decompressed output becomes a 2× larger hex string
back. In wasm's 32-bit address space (4 GB hard cap, browser
limits typically ~2 GB), this ceilings out at roughly 500 MB
decompressed data per call. Fine for COTTAS files up to that
size; catastrophic beyond. Worth measuring before promising
anything to users. Real fix would be chunked
caml_parquet_zstd_decompress_hex_streaming, which is a bigger
refactor of the F*–C boundary.
Exit criterion: ./build-ocaml.sh wasm produces a bundle
that answers the same SPARQL query over the same small Parquet
file as Phase 2, running under wasm_of_ocaml.
tests/ — a tiny Parquet
file (say, 10 quads) committed as binary under
tests/fixtures/cottas/. Document how it was generated
(which emitter, encoding settings).libzstd-dev (or
install it in the workflow).docs/designissues/cottas-native-backend.md — drop the
"no implementation yet" caveats, cross-link
Parquet.Footer.fst.docs/designissues/2026-04-19-hdt-fstar-status.md — remove
the "build-wiring caveat" section once Phase 1 lands.Estimate: 1–2 hours.
Realistic plan-to-ship total: one agent-day, two at most if the Zstd-in-wasm path needs Option B or the hex-ceiling requires a chunked primitive. Single-day outcomes are plausible if Phase 3 uses Option A without surprises.
The big unknowns (in order of likelihood to bite):
zstd binding compatibility with the pinned F*
opam switch (Phase 1). If the binding's OCaml version
requirements clash with ocaml-base-compiler.4.14.1,
fall back to a lower-level Ctypes binding or a hand-rolled
FFI. Adds ~2 hours.TextDecoder/String.fromCharCode tricks is too slow in
practice, we need to bypass hex and pass Uint8Array
directly, which means changing the F* side to accept
FStar.Bytes rather than string — a real refactor of
Parquet.Footer.fst. Adds ~1 day. Defer unless measurement
forces it..wat →
JS-primitive wiring has some undocumented snag specific to
wasm_of_ocaml's current version, could eat a few hours
debugging the bridge. Fallback is Option B (emscripten),
which costs ~1 day but avoids the JS↔wasm bridge entirely.Parquet.Footer.fst, not build-plumbing work.--data-parquet as a generic flag. The plan wires
COTTAS-shaped Parquet (4 columns: subject, predicate,
object, graph) into the Store layer. Generic Parquet-to-RDF
mapping (arbitrary schemas, typed columns) is a separate
project.Candidate GitHub issue titles if/when this moves:
Recommend opening them only once Phase 1 is in flight, so the concrete shape of the commit informs the later issue text.