Hands-on review of the in-browser demo at https://danbri.github.io/factoidal/fstar-extracted/. Every dataset × query combination listed in the dropdowns was executed through Playwright; a handful of combinations were also driven via Chrome DevTools for a performance trace. The wasm and js engines were compared on the same queries.
none, RDFS, OWL-RL on the schema.org dataset.js (js_of_ocaml) and wasm (wasm_of_ocaml).SELECT *, typo'd keyword, malformed Turtle) on the People dataset.tmp/perf-trace-music-propertypath-js.json).Non-ASCII literals are mangled in the rendered results table when the js engine is selected — but not when the wasm engine is selected. The input Turtle textarea always holds the correct bytes; the damage is on the output path.
Latin-script accents (known at first pass):
| Source literal | js-engine output | wasm-engine output |
|---|---|---|
"Eve Müller"@de |
"Eve M����r"@de |
"Eve Müller"@de ✓ |
"Cafetière" |
"Cafeti鳥" |
(correct) |
"Ralf Hütter" |
"Ralf H����r" |
(correct) |
"Gödel, Escher, Bach" |
"G����, Escher, Bach" |
(correct) |
RTL + mixed-script + emoji stress (added 2026-04-21):
| Source literal | js-engine output | wasm |
|---|---|---|
"שלום עולם"@he (Hebrew) |
"ёӓՕ"@he |
✓ |
"مرحبا بالعالم"@ar (Arabic) |
"E1-(' ('D9'DE"@ar |
✓ |
"Hello שלום world"@en (LTR+RTL mix) |
"Hello 靕ݠworld"@en |
✓ |
"混合 español عربي"@mul (CJK+Latin+Arabic) |
"����spa91(J"@mul |
✓ |
"كافيه Café ☕"@mul (Arabic+accent+emoji) |
"C'AJG Caf頕"@mul |
✓ |
"אבגדהו"@he (Hebrew) |
missing — not in result set | ✓ |
On the mixed-script dataset the js engine returned 5 rows when 6 were expected. So the corruption is not merely display-layer; a row was silently dropped from the result set. Most likely cause: the ORDER BY comparator or a DISTINCT-like collapse hits invalid UTF-8 during comparison and either errors out of a row quietly or treats two rows as equal. Either way this is a correctness bug, not just a rendering glitch.
Because the wasm build produces correct output and correct cardinality from
the same F* sources, the bug lives in the js extraction / JS-side string
handling, not in .fst logic. Most likely: a byte-length-indexed slice or
String.sub path is producing invalid UTF-8 fragments that the DOM then
renders as U+FFFD replacement characters, and a comparator on those bytes
sometimes produces equality where the source strings differed. The fact that
UCASE also corrupts in the js engine ("EVE M����R") confirms the bad
bytes flow all the way through string operations.
This is the single biggest demo-facing bug — and now the biggest correctness bug in the js build too.
Even on the wasm engine where UTF-8 is preserved, UCASE("Eve Müller")
returns "EVE MüLLER" — the ü is not uppercased. This matches the known
limitation in CLAUDE.md rule #10 (OCaml Str/byte operations on UTF-8).
Worth calling out in the demo copy or the review doc because users will
notice. A Unicode-aware casing helper — or an assume val boundary to a
Pcre/Re-backed stub — would fix it.
Typing SELEKT for SELECT:
SPARQL parse error: expected SELECT, ASK, CONSTRUCT, or DESCRIBE
SPARQL parse error: Failure("Error: __factoidal_exit__")
The second line is raw OCaml internals and should be filtered out of the UI surface. The first line is fine.
Pasting @prefix ex: <http://example.org/> . ex:a ex:p "unterminated into the
Turtle textarea and running a query returns Done (js) / 0 results with no
parse error shown. The Turtle parser (or its driver) is swallowing the failure
and handing an empty graph to the evaluator. Should surface as an error like
the SPARQL parse path does.
ORDER BY ?name on the People dataset interleaves groups by literal kind
before alphabetising within each group:
Bob Nguyen @en
Dave Patel @en
Eve M����r @de ← language-tagged literals first
Frank O'Connor @en
Heidi Klum @de
Alice Carter ← then plain xsd:string literals
Carol Diaz
Grace Hopper
SPARQL lets implementations choose the order across incomparable literal
types, so this is technically spec-legal, but it makes the demo look
buggy: users expect Alice to come before Bob. A note in the demo, or an
implementation tweak that compares lexical form across the two groups,
would help.
AVG on the books category returns "31.666666666666"^^xsd:decimal (13 sixes,
truncated) instead of 31.666666666667 (rounded) or a fully-preserved
repeating representation. Minor but off-spec if the intent was correctly
rounded decimal.
The TriG dataset declares @prefix ex: <http://example.org/lib/> and uses
ex:central. In the results table the IRI renders as ex:lib/central,
which is what you'd get if the output formatter were using a different
prefix table (ex: → http://example.org/) than the input. Display-only
cosmetic, but confusing.
SELECT * column order (LOW)#SELECT * WHERE { ?s foaf:name ?n } produces ?n ?s — reverse of the
textual order. SPARQL 1.1 §16.2 says * projects all in-scope variables
but doesn't pin column order; many engines sort by first-appearance. Not
a bug, but worth noting that the order is stable / deterministic but not
aligned with user expectation.
People & friendships:
Product catalogue:
Music:
Places:
TriG libraries:
schema.org hierarchy:
a/rdfs:subClassOf* with entailment=none → 4 rows ✓Measured wall-clock from button-click to Done text, using performance.now()
in the page; 10 ms polling granularity. Small queries bottom out around 15 ms
so treat sub-30-ms numbers as noise.
| Test | js (ms) | wasm (ms) | wasm speedup |
|---|---|---|---|
| schemaOrg All Places, OWL-RL | 43.7 | 27.0 | 1.6× |
| schemaOrg All Places, RDFS | 30.1 | 27.7 | 1.1× |
| schemaOrg All Places, none | 21.2 | 15.7 | 1.4× |
| music path-performers, none | 62.2 | 25.6 | 2.4× |
| people count-friends, none | 19.2 | 21.0 | 0.9× |
| libraries by-library, none | 28.6 | 15.9 | 1.8× |
wasm is 1.4–2.4× faster than js on the heavier queries; on tiny queries they tie. Both are well under the 300 ms latency ceiling any user would notice.
Saved to tmp/perf-trace-music-propertypath-js.json (music property-path
query under js engine). Total trace span ≈ 24 s including idle time before and
after the click. The actual query took ~60 ms and fit in one frame; there
was no layout shift or long task worth analysing at this data scale. The demo
data is simply too small to stress the engine — to produce a meaningful
trace we'd need a synthetic dataset in the thousands-of-triples range and,
ideally, a slow query (rule #20 — send that to a subagent, not the main
loop). Noted as follow-up.
Clean on both engines. The only console error across all runs was
Failed to load resource: the server responded with a status of 404 () for
/favicon.ico — unrelated to the evaluator, easy to add a favicon or a
1-byte stub.
Ranked by demo impact:
String.sub / byte-index slice in the result formatter.
Trace where a literal's lexical form is converted for DOM insertion in
the js build and compare with the wasm build's path.Failure("... __factoidal_exit__") line in the
status area (finding #3).ex:central not ex:lib/central.Run query button.document.querySelector('main').innerText in the page.