Post 18 named toCottas/openCottas as
the browser's write half of the on-disk COTTAS store, but never called
them from a live cell. Post 24
ran a binary store's reader live, against a real fixture, with timed
queries — but that fixture was 343 triples. This page runs the COTTAS
equivalent at a larger scale: a four-graph corpus, loaded once into the
same in-memory bytes store the CLI's --data-cottas-mem flag serves,
queried three different shapes, and timed against the same queries run
over an ordinary parsed dataset instead.
cottas-corpus.trig holds four
named graphs — people, organizations, projects, and the affiliations
linking them — generated by
gen-cottas-corpus.mjs, a
pure function of an integer index (no randomness, so every id's fields
are computable, not just storable):
| Graph | Records | Triples/record | Triples |
|---|---|---|---|
ex:gPeople |
500 people | 5 (type, name, age, mbox, memberSince) | 2,500 |
ex:gOrgs |
50 orgs | 4 (type, name, location, founded) | 200 |
ex:gProjects |
100 projects | 3 (type, name, budget) | 300 |
ex:gAffiliations |
500 links | 2 (worksFor, contributesTo) | 1,000 |
| Total | 4,000 |
The task behind this page targeted ~100,000 triples. Measured
2026-08-25, against the exact js_of_ocaml engine bundle this page
loads (in-process, no OS subprocess — npm/factoidal/lib/engine-js.js
drives the same bundle browser.js does): parsing a 100,000-triple
corpus through this bundle took over two minutes and did not finish
toCottas() within a further two-minute cap; even 10,000 triples cost
7.5 s to parse and 14.3 s to serialize to COTTAS bytes. That cost comes
from the call path at this corpus size, measured directly — not from
the fixture's file size (the committed .trig stays at 131 KB either
way, far under the 6 MB budget). 4,000 triples keeps every cell
comfortably inside the site's headless-Chromium sweep budget while
still holding four named graphs with real cross-graph links.
fixture = {
const t0 = performance.now();
const text = await fetch("../assets/data/cottas-corpus.trig").then((r) => r.text());
const fetchMs = performance.now() - t0;
return { text, fetchMs, textBytes: text.length };
}
parsed = {
const t0 = performance.now();
const dataset = await fn.parse(fixture.text, { format: "trig" });
const parseMs = performance.now() - t0;
return { dataset, parseMs, tripleCount: dataset.size };
}
parsed.tripleCount should read 4000 — the fixture's own labelled
count, round-tripped through the real F*-extracted TriG parser.
toCottas#fn.toCottas runs the SAME pure Tot serializer
(RDF.CottasStore.BaseWriter.serialize_cottas_v2) the native
factoidal compact --native-writer CLI path calls — not a parallel
browser-only encoder (post 18's finding, restated here at four times
the graph count):
store = {
const t0 = performance.now();
const bytes = await fn.toCottas(parsed.dataset);
const toCottasMs = performance.now() - t0;
const t1 = performance.now();
const handle = await fn.openCottas(bytes);
const openMs = performance.now() - t1;
return {
bytes, handle, toCottasMs, openMs,
storeBytes: bytes.length,
textBytes: fixture.textBytes,
};
}
store.storeBytes compared with store.textBytes is the COTTAS
artifact's size against the .trig text it was built from — a
Parquet-based columnar encoding of a fairly compact, prefix-abbreviated
serialization, so the win here is modest, not the order-of-magnitude
difference the format shows on larger or less-abbreviated corpora.
Three SPARQL 1.1 queries, all against the same four graphs:
GRAPH join — three triple patterns, each pinned to
a different named graph, joined on shared variables (?person's
name from gPeople, their org from gAffiliations, that org's name
from gOrgs).Each runs against store.handle (the COTTAS bytes store) and against
parsed.dataset (the same triples, held as parsed text). The store
side is timed as the median of three runs — labelled below — since a
COTTAS handle's first query on a given predicate pays a one-time lazy
dictionary-populate cost that later queries on the same handle do not;
the table below shows all three per-query times, not only the median,
so that first-call cost stays visible. The in-memory side is timed
once: fn.query() serializes parsed.dataset back to text and
reparses it through the CLI bundle on every single call — a fixed cost
unrelated to which of the three queries runs, so three repeats would
add three times the wall-clock cost without changing the comparison.
queries = ({
point: `# Look up one person's name: one bound subject and predicate,
# in one named graph.
PREFIX ex: <http://example.org/factoidal/> PREFIX foaf: <http://xmlns.com/foaf/0.1/>
SELECT ?name WHERE { GRAPH ex:gPeople { ex:person375 foaf:name ?name } }`,
star: `# List every property/value pair one person has, in one named graph.
PREFIX ex: <http://example.org/factoidal/>
SELECT ?p ?o WHERE { GRAPH ex:gPeople { ex:person375 ?p ?o } }`,
crossGraph: `# Three-graph join: a person's name from one graph, the
# organisation they work for from a second, and that org's name from a third.
PREFIX ex: <http://example.org/factoidal/> PREFIX foaf: <http://xmlns.com/foaf/0.1/> PREFIX org: <http://www.w3.org/ns/org#>
SELECT ?personName ?orgName WHERE {
GRAPH ex:gPeople { ex:person375 foaf:name ?personName }
GRAPH ex:gAffiliations { ex:person375 org:worksFor ?org }
GRAPH ex:gOrgs { ?org foaf:name ?orgName }
}`,
})
storeTimings = {
function median(a) { const s = [...a].sort((x, y) => x - y); return s[Math.floor(s.length / 2)]; }
const labels = { point: "point lookup", star: "star join", crossGraph: "cross-graph GRAPH join" };
const out = [];
for (const [key, sparql] of Object.entries(queries)) {
const times = [];
let rows;
for (let i = 0; i < 3; i++) {
const t0 = performance.now();
rows = await fn.queryCottas(store.handle, sparql);
times.push(performance.now() - t0);
}
out.push({
query: labels[key], medianMs: Math.round(median(times) * 100) / 100,
runs: times.map((t) => Math.round(t * 100) / 100), rows: rows.length,
});
}
return out;
}
storeTimings above — median of 3, your browser, fn.queryCottas
against store.handle.
memoryTimings = {
const labels = { point: "point lookup", star: "star join", crossGraph: "cross-graph GRAPH join" };
const out = [];
for (const [key, sparql] of Object.entries(queries)) {
const t0 = performance.now();
const rows = await fn.query(parsed.dataset, sparql);
out.push({ query: labels[key], ms: Math.round((performance.now() - t0) * 100) / 100, rows: rows.length });
}
return out;
}
memoryTimings above — one run each, your browser, fn.query against
parsed.dataset.
The same three queries, one row per query, store time against memory time:
comparisonTable = storeTimings.map((s, i) => ({
query: s.query,
storeMedianMs: s.medianMs,
memoryMs: memoryTimings[i].ms,
storeRows: s.rows,
memoryRows: memoryTimings[i].rows,
answersMatch: s.rows === memoryTimings[i].rows,
}))
answersMatch reads true on every row — the store and the parsed
dataset answer the same three queries with the same row counts, over
the same 4,000 triples. storeMedianMs and memoryMs are not
comparable as "COTTAS vs. a heap dataset" in the abstract: the memory
side's cost is dominated by fn.query()'s per-call re-serialize/
re-parse round trip through the CLI bundle, an architectural cost of
this particular Node/browser adapter shape, not of holding triples in
memory in general. What the table does show directly: once
store.handle is open, a query against it does not pay that round
trip again.
store.bytes is a Uint8Array — the caller's to keep (IndexedDB,
OPFS, a download). Opening it again is opening the SAME COTTAS
artifact, not re-parsing the original 4,000 triples:
reopened = {
const t0 = performance.now();
const handle2 = await fn.openCottas(store.bytes);
const openMs = performance.now() - t0;
const rows = await fn.queryCottas(handle2, `# Count every triple across all named graphs in the reopened store.
SELECT (COUNT(*) AS ?n) WHERE { GRAPH ?g { ?s ?p ?o } }`);
await fn.closeCottas(handle2);
return { openMs, count: Number(rows[0].get("n").value) };
}
reopened.openMs should read close to store.openMs — opening a
second handle over the same bytes costs about what opening the first
one did, nothing like parsed.parseMs. reopened.count reads 4000:
the reopened handle answers over the full corpus, not a truncated or
cached-subset view.
Issue #595 tracks
the persistence program this page is one stop on: getting COTTAS bytes
that survive a page reload, not just a page session. The store this
page opens is the in-memory bytes tier of the same reader/writer the
CLI's on-disk --data-cottas store serves
(skills/disk-storage-format/SKILL.md
§6) — the difference is only where the bytes live (a Uint8Array in
this tab vs. a file on disk), never the parser, the query executor, or
the serializer.
The corpus-size finding above — parse and toCottas cost growing
faster than the triple count through this call path, well before
100,000 triples — has no GitHub issue of its own yet.
The live cells above are pinned in
tests/hub/post42_test.mjs —
the exact same source, executed against the real npm/factoidal typed
API (fn === factoidal, same shape as the browser adapter) with the
committed fixture read from disk instead of fetched.