Live mode — this page may load external map tiles and query remote endpoints; the standard hub is fully self-contained (same cells, no network).

Post 18 named toCottas/openCottas as the browser's write half of the on-disk COTTAS store, but never called them from a live cell. Post 24 ran a binary store's reader live, against a real fixture, with timed queries — but that fixture was 343 triples. This page runs the COTTAS equivalent at a larger scale: a four-graph corpus, loaded once into the same in-memory bytes store the CLI's --data-cottas-mem flag serves, queried three different shapes, and timed against the same queries run over an ordinary parsed dataset instead.

The corpus#

cottas-corpus.trig holds four named graphs — people, organizations, projects, and the affiliations linking them — generated by gen-cottas-corpus.mjs, a pure function of an integer index (no randomness, so every id's fields are computable, not just storable):

Graph Records Triples/record Triples
ex:gPeople 500 people 5 (type, name, age, mbox, memberSince) 2,500
ex:gOrgs 50 orgs 4 (type, name, location, founded) 200
ex:gProjects 100 projects 3 (type, name, budget) 300
ex:gAffiliations 500 links 2 (worksFor, contributesTo) 1,000
Total 4,000

The task behind this page targeted ~100,000 triples. Measured 2026-08-25, against the exact js_of_ocaml engine bundle this page loads (in-process, no OS subprocess — npm/factoidal/lib/engine-js.js drives the same bundle browser.js does): parsing a 100,000-triple corpus through this bundle took over two minutes and did not finish toCottas() within a further two-minute cap; even 10,000 triples cost 7.5 s to parse and 14.3 s to serialize to COTTAS bytes. That cost comes from the call path at this corpus size, measured directly — not from the fixture's file size (the committed .trig stays at 131 KB either way, far under the 6 MB budget). 4,000 triples keeps every cell comfortably inside the site's headless-Chromium sweep budget while still holding four named graphs with real cross-graph links.

fixture = {
  const t0 = performance.now();
  const text = await fetch("../assets/data/cottas-corpus.trig").then((r) => r.text());
  const fetchMs = performance.now() - t0;
  return { text, fetchMs, textBytes: text.length };
}
parsed = {
  const t0 = performance.now();
  const dataset = await fn.parse(fixture.text, { format: "trig" });
  const parseMs = performance.now() - t0;
  return { dataset, parseMs, tripleCount: dataset.size };
}

parsed.tripleCount should read 4000 — the fixture's own labelled count, round-tripped through the real F*-extracted TriG parser.

Bytes in, bytes out: toCottas#

fn.toCottas runs the SAME pure Tot serializer (RDF.CottasStore.BaseWriter.serialize_cottas_v2) the native factoidal compact --native-writer CLI path calls — not a parallel browser-only encoder (post 18's finding, restated here at four times the graph count):

store = {
  const t0 = performance.now();
  const bytes = await fn.toCottas(parsed.dataset);
  const toCottasMs = performance.now() - t0;
  const t1 = performance.now();
  const handle = await fn.openCottas(bytes);
  const openMs = performance.now() - t1;
  return {
    bytes, handle, toCottasMs, openMs,
    storeBytes: bytes.length,
    textBytes: fixture.textBytes,
  };
}

store.storeBytes compared with store.textBytes is the COTTAS artifact's size against the .trig text it was built from — a Parquet-based columnar encoding of a fairly compact, prefix-abbreviated serialization, so the win here is modest, not the order-of-magnitude difference the format shows on larger or less-abbreviated corpora.

Three query shapes, two ways#

Three SPARQL 1.1 queries, all against the same four graphs:

Each runs against store.handle (the COTTAS bytes store) and against parsed.dataset (the same triples, held as parsed text). The store side is timed as the median of three runs — labelled below — since a COTTAS handle's first query on a given predicate pays a one-time lazy dictionary-populate cost that later queries on the same handle do not; the table below shows all three per-query times, not only the median, so that first-call cost stays visible. The in-memory side is timed once: fn.query() serializes parsed.dataset back to text and reparses it through the CLI bundle on every single call — a fixed cost unrelated to which of the three queries runs, so three repeats would add three times the wall-clock cost without changing the comparison.

queries = ({
  point: `# Look up one person's name: one bound subject and predicate,
    # in one named graph.
    PREFIX ex: <http://example.org/factoidal/> PREFIX foaf: <http://xmlns.com/foaf/0.1/>
    SELECT ?name WHERE { GRAPH ex:gPeople { ex:person375 foaf:name ?name } }`,
  star: `# List every property/value pair one person has, in one named graph.
    PREFIX ex: <http://example.org/factoidal/>
    SELECT ?p ?o WHERE { GRAPH ex:gPeople { ex:person375 ?p ?o } }`,
  crossGraph: `# Three-graph join: a person's name from one graph, the
    # organisation they work for from a second, and that org's name from a third.
    PREFIX ex: <http://example.org/factoidal/> PREFIX foaf: <http://xmlns.com/foaf/0.1/> PREFIX org: <http://www.w3.org/ns/org#>
    SELECT ?personName ?orgName WHERE {
      GRAPH ex:gPeople { ex:person375 foaf:name ?personName }
      GRAPH ex:gAffiliations { ex:person375 org:worksFor ?org }
      GRAPH ex:gOrgs { ?org foaf:name ?orgName }
    }`,
})
storeTimings = {
  function median(a) { const s = [...a].sort((x, y) => x - y); return s[Math.floor(s.length / 2)]; }
  const labels = { point: "point lookup", star: "star join", crossGraph: "cross-graph GRAPH join" };
  const out = [];
  for (const [key, sparql] of Object.entries(queries)) {
    const times = [];
    let rows;
    for (let i = 0; i < 3; i++) {
      const t0 = performance.now();
      rows = await fn.queryCottas(store.handle, sparql);
      times.push(performance.now() - t0);
    }
    out.push({
      query: labels[key], medianMs: Math.round(median(times) * 100) / 100,
      runs: times.map((t) => Math.round(t * 100) / 100), rows: rows.length,
    });
  }
  return out;
}

storeTimings above — median of 3, your browser, fn.queryCottas against store.handle.

memoryTimings = {
  const labels = { point: "point lookup", star: "star join", crossGraph: "cross-graph GRAPH join" };
  const out = [];
  for (const [key, sparql] of Object.entries(queries)) {
    const t0 = performance.now();
    const rows = await fn.query(parsed.dataset, sparql);
    out.push({ query: labels[key], ms: Math.round((performance.now() - t0) * 100) / 100, rows: rows.length });
  }
  return out;
}

memoryTimings above — one run each, your browser, fn.query against parsed.dataset.

Store vs. memory#

The same three queries, one row per query, store time against memory time:

comparisonTable = storeTimings.map((s, i) => ({
  query: s.query,
  storeMedianMs: s.medianMs,
  memoryMs: memoryTimings[i].ms,
  storeRows: s.rows,
  memoryRows: memoryTimings[i].rows,
  answersMatch: s.rows === memoryTimings[i].rows,
}))

answersMatch reads true on every row — the store and the parsed dataset answer the same three queries with the same row counts, over the same 4,000 triples. storeMedianMs and memoryMs are not comparable as "COTTAS vs. a heap dataset" in the abstract: the memory side's cost is dominated by fn.query()'s per-call re-serialize/ re-parse round trip through the CLI bundle, an architectural cost of this particular Node/browser adapter shape, not of holding triples in memory in general. What the table does show directly: once store.handle is open, a query against it does not pay that round trip again.

Reopening the same bytes#

store.bytes is a Uint8Array — the caller's to keep (IndexedDB, OPFS, a download). Opening it again is opening the SAME COTTAS artifact, not re-parsing the original 4,000 triples:

reopened = {
  const t0 = performance.now();
  const handle2 = await fn.openCottas(store.bytes);
  const openMs = performance.now() - t0;
  const rows = await fn.queryCottas(handle2, `# Count every triple across all named graphs in the reopened store.
    SELECT (COUNT(*) AS ?n) WHERE { GRAPH ?g { ?s ?p ?o } }`);
  await fn.closeCottas(handle2);
  return { openMs, count: Number(rows[0].get("n").value) };
}

reopened.openMs should read close to store.openMs — opening a second handle over the same bytes costs about what opening the first one did, nothing like parsed.parseMs. reopened.count reads 4000: the reopened handle answers over the full corpus, not a truncated or cached-subset view.

What this is the browser tier of#

Issue #595 tracks the persistence program this page is one stop on: getting COTTAS bytes that survive a page reload, not just a page session. The store this page opens is the in-memory bytes tier of the same reader/writer the CLI's on-disk --data-cottas store serves (skills/disk-storage-format/SKILL.md §6) — the difference is only where the bytes live (a Uint8Array in this tab vs. a file on disk), never the parser, the query executor, or the serializer.

What's next#

The corpus-size finding above — parse and toCottas cost growing faster than the triple count through this call path, well before 100,000 triples — has no GitHub issue of its own yet.

The live cells above are pinned in tests/hub/post42_test.mjs — the exact same source, executed against the real npm/factoidal typed API (fn === factoidal, same shape as the browser adapter) with the committed fixture read from disk instead of fetched.