Purpose: make Lean 4 block-engine measurements comparable and
reproducible across sessions and machines. This is a metadata
catalogue with retrieval rules — no large data is committed to this
repository. Requested in tmp_podman_codexfablesync.txt (Codex,
2026-08-31/09-01); written by the review session.
tools/. No number is written from a
website's claim (counting-coverage discipline).third_party/testing/ — term-shape coverage (lists, typed and
language-tagged literals) now exercised through the persisted path
(da5886c42, 23941fe7e, de85aec57). The persisted-path
executability census over these suites is
tools/w3c-persisted-census.sh
(results).258cbac9c) — controlled shape/skew
fixtures, seed-stable.4d66b1073):
44 statements, six datatypes, five language tags, blank nodes both
positions, the "1"/"01" term-identity sentinel, and the full
update → compaction → activation → re-query cycle asserted
(worknote). This closes the
rung-1 heterogeneity gap named below for rung 2; rung 3 is now
about SCALE with heterogeneity, not first heterogeneity.tools/corpus-profile.sh (bytes, SHA-256, engine-parsed counts,
predicate/object/datatype/language histograms).examples/wikidata/subsets/lifesci-kgx/data/gene.ttl — 888,949
triples (measured), Wikidata-derived, CC0. Predicate-skew-heavy,
weak on language tags.be added, hashes recorded at first retrieval)
| Candidate | Why | License | Scale (order) | Notes |
|---|---|---|---|---|
| UK Parliament curated store dump | Already a project corpus (tools/bench_ukpar_*); ontology-rich; dates, gYear, lang strings |
Open Parliament Licence | ~10^6–10^7 | The natural first heterogeneous rung; pipeline exists |
| Wikidata lexemes subset | Extreme language-tag diversity, CC0 | CC0 | 10^6+ (subset recipe) | Cut via existing lifesci-kgx-style extraction; record the query + dump date |
| DBLP RDF dump | Real bibliographic heterogeneity; stable monthly snapshots with checksums | CC0 | 10^8 full; slice to 10^6–10^7 | Good literal/typed-value mix |
| Nobel Prize Linked Data | Small but genuinely multi-lingual, multi-class | CC0 | ~10^5 | Cheap heterogeneity smoke |
| WatDiv generator | Deterministic synthetic skew/structuredness control | Academic open | any | For controlled physical-layout experiments; seed recorded |
URLs deliberately not written from memory — the fetch script records the URL it actually used, next to the sha256, at first retrieval.
| Candidate | Why | License | Scale | Notes |
|---|---|---|---|---|
| Wikidata truthy subset, recipe-defined | The stated project target; CC0 | CC0 | 10^7–10^8 by recipe | Define by an extraction query committed to the repo, not by a frozen file; record dump date + sha256 of the extract |
| UniProt RDF (taxonomy or citations slice) | Real datatypes at scale | CC BY 4.0 | 10^7+ | Attribution required; measure-only until attribution added |
| Full DBLP | Real, single-file, checksummed upstream | CC0 | ~4×10^8 | Ingest-scaling rung |
Base: the six competitive-bench queries
(docs/test-results/competitive-bench.json) plus, per the SBM6
work: object-bound lookups (IRI and literal objects, language-tagged
and typed), negative object/subject lookups, and a two-pattern join
driven from each side. Every run compares rows (not counts) against
the in-memory Lean evaluator on the same data. Updates: an INSERT
DATA / DELETE DATA / compaction / re-query cycle at rungs 1–3.
tools/corpus-fetch.sh <name> — downloads, hashes, and appends
the measured record to this file (script to be written when the
first Rung-3 corpus is pulled; owner approval for anything over
~100 MB).tools/corpus-profile.sh emitting
predicate counts, object cardinality per predicate, literal
datatype/lang histograms, via the existing engines.