How to read this page
- Green — full pass: every runnable test in the suite passes.
- Amber — partial: at least one fail, but every residual fail is diagnosed in writing (a link sits next to the row).
- Grey — not measured this run, or out of scope for this suite (e.g. skipped fixtures, a roadmap item with no runner yet).
Collapsed rows show a condensed score in pass/fail/skip of total order (e.g. “461/0/6 of 467” means 461 pass, 0 fail, 6 skip out of 467) — tap a row to expand it for the fully-labelled sentence and a stacked pass/fail/skip meter, measured against that suite's own total so proportions are comparable across suites of very different sizes. A family row with an amber N gaps badge has a “Remaining work” list inside it; an expanded family with no badge and no fails is marked complete against the vendored suites.
Harness diagnostics — escape branches (#316)
A conformance number is only as strong as its weakest comparator. These are the branches where a test leaves the plain run-it-and-compare-it path. budget exceeded = the RDFC-1.0 canonicalizer hit its Hash-N-Degree-Quads work budget, so graph isomorphism could not be decided; that test is scored FAIL (before 2026-07-29 it fell back to a blank-node-collapsing multiset comparison that could score it PASS). GSP seed = the Graph Store Protocol pre-state seeding branch fired. no manifest / zero tests = a suite discovered nothing; either makes the run exit 2 rather than report green.
| mode | tests discovered | budget exceeded | GSP seed | no manifest | zero tests | skip + unsupported |
|---|---|---|---|---|---|---|
sparql | 631 | 0 | 1 | 0 | 0 | 0 |
rdf | 1031 | 0 | 0 | 0 | 0 | 1 |
rdf12 | 242 | 0 | 0 | 0 | 0 | 0 |
rdf12c14n | 82 | 0 | 0 | 0 | 0 | 0 |
rdf12entail | 47 | 0 | 0 | 0 | 0 | 3 |
sparql12 | 254 | 0 | 0 | 0 | 0 | 0 |
No comparator escapes, no discovery faults. Every score above came from the strict RDFC-1.0 comparator on a suite that discovered its tests.
OWL 2 catalogs — closure cap escapes
The OWL runner puts a per-test CPU cap on the RL/DL closure, on the inconsistency-marker scan, and on the clash-seeking tableau, so one hard ontology cannot hang a catalog. On a cap-trip it falls back to a less-closed graph or to "nothing found". closure and silent count closure-stage escapes (the second column is the catch-all arms, which now log too); marker counts abandoned inconsistency-marker scans; refuter counts abandoned clash searches; unsupported (cap-escape) counts the Consistency / NegativeEntailment verdicts those escapes moved off PASS (#326). unsupported (indeterminate) counts the same two kinds moved off PASS for a different reason: the clash-seeking tableau finished inside its budget and returned don't-know, because the calculus cannot decide that input. A cap escape means "ran out of time" and a bigger budget may resolve it; an indeterminate means "our calculus cannot decide this" and no budget will. Either way, "we did not find X" is not published as "X is not there".
| catalog | closure | silent | marker | refuter | tests affected | unsupported (cap-escape) | unsupported (indeterminate) | which |
|---|---|---|---|---|---|---|---|---|
profile_rl | 0 | 0 | 0 | 0 | 0 | 0 | 0 | — |
type_positive_entailment | 8 | 0 | 2 | 0 | 5 | 5 | 0 | WebOnt-description-logic-661,WebOnt-description-logic-662 WebOnt-description-logic-663,WebOnt-description-logic-664 WebOnt-miscellaneous-011 |
type_negative_entailment | 0 | 0 | 0 | 0 | 0 | 0 | 0 | — |
type_consistency | 14 | 0 | 2 | 5 | 12 | 12 | 0 | WebOnt-Thing-004,WebOnt-description-logic-501 WebOnt-description-logic-661,WebOnt-description-logic-662 WebOnt-description-logic-663,WebOnt-description-logic-664 WebOnt-description-logic-905,WebOnt-description-logic-906 WebOnt-description-logic-907,WebOnt-miscellaneous-001 WebOnt-miscellaneous-002,WebOnt-miscellaneous-011 |
type_inconsistency | 0 | 0 | 0 | 1 | 1 | 0 | 0 | WebOnt-description-logic-910 |
profile_el | 1 | 0 | 0 | 0 | 1 | 1 | 0 | WebOnt-Thing-004 |
profile_ql | 0 | 0 | 0 | 0 | 0 | 0 | 0 | — |
semantics_direct | 14 | 0 | 2 | 6 | 13 | 12 | 0 | WebOnt-Thing-004,WebOnt-description-logic-501 WebOnt-description-logic-661,WebOnt-description-logic-662 WebOnt-description-logic-663,WebOnt-description-logic-664 WebOnt-description-logic-905,WebOnt-description-logic-906 WebOnt-description-logic-907,WebOnt-description-logic-910 WebOnt-miscellaneous-001,WebOnt-miscellaneous-002 WebOnt-miscellaneous-011 |
Listed tests were scored on reasoning the runner abandoned part-way. Since #326 the two kinds that pass on the ABSENCE of a derived fact (ConsistencyTest, NegativeEntailmentTest) score unsupported rather than PASS when that happens — a truncated closure satisfies them trivially, so their pass would have been an artifact of giving up. The other two kinds (PositiveEntailment, Inconsistency) are fail-safe under truncation: a missing derivation shows up as a FAIL.
Local overrides of upstream fixtures
A fixture whose upstream expectation we dispute in writing, under tests/local-overrides/. The runner scores it in its own bucket rather than as a pass or a fail. Listed here because folding it into “skip” hid a deliberate disagreement with the spec suite.
| suite | overrides | which |
|---|---|---|
rif_core | 1 | corpus:RDF_Combination_Constant_Equivalence_4 |
Known defects — found, filed, not yet fixed
Known defects (found, filed, unfixed) 2/3/0 of 5
2 known defects reproduce on this build, as expected; 3 unexpectedly gone; 0 probe errors (out of 5).
Defects we have found, reproduced and filed, and have not yet fixed. Each row is an executable probe run on every build, not a note in a tracker.
2 still reproduce (expected), 3 unexpectedly gone, 0 probe errors (out of 5).
Every one of these lives underneath a suite that
is green. SR-1 and SR-2 sit inside SPARQL 1.1 at
631 pass, 0 fail — a conformance suite measures the fixtures it ships,
and these defects have no fixture. When a row flips to XPASS
the known-defects suite fails on purpose: either it was fixed and the row
should go, or the probe drifted and stopped measuring what it claims.
| ID | Issue | Defect | State | Observed |
|---|---|---|---|---|
SR-1 | #336 | SELECT DISTINCT returns duplicate rows | XPASS | DISTINCT now returns 1 row — defect appears FIXED |
SR-2 | #337 | OPTIONAL misses a match the BGP finds (lang-tag case) | XPASS | BGP and OPTIONAL now agree — defect appears FIXED |
SE-1 | #324 | sameTerm case-folds language tags (term identity) | XFAIL | sameTerm("x"@en,"x"@EN) is TRUE; identity says they differ |
TTL-PFX | #334 | Turtle drops undeclared-prefix statements silently | XPASS | document now rejected (exit 1) — defect appears FIXED |
JSONLD-MSG | #275 | syntax error and missing loader share one message | XFAIL | inline-context syntax error still blamed on remote contexts |
W3C Recommendations
RDF 1.1 core 1030/0/0 of 1031 1 gap
1030 pass, 0 fail, 0 skip (of 1031) across N-Triples, Turtle, N-Quads, TriG, RDF/XML, and RDF 1.1 Semantics (rdf-mt, which carries the RDFS entailment tests).
RDF 1.1 N-Triples
rdf-n-triples70/0/0 of 70
70 pass, 0 fail, 0 skip (out of 70)
RDF 1.1 Turtle
rdf-turtle313/0/0 of 313
313 pass, 0 fail, 0 skip (out of 313)
RDF 1.1 N-Quads
rdf-n-quads87/0/0 of 87
87 pass, 0 fail, 0 skip (out of 87)
RDF 1.1 TriG
rdf-trig356/0/0 of 356
356 pass, 0 fail, 0 skip (out of 356)
RDF 1.1 RDF/XML
rdf-xml166/0/0 of 166
166 pass, 0 fail, 0 skip (out of 166)
RDF 1.1 Semantics
rdf-mt38/0/1 of 39
38 pass, 0 fail, 1 skip (out of 39)
Every residual in this family is named, explained, and dispositioned on the RDF conformance page, which also defines the entailment regimes (simple / RDF / RDFS / D), says which one each runner dispatches, and separates proved from measured.
Remaining work
- some very deeply nested RDF/XML documents can overflow the parser — a parser-depth gap, not a semantics gap
RDF Dataset Canonicalization (RDFC-1.0) 86/0/0 of 86
86 pass, 0 fail, 0 skip (of 86) on the W3C rdf-canon suite.
RDF Dataset Canonicalization (RDFC-1.0)
RDFC-1.0 (eval + Map + NegEval)86/0/0 of 86
86 pass, 0 fail, 0 skip (out of 86)
Runner: bin/rdfc10-runner (bin/linux-x86_64/rdfc10_runner) ·
Suite: third_party/testing/rdf-canon/ ·
Algorithm: F* formal/fstar/RDF.Canonical.fst (Hash First Degree Quads +
full Hash N-Degree Quads permutation enumeration), verified with no
--lax and no --admit_smt_queries. Eval tests compare the
canonical N-Quads form bytewise; Map tests compare the bnode→canonical-id
mapping structurally; NegEval tests are bounded by an HNDQ work budget.
Scored alongside the other RDF suites, with its assurance level stated, on the RDF conformance page.
Complete against the vendored suites.
SPARQL 1.1 631/0/0 of 631 2 gaps
631 pass, 0 fail, 0 skip (of 631) across Query, Update, Protocol, Federated Query, Service Description, and Entailment Regimes.
SPARQL 1.1 Query Language
aggregates47/0/0 of 47
47 pass, 0 fail, 0 skip (out of 47)
bind10/0/0 of 10
10 pass, 0 fail, 0 skip (out of 10)
bindings11/0/0 of 11
11 pass, 0 fail, 0 skip (out of 11)
cast6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
construct7/0/0 of 7
7 pass, 0 fail, 0 skip (out of 7)
csv-tsv-res6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
exists6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
functions75/0/0 of 75
75 pass, 0 fail, 0 skip (out of 75)
grouping6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
json-res4/0/0 of 4
4 pass, 0 fail, 0 skip (out of 4)
negation12/0/0 of 12
12 pass, 0 fail, 0 skip (out of 12)
project-expression7/0/0 of 7
7 pass, 0 fail, 0 skip (out of 7)
property-path33/0/0 of 33
33 pass, 0 fail, 0 skip (out of 33)
subquery14/0/0 of 14
14 pass, 0 fail, 0 skip (out of 14)
syntax-query94/0/0 of 94
94 pass, 0 fail, 0 skip (out of 94)
SPARQL 1.1 Update
add8/0/0 of 8
8 pass, 0 fail, 0 skip (out of 8)
basic-update13/0/0 of 13
13 pass, 0 fail, 0 skip (out of 13)
clear4/0/0 of 4
4 pass, 0 fail, 0 skip (out of 4)
copy6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
delete19/0/0 of 19
19 pass, 0 fail, 0 skip (out of 19)
delete-data6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
delete-insert17/0/0 of 17
17 pass, 0 fail, 0 skip (out of 17)
delete-where6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
drop4/0/0 of 4
4 pass, 0 fail, 0 skip (out of 4)
move6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
syntax-update-154/0/0 of 54
54 pass, 0 fail, 0 skip (out of 54)
syntax-update-21/0/0 of 1
1 pass, 0 fail, 0 skip (out of 1)
update-silent13/0/0 of 13
13 pass, 0 fail, 0 skip (out of 13)
SPARQL 1.1 Protocol
http-rdf-update19/0/0 of 19
19 pass, 0 fail, 0 skip (out of 19)
protocol34/0/0 of 34
34 pass, 0 fail, 0 skip (out of 34)
SPARQL 1.1 Federated Query
service7/0/0 of 7
7 pass, 0 fail, 0 skip (out of 7)
syntax-fed3/0/0 of 3
3 pass, 0 fail, 0 skip (out of 3)
SPARQL 1.1 Service Description
service-description3/0/0 of 3
3 pass, 0 fail, 0 skip (out of 3)
SPARQL 1.1 Entailment Regimes
entailment (SPARQL 1.1 regime — RDFS / D-entailment, 70 tests)70/0/0 of 70
70 pass, 0 fail, 0 skip (out of 70)
Every residual in this family is named, explained, and dispositioned on the SPARQL conformance page, which also says what these suites do NOT test (§18 algebra conformance is not the same as query-result conformance) and separates proved from measured.
Remaining work
- wire-level replay: the 53 W3C Protocol + Graph Store tests pass by dispatching each test's HTTP request through the verified protocol modules in-process; replaying the same suite against a live factoidal-http socket (which already serves query/update and, with --rw plus a delta log, Graph Store writes) needs a small harness and is not yet automated
- Graph Store writes over HTTP require the explicit --rw + --delta-log durable mode; the read-only default answers PUT/DELETE with 405 by design — document-and-test coverage for both modes belongs in the wire harness above
OWL 2 30/0/0 of 30 2 gaps
profile-RL PositiveEntailmentTests: 30 pass, 0 fail (of 30); six further OWL 2 catalogs (NegEnt/Cons/Inc + 4 DL catalogs) scored below, RIF Core scored in its own family.
OWL 2 30+ pass via owl_runner across 8 catalogs (profile-RL/EL/QL + 4 DL + syntax-dl species) · live scoring
Scope. The OWL 2 W3C Test Cases catalog (~3378
test:TestCase entries across 9 categories) is vendored
under third_party/testing/owl/. As of Phase 2.3
(2026-05-08), seven catalogs run live through
owl_runner with PositiveEntailment / NegativeEntailment /
Consistency / Inconsistency scoring: profile-RL.rdf,
profile-EL.rdf, profile-QL.rdf,
type-positive-entailment.rdf,
type-negative-entailment.rdf,
type-consistency.rdf, type-inconsistency.rdf,
and semantics-direct.rdf. The ninth catalog
(syntax-dl.rdf) scores through the F\*
OWL2.SyntaxDL species checker (its row appears above).
Every remaining failure is named, explained, and dispositioned on the
OWL 2 conformance page.
Tableau on the live codepath. The F\*
Tableau.tableau_materialise module (0
assume val, 0 --lax) drives the SPARQL
entailment regime codepath via w3c_runner.ml:
parent4/5/6/7, simple7/8, sparqldl-01…12, etc. — the
SPARQL 1.1 Entailment Regimes row above
passes (see its live score) because Tableau drives the membership
check. Phase 2.3d (2026-07-09) wires the same Tableau
materialisation into owl_runner via
--regime dl: each DL catalog row below is tagged
[DL] (Tableau) or [RL] (Datalog closure)
so the two regimes are never silently mixed. DL runs RL-closure
→ tableau_materialise → RL-closure, falling
back to the RL closure on a per-test cap-trip, so every DL row
scores ≥ its RL baseline. Tableau is positive-sound, so the
RL→DL flips it produces are gains (type-inconsistency and
positive-entailment rows moved up), never wrong answers.
Pass-rate context. Across the scored catalogs,
the bulk of remaining failures fall into two categories:
(1) tests that use OWL Functional-Style Syntax (not RDF/XML)
trigger FAIL/no-premise — these need fuller FSS parser
coverage before scoring is meaningful; (2) Inconsistency /
entailment failures beyond the tableau's decided fragment. The
refutation side landed 2026-07-10 (Tableau.Refute.fst:
NNF, lazy TBox unfolding, disjunction branching, ∃-witnesses,
complement / min-max / counting / bottom-property clashes — each
Some false carries a per-rule Direct Semantics
soundness argument, and the DL rows below score it), so what
remains on the inconsistency side is the undecided residue:
nominals (owl:oneOf), datatype facets, inverse-role
interaction, and searches that exhaust the refuter's linear work
budget (deep propositional encodings) — indeterminate results fall
back to the RL verdict, never below it.
OWL DL entailment — Tableau (Tableau.fst · tableau_materialise, live in w3c_runner — SPARQL 1.1 entailment-regimes suite)70/0/0 of 70
70 pass, 0 fail, 0 skip (out of 70)
profile-RL PosEnt30/0/0 of 30
30 pass, 0 fail, 0 skip (out of 30)
profile-RL NegEnt6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
profile-RL Consistency76/0/0 of 76
76 pass, 0 fail, 0 skip (out of 76)
profile-RL Inconsistency14/0/0 of 14
14 pass, 0 fail, 0 skip (out of 14)
type-PosEnt PE [DL]195/9/0 of 204
195 pass, 9 fail, 0 skip (out of 204)
type-PosEnt Cons [DL]199/5/0 of 204
199 pass, 5 fail, 0 skip (out of 204)
type-NegEnt PE [DL]0/0/0 of 0
0 pass, 0 fail, 0 skip (out of 0)
type-NegEnt NE [DL]23/0/0 of 23
23 pass, 0 fail, 0 skip (out of 23)
type-NegEnt Cons [DL]23/0/0 of 23
23 pass, 0 fail, 0 skip (out of 23)
type-Cons PE [DL]195/9/0 of 204
195 pass, 9 fail, 0 skip (out of 204)
type-Cons NE [DL]23/0/0 of 23
23 pass, 0 fail, 0 skip (out of 23)
type-Cons Cons [DL]340/12/0 of 352
340 pass, 12 fail, 0 skip (out of 352)
type-Inc PE [DL]0/0/0 of 0
0 pass, 0 fail, 0 skip (out of 0)
type-Inc Inc [DL]126/1/0 of 127
126 pass, 1 fail, 0 skip (out of 127)
profile-EL PE [RL]29/0/0 of 29
29 pass, 0 fail, 0 skip (out of 29)
profile-EL NE [RL]6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
profile-EL Cons [RL]71/1/0 of 72
71 pass, 1 fail, 0 skip (out of 72)
profile-EL Inc [RL]13/0/0 of 13
13 pass, 0 fail, 0 skip (out of 13)
profile-QL PE [RL]20/0/0 of 20
20 pass, 0 fail, 0 skip (out of 20)
profile-QL NE [RL]3/0/0 of 3
3 pass, 0 fail, 0 skip (out of 3)
profile-QL Cons [RL]58/0/0 of 58
58 pass, 0 fail, 0 skip (out of 58)
profile-QL Inc [RL]6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
sem-Direct PE [DL]195/9/0 of 204
195 pass, 9 fail, 0 skip (out of 204)
sem-Direct NE [DL]23/0/0 of 23
23 pass, 0 fail, 0 skip (out of 23)
sem-Direct Cons [DL]339/12/0 of 351
339 pass, 12 fail, 0 skip (out of 351)
sem-Direct Inc [DL]126/1/0 of 127
126 pass, 1 fail, 0 skip (out of 127)
syntax-dl species DL-vs-FULL [syntactic]319/2/2 of 323
319 pass, 2 fail, 2 skip (out of 323)
OWL 2 (W3C conformance):
We vendor the full W3C OWL 2 Test Cases at
third_party/testing/owl/ (10 catalog files, ~2500
test:TestCase entries after overlap). After Phase 2.3
(2026-05-08), all 8 main catalogs run live through
owl_runner: 7 with PE/NE/Cons/Inc scoring, plus (as of
2026-07-10) syntax-dl.rdf with species identification.
The runner applies owl_rl_closure_with_reflexivity
(fuel 100) and for entailment tests checks the conclusion’s
triples against the closure (relaxed bnode match); for consistency
tests it consults RDF_Graph_Executable.is_inconsistent
against the same closure; for species identification it consults the
syntax-directed F* checker OWL2.SyntaxDL.species_is_dl
(no reasoning) over premise + conclusion documents.
Other suites once listed here as roadmap items — GeoSPARQL, JSON-LD 1.1,
CSVW, ShEx, DID, VC, RML — now have live scored suite nodes of their
own elsewhere on this page.
Remaining work
- the DL entailment catalogs (EL, QL, and the general DL profile) are wired through the tableau reasoner (DL regime = RL closure -> tableau_materialise -> RL closure, landed 2026-07-09) but still carry fails — some because the parser doesn't yet cover OWL Functional-Style Syntax fixtures, and some because the reasoner materialises entailments but doesn't yet refute the negative (inconsistency) side
- large graphs using owl:sameAs extensively can hit a performance blowup during closure computation — a known scaling gap, not a correctness gap
SHACL (Shapes Constraint Language) 120/0/0 of 120
120 pass, 0 fail, 0 skip (out of 120) across SHACL Core and SHACL SPARQL-based Constraints.
SHACL Core98/0/0 of 98
98 pass, 0 fail, 0 skip (out of 98)
Runner: bin/shacl-runner (bin/linux-x86_64/shacl_runner) · Suite: third_party/testing/shacl/data-shapes-test-suite/tests/core/
SHACL SPARQL-based Constraints22/0/0 of 22
22 pass, 0 fail, 0 skip (out of 22)
Runner: bin/shacl-runner (bin/linux-x86_64/shacl_runner tests/sparql/manifest.ttl) · Suite: third_party/testing/shacl/data-shapes-test-suite/tests/sparql/
See also ShEx, the other shapes-constraint language for RDF — a W3C Community Group specification, not a Recommendation.
Complete against the vendored suites.
CSVW (CSV on the Web) 270/0/0 of 270 8 gaps
270 pass, 0 fail, 0 skip (out of 270) on the vendored csv2rdf corpus.
CSVW csv2rdf270/0/0 of 270
270 pass, 0 fail, 0 skip (out of 270)
Runner: bin/csvw-runner (bin/linux-x86_64/csvw_runner) · Suite: vendored W3C csv2rdf corpus (ToRdf / ToRdfWithWarnings / NegativeRdf)
Remaining work
- metadata discovery over HTTP (Link headers, a metadata directory, the /.well-known/csvm well-known path) is not implemented — this is a networked-protocol feature, parked by owner directive alongside SPARQL Protocol
- some metadata-normalization warnings are not produced: @base/@type context edge cases (invalid-property pruning and container-shape graceful degradation landed 2026-07-15)
- schema compatibility and validation gaps: the header=false option, column-name restriction checks, and title-language intersection checks (required/null column handling landed 2026-07-15)
- complex fixtures combining virtual columns, multiple subjects per row, referenced schemas, and inherited-property combinations are not fully handled
- rowTitles is not implemented
- table-level separator handling and ordered-list values are not implemented
- the duration format pattern needs a regex engine this checker doesn't yet call out to
- 2 cross-table foreign-key validation tests (NegativeRdf) are not handled
JSON-LD 1.1 1286/64/2 of 1352 · 7 of 7 suites 1 gap
1286 pass, 64 fail, 2 skip (of 1352) across the W3C JSON-LD 1.1 toRdf, fromRdf, expand, compact and flatten manifests.
JSON-LD 1.1 toRdf467/0/0 of 467
467 pass, 0 fail, 0 skip (out of 467)
Runner: bin/jsonld-runner (bin/linux-x86_64/jsonld_runner) · Suite: third_party/testing/json-ld/ (toRdf manifest)
JSON-LD 1.1 fromRdf53/0/1 of 54
53 pass, 0 fail, 1 skip (out of 54)
Runner: bin/jsonld-fromrdf-runner (bin/linux-x86_64/jsonld_fromrdf_runner) · Suite: W3C JSON-LD 1.1 fromRdf manifest (Serialize RDF as JSON-LD)
JSON-LD 1.1 expand385/0/0 of 385
385 pass, 0 fail, 0 skip (out of 385)
Runner: bin/jsonld-expand-runner (bin/linux-x86_64/jsonld_expand_runner) · Suite: W3C JSON-LD 1.1 expand manifest (Expansion Algorithm)
JSON-LD 1.1 compact245/0/1 of 246
245 pass, 0 fail, 1 skip (out of 246)
Runner: bin/jsonld-compact-runner (bin/linux-x86_64/jsonld_compact_runner) · Suite: W3C JSON-LD 1.1 compact manifest (Compaction Algorithm)
JSON-LD 1.1 flatten58/0/0 of 58
58 pass, 0 fail, 0 skip (out of 58)
Runner: bin/jsonld-flatten-runner (bin/linux-x86_64/jsonld_flatten_runner) · Suite: W3C JSON-LD 1.1 flatten manifest (Node Map Generation + Flattening Algorithm)
JSON-LD 1.1 HTML50/0/0 of 50
50 pass, 0 fail, 0 skip (out of 50)
Runner: bin/jsonld-html-runner (bin/linux-x86_64/jsonld_html_runner) · Suite: W3C JSON-LD 1.1 html manifest (JSON-LD embedded in HTML <script> elements)
JSON-LD 1.1 Framing28/64/0 of 92
28 pass, 64 fail, 0 skip (out of 92)
Runner: bin/jsonld-frame-runner (bin/linux-x86_64/jsonld_frame_runner) · Suite: W3C json-ld-framing manifest (Framing Algorithm)
Remaining work
- frame suite (Framing Algorithm): first-cut F* algorithm — passes the basic match/embed cases; residual fails are the un-implemented frame features (@explicit, @default, @omitDefault, @embed @always/@never/@link, @reverse, @included, named graphs, value-level matching). See jsonld-frame.yaml.
Verifiable Credentials 2.0 117/0/0 of 117 3 gaps
117 pass, 0 fail, 0 skip (out of 117) on the structural Stage 1 suite. Signing and verifying credentials (the Data Integrity eddsa-rdfc-2022 proof: RDFC-1.0 canonicalization + SHA-256 + Ed25519 via HACL*) runs in the native binary, Node, and the browser bundle; full conformance testing of that signing layer against the official W3C proof-options test suite is still pending.
VC Data Model 2.0 — structural (Stage 1)117/0/0 of 117
117 pass, 0 fail, 0 skip (out of 117)
Runner: bin/vc-runner (bin/linux-x86_64/vc_runner) · Suite: third_party/testing/vc/tests/input/ (structural fixtures, filename-encoded verdicts)
Data Integrity eddsa-rdfc-2022 (W3C vc-di-eddsa-test-suite, via shim)31/0/0 of 31
31 pass, 0 fail, 0 skip (out of 31)
Runner: tests/vc-di-eddsa/run.sh · Suite: third_party/testing/vc-di-eddsa/tests/{05-di-rdfc-create,15-di-rdfc-verify}.js against bin/vc-api-shim/server.mjs
VC Data Model 2.0 — issuer/verifier HTTP suite (via shim)59/0/0 of 59
59 pass, 0 fail, 0 skip (out of 59)
Runner: tests/vc20-api/run.sh · Suite: third_party/testing/vc/tests/{1.03-conformance,4.03-contexts,...,7-algorithms}.js (14 files) against bin/vc-api-shim/server.mjs
EECC VC/DID interop fixtures (vendored vc-verifier-rules + webuild-attestations)4/0/51 of 55
4 pass, 0 fail, 51 skip (out of 55)
Runner: bin/eecc-runner (bin/linux-x86_64/eecc_runner) · Fixtures: third_party/testing/eecc/ (vendored real-world GS1 credentials, Apache-2.0)
VC community compatibility (canivc.com)snapshot 2026-07-10
Digital Bazaar's canivc.com aggregates the same official W3C/CCG test-suite reports across ~20 implementations. Per-suite comparison (community min/median/max % across registered implementations, from a snapshot vendored at third_party/testing/canivc-snapshot/; full scoping + gap rationale in the canivc.com integration doc):
| suite | Factoidal | community min / median / max | status |
|---|---|---|---|
| eddsa-rdfc-2022 (create+verify, 31 tests) | 83.9% (26/31) | 0% / 41.9% / 100% | measured — above median |
| VC Data Model 2.0 (issuer/verifier HTTP) | 37.3% (22/59) | 22.0% / 54.7% / 96.6% | measured — below median (structural validation gap, see row above) |
| did:key | not measured via HTTP resolver | 25.0% / 25.0% / 87.5% | assessed, not implemented — see scoping doc |
| Data Integrity ECDSA | not implemented | 0% / 90.0% / 100% | out of scope — no ECDSA cryptosuite wired (HACL* has P-256 source, not adopted) |
| Ed25519Signature2020 | not implemented | 0% / 78.6% / 100% | out of scope — legacy cryptosuite, assessed distance from eddsa-rdfc-2022, not attempted |
| Data Integrity BBS | not implemented | 2.3% / 78.5% / 100% | out of scope — no BBS signature scheme anywhere in the codebase |
| VC JOSE/COSE | not implemented | 68.6% / 100% / 100% | out of scope — different (JWT/CWT) securing stack |
| VC Bitstring Status List | not implemented | 3.8% / 32.7% / 96.2% | out of scope — credentialStatus processing not implemented |
| VC-API Issuer (full) | not implemented | 0% / 42.9% / 96.4% | out of scope — full VC-API surface far beyond issue/verify |
| VC-API Verifier (full) | not implemented | 0% / 55.6% / 97.2% | out of scope — same reason |
Remaining work
- context-driven type redefinition is not checked — it needs real JSON-LD term resolution this stage does not implement
- full W3C VC Data Integrity conformance against the official corpus is now measured (task #88, 2026-07-10): eddsa-rdfc-2022 create+verify scores 26 pass, 5 fail (out of 31) via a new HTTP shim (bin/vc-api-shim) — see .github/test-suites/vc-di-eddsa.yaml and docs/designissues/2026-07-10-canivc-community-compat.md. Not yet measured: inside the browser/JS bundle specifically (the shim runs under Node only), and eddsa-jcs-2022 (JCS canonicalization is a different transform path, not implemented)
- the VC Data Model 2.0 suite's own issuer/verifier HTTP tests (distinct from this suite's structural-fixture check) now also run against the shim: 22 pass, 37 fail (out of 59), see .github/test-suites/vc20-api.yaml — the shim has no structural validator of its own, so almost every failure is a missing-rejection-of-malformed-input test
DID (Decentralized Identifiers) 8/0/0 of 8 3 gaps
8 pass, 0 fail, 0 skip (out of 8) on did:key resolution.
DID did:key8/0/0 of 8
8 pass, 0 fail, 0 skip (out of 8)
Runner: bin/did-runner (bin/linux-x86_64/did_runner) · Suite: did:key resolution + multibase/multicodec accept/reject cases
Remaining work
- this suite covers only did:key — did:web (or any other DID method) is not implemented
- DID resolution metadata is not produced — a full DID Core conformance target still needs to be enumerated, not just did:key vectors
- assessed against the official w3c-ccg/did-key-test-suite (task #88, 2026-07-10) and not implemented: it is a resolver-conformance suite (GET /didResolvers/<did> returning a full DID Resolution Result envelope with didDocument/didDocumentMetadata/didResolutionMetadata and structured error codes) rather than the vector-accept/reject shape this suite covers; didKeyResolve's npm export returns bare RDF, not that envelope. See docs/designissues/2026-07-10-canivc-community-compat.md's did:key section for the specific gap.
Rules: RIF Core 46/0/4 of 50 3 gaps
46 pass, 0 fail, 4 skip (of 50) — Part 1 (4 vendored SPARQL-manifest cases) + Part 2 (46-test W3C RIF Core dialect corpus).
RIF Core — combined (Part 1 + Part 2)46/0/4 of 50
46 pass, 0 fail, 4 skip (out of 50)
Runner: bin/rif-runner (bin/linux-x86_64/rif_runner) · Suites: third_party/testing/rif/tc/ + third_party/testing/rif-core-suite/
RIF Core — Part 1 (4 vendored SPARQL-manifest cases)4/0/0 of 4
4 pass, 0 fail, 0 skip (out of 4)
Runner: bin/rif-runner · Suite: third_party/testing/rif/tc/ (the same 4 cases are also scored independently as the “RIF Core” row under SPARQL 1.1 Entailment Regimes above, via w3c_runner's entailment dispatch — both pipelines agree)
RIF Core — Part 2 (W3C Core_v1.22 corpus)42/0/4 of 46
42 pass, 0 fail, 4 skip (out of 46)
Runner: bin/rif-runner · Suite: third_party/testing/rif-core-suite/Core_v1.22/Approved/
Remaining work
- 1 local-override: RDF_Combination_Constant_Equivalence_4 is a defect in the vendored corpus data itself (a malformed xsd:string datatype IRI), not an engine bug -- dispositioned via tests/local-overrides/rif/ (distinctly counted, exit 0)
- 2 skips: RIF List terms (Builtins_List, NestedListsAreNotFlatLists) are not modelled in RIF.Core.Syntax -- a syntax/translation/eval extension, not attempted
- 1 skip: the full date/time/duration builtin family (~85 builtins; Builtins_Time) is not implemented -- only the EBusiness_Contract dateTime slice exists; the skip names the first unimplemented builtin (is-literal-dateTimeStamp)
GRDDL 18/50/0 of 68 5 gaps
18 pass, 50 fail, 0 skip (out of 68) on the local GRDDL Stage 1 subset.
GRDDL Stage 1 (local subset)18/50/0 of 68
18 pass, 50 fail, 0 skip (out of 68)
Runner: bin/grddl-runner (bin/linux-x86_64/grddl_runner) · Suite: local subset of the GRDDL Test Cases
the fails are graph-mismatch cases in the local subset; the skips are fixtures that need a live network fetch of a transformation profile, out of scope for an offline conformance run
Remaining work
- 30 fail-graph-mismatch: consumed XSLT-engine fidelity gaps (e.g. grokSheet.xsl Gnumeric, hcard2rdf.xsl), HTTP content-negotiation (langconneg/multipleRepresentations), namespace-transformation recursion+merge, and XInclude posture — engine bugs are reported, not fixed in the GRDDL layer
- 6 fail-known-gap-xslt-feature: inline-rdf xml:base/xml:lang handling in embeddedRDF.xsl
- 10 fail-no-transformation-discovered: namespace/profile-document tests whose SECOND document is not vendored offline (two/three/four-transforms, ns-*, sq1/sq2)
- 7 fail-transform-not-vendored: a discovered stylesheet (.sxsl, or a redirect-base path) is not present in the offline docroot
- Recursion + loop-detection (loop/loopx/ns-* chains) is single-level only; HTTP redirect-base and content-negotiation are not simulated; Stage 3 (HTML tag-soup input) is not implemented
XSLT 1.0 87/0/1 of 88
87 pass, 0 fail, 1 skip (out of 88) on XSLT 1.0 transform conformance.
XSLT 1.087/0/1 of 88
87 pass, 0 fail, 1 skip (out of 88)
Runner: bin/xslt-runner (bin/linux-x86_64/xslt_runner) · Suite: third_party/testing/xslt/manifest.json
Complete against the vendored suites.
XML 1.0 1447/0/1138 of 2585
1447 pass, 0 fail, 1138 skip (out of 2585) on XML 1.0 well-formedness conformance.
XML 1.0 conformance1447/0/1138 of 2585
1447 pass, 0 fail, 1138 skip (out of 2585)
Runner: bin/xml-runner (bin/linux-x86_64/xml_runner) · Suite: third_party/testing/xml/xmlconf (OASIS/W3C XML conformance) · skips are DTD-boundary / out-of-XML-1.0-profile fixtures plus the NOT-APPLICABLE fixtures whose own EDITION attribute excludes the 5th edition this parser targets (correct scoping, not a gap) — all decomposed in the breakdown below
XML conformance — integrity breakdown (real rejects vs skips)
-- HONEST BREAKDOWN (PART 1 integrity accounting) --
not-wf REAL passes (parser rejected for the tested construct): 720
valid wf-accept passes (DOCTYPE internal subset now parsed): 727
--> real pass total: 1447
not-wf FAILs (parser wrongly ACCEPTED a wf-1.0 not-wf doc): 0
valid FAILs (parser wrongly REJECTED a valid doc): 0
--> fail total: 0
SKIP: not-wf DTD-internal (accepted; violation in the DTD
subset, parsed-but-not-validated — Stage-A scope limit): 331
SKIP: valid DTD-boundary (rejected; markup entity / external
subset / DTD construct beyond the WF slice): 58
SKIP: out-of-profile (not-wf only under XML 1.1 / Namespaces,
parser is XML 1.0 non-namespace): 37
SKIP: vacuous (RETIRED by Stage-A DTD support; must be 0): 0
NOT-APPLICABLE: fixture's own EDITION attribute excludes the 5th
edition, which this parser targets — the rule it asserts
was retired by the spec we implement, so it is correct
scoping, NOT a gap (not-wf 312, valid 0, other 1): 313
SKIP: other (DOCTYPE-external-on-valid, encoding, invalid/error
by design, external-entity exemption, file-not-found): 399
--> skip total (INCLUDES the not-applicable count above): 1138
(side-check: PassVacuous counter = 0, must equal vacuous skip = 0)
XML_Wellformedness.is_valid_ncname (informational only, see module comment):
checked 1458 element tags across accepted documents; 12 would have been
rejected by the NCName production (expected — it excludes ':' from Name
start/continue characters, which is an RDF/XML-domain rule, not generic
XML conformance; plain XML 1.0 legally allows ':' in Names).
========================================
Complete against the vendored suites.
MathML 3 81/0/0 of 81
81 pass, 0 fail, 0 skip (out of 81) on Content MathML evaluation.
MathML 3 content81/0/0 of 81
81 pass, 0 fail, 0 skip (out of 81)
Runner: bin/mathml-runner (bin/linux-x86_64/mathml_runner) · Suite: third_party/testing/mathml/manifest.json (Content MathML evaluation)
Complete against the vendored suites.
W3C Working Drafts (emerging)
RDF 1.2 (N-Triples / N-Quads / Turtle / TriG) 242/0/0 of 242
242 pass, 0 fail, 0 skip (of 242) — W3C Working Draft, run via w3c_runner --rdf12.
RDF 1.2 (N-Triples / N-Quads / Turtle / TriG)
rdf-n-triples/syntax29/0/0 of 29
29 pass, 0 fail, 0 skip (out of 29)
rdf-n-quads/syntax27/0/0 of 27
27 pass, 0 fail, 0 skip (out of 27)
rdf-turtle/syntax67/0/0 of 67
67 pass, 0 fail, 0 skip (out of 67)
rdf-turtle/eval29/0/0 of 29
29 pass, 0 fail, 0 skip (out of 29)
rdf-trig/syntax35/0/0 of 35
35 pass, 0 fail, 0 skip (out of 35)
rdf-trig/eval25/0/0 of 25
25 pass, 0 fail, 0 skip (out of 25)
rdf-xml/eval30/0/0 of 30
30 pass, 0 fail, 0 skip (out of 30)
Residuals for this suite are named and dispositioned on the RDF conformance page.
Complete against the vendored suites.
RDF 1.2 Canonicalization (N-Triples / N-Quads) 82/0/0 of 82
82 pass, 0 fail, 0 skip (of 82) — W3C Working Draft, run via w3c_runner --rdf12c14n.
RDF 1.2 Canonicalization (N-Triples / N-Quads)
rdf-n-triples/c14n41/0/0 of 41
41 pass, 0 fail, 0 skip (out of 41)
rdf-n-quads/c14n41/0/0 of 41
41 pass, 0 fail, 0 skip (out of 41)
Residuals for this suite are named and dispositioned on the RDF conformance page.
Complete against the vendored suites.
RDF 1.2 Semantics (simple entailment) 41/3/3 of 47
41 pass, 3 fail, 3 skip (of 47) — W3C Working Draft, run via w3c_runner --rdf12entail.
RDF 1.2 Semantics (simple entailment)
rdf-semantics41/3/3 of 47
41 pass, 3 fail, 3 skip (out of 47)
Residuals for this suite are named and dispositioned on the RDF conformance page.
SPARQL 1.2 Query 254/0/0 of 254
254 pass, 0 fail, 0 skip (of 254) — W3C Working Draft, run via w3c_runner --sparql12.
SPARQL 1.2 Query
codepoint-escapes8/0/0 of 8
8 pass, 0 fail, 0 skip (out of 8)
eval-triple-terms41/0/0 of 41
41 pass, 0 fail, 0 skip (out of 41)
expression1/0/0 of 1
1 pass, 0 fail, 0 skip (out of 1)
grouping1/0/0 of 1
1 pass, 0 fail, 0 skip (out of 1)
lang-basedir11/0/0 of 11
11 pass, 0 fail, 0 skip (out of 11)
rdf113/0/0 of 3
3 pass, 0 fail, 0 skip (out of 3)
syntax2/0/0 of 2
2 pass, 0 fail, 0 skip (out of 2)
syntax-triple-terms-negative65/0/0 of 65
65 pass, 0 fail, 0 skip (out of 65)
syntax-triple-terms-positive113/0/0 of 113
113 pass, 0 fail, 0 skip (out of 113)
version9/0/0 of 9
9 pass, 0 fail, 0 skip (out of 9)
Residuals for this suite are named and dispositioned on the SPARQL conformance page.
Complete against the vendored suites.
W3C Community Group / Notes / Submissions
ShEx (Shape Expressions) 1182/0/0 of 1182
1182 pass, 0 fail, 0 skip (out of 1182) on the shexSpec/shexTest validation manifest.
ShEx Validation1182/0/0 of 1182
1182 pass, 0 fail, 0 skip (out of 1182)
Runner: bin/shex-runner (bin/linux-x86_64/shex_runner) · Suite: third_party/testing/shex/ (shexSpec/shexTest, ShExJ-first)
ShEx negativeSyntax (ShExC grammar-reject)not measured this run
not measured this run
Runner: bin/shex-runner --negative-syntax · Suite: third_party/testing/shex/negativeSyntax/ — every fixture MUST fail to parse; pass = Parser.ShExC (F*) rejects it
See also SHACL, the other shapes-constraint language for RDF — a W3C Recommendation.
Complete against the vendored suites.
HDT not measured this run
No format-conformance suite to measure; see the backend parity check below.
HDT (binary RDF format)no format-conformance suite vendored
no format-conformance suite vendored
No HDT-specific conformance suite is vendored in this repository (the HDT Member Submission publishes a binary format spec, not a test suite). Correctness of this project's HDT backend is instead checked by the HDT stage-4 backend parity check below, which compares HDT-backed query results against the in-memory backend byte-for-byte on unbound/bound-S/P/O/ASK/COUNT queries.
RML (rml-core / rml-io) 93/1/55 of 149 4 gaps
93 pass, 1 fail, 55 skip (out of 149) across RML rml-core and RML rml-io source-tests.
RML rml-core76/0/0 of 76
76 pass, 0 fail, 0 skip (out of 76)
Runner: bin/rml-runner (bin/linux-x86_64/rml_runner) · Suite: third_party/testing/rml-modules/rml-core/
RML rml-io (source-tests)17/1/55 of 73
17 pass, 1 fail, 55 skip (out of 73)
Runner: bin/linux-x86_64/rml_runner --io · Suite: third_party/testing/rml-modules/rml-io/ (RMLSTC0* source tests) · rml-cc (content-container) has no row: the committed runner exposes no --cc mode yet (tracked with the rml program plan)
Remaining work
- rml-io's logical-target section (writing to a target RDF dataset, RMLTTC0* fixtures) is not scored — only the source-tests section is measured
- rml-cc (content-container: gather/gatherAs) has no test row — the committed runner doesn't expose a --cc mode yet
- rml-fnml (function-map dispatch) and rml-star (quoted-triple mapping) are vendored but not wired to any runner
- relational/SQL logical sources are out of scope for this stage — only JSON and CSV logical sources are evaluated
Other standards (OGC / ISO / independent)
GeoSPARQL (OGC) 37/0/0 of 37
37 pass, 0 fail, 0 skip (out of 37) on geof: topology + WKT unit fixtures.
GeoSPARQL (geof: topology + WKT)37/0/0 of 37
37 pass, 0 fail, 0 skip (out of 37)
OGC standard (not a W3C Recommendation) · Runner: tests/unit/run-all.sh geosparql_v0_unit · local fixtures, not the OGC compliance suite
Complete against the vendored suites.
SPARQL extras: entailment regimes 70/0/0 of 70
A duplicate presentation of the SPARQL 1.1 Entailment Regimes score, surfaced here for readers looking for reasoning-adjacent SPARQL capability.
SPARQL 1.1 Entailment Regimes70/0/0 of 70
70 pass, 0 fail, 0 skip (out of 70)
Same measurement as the SPARQL 1.1 family's Entailment Regimes row · driven by Tableau.fst's tableau_materialise via w3c_runner
Complete against the vendored suites.
SPARQL full-text search (jena-text convention) not measured this run
Implemented (text:query magic property); no vendorable conformance corpus exists, so no score is claimed.
Jena rdf:text (full-text search)no test data vendored
no test data vendored
This project implements the jena-text text:query magic property (Slice 1: exact/token AND-match, no BM25 ranking) in formal/fstar/SPARQL.FullText.fst, exercised by hand-written local fixtures (tests/local/fulltext_slice1.sh, hub post 20) rather than an official conformance suite — Apache Jena does not publish a vendorable rdf:text/full-text-search test corpus, and none is checked into this repository. No pass/fail score is reported here because there is nothing to check it against.
QUDT (units of measure) 9/0/29 of 38
9 pass, 0 fail, 29 skip (out of 38) across QUDT's own shipped SHACL rulesets (v3.4.0). An independent ontology suite (qudt.org), not W3C — no official conformance suite exists, so these are the scoping doc's Layer-A targets.
QUDT integrity (contributor ruleset vs distribution)0/0/29 of 29
0 pass, 0 fail, 29 skip (out of 29)
Runner: bin/qudt-runner (bin/linux-x86_64/qudt_runner --integrity) · Data: third_party/qudt/QUDT-all-in-one-SHACL.ttl (131k triples) vs COLLECTION_QUDT_QA_TESTS_ALL.ttl; one entry per ruleset shape; upstream data findings annotated in tests/qudt/dispositions.tsv, never patched
QUDT user shapes (deprecation + consistency fixtures)9/0/0 of 9
9 pass, 0 fail, 0 skip (out of 9)
Runner: bin/qudt-runner (bin/linux-x86_64/qudt_runner --fixtures) · Fixtures: tests/qudt/fixtures/ (-ok/-viol verdicts) vs COLLECTION_QUDT_USER_TESTS.ttl · remaining: Layer B exact-rational conversion + dimension algebra in F*, Layer C qudtf: SPARQL functions
Complete against the vendored suites.
JSON Schema 770/0/0 of 770
770 pass, 0 fail, 0 skip (out of 770) on the JSON-Schema-Test-Suite draft-07 battery. An independent specification (json-schema.org), not W3C.
JSON Schema draft-07770/0/0 of 770
770 pass, 0 fail, 0 skip (out of 770)
Runner: bin/jsonschema-runner (bin/linux-x86_64/jsonschema_runner) · Suite: third_party/testing/jsonschema/ (JSON-Schema-Test-Suite draft7) · skips are optional/format vocabularies out of scope for the core validator
Complete against the vendored suites.
ISO Schematron 8/0/0 of 8
8 pass, 0 fail, 0 skip (out of 8) on rule-based assertion cases. An ISO standard (ISO/IEC 19757-3), not a W3C specification.
ISO Schematron8/0/0 of 8
8 pass, 0 fail, 0 skip (out of 8)
Runner: bin/schematron-runner (bin/linux-x86_64/schematron_runner) · Suite: third_party/testing/schematron/ (rule-based assertion cases)
Complete against the vendored suites.
Factoidal internal suites (engine end-to-end, parity, regressions)
Storage backend & JS runtime: HDT parity / hub / npm 761/1/2 of 764
761 pass, 1 fail, 2 skip (of 764) across HDT backend parity and the browser/npm bundle suites (node --test).
HDT stage-4 backend parity6/0/0 of 6
6 pass, 0 fail, 0 skip (out of 6)
Runner: tests/local/hdt_stage4_parity.sh · Fixture: third_party/testing/hdt/rml-core-ontology.hdt vs ground-truth .nt (unbound/bound-S/P/O/ASK/COUNT, byte-identical)
hub browser-bundle cells441/1/0 of 442
441 pass, 1 fail, 0 skip (out of 442)
Runner: node --test tests/hub/*.mjs · live cells for every docs-hub post, run against the JS/wasm bundle
npm package suite314/0/2 of 316
314 pass, 0 fail, 2 skip (out of 316)
Runner: node --test npm/factoidal/test/*.test.js · the 7 engine FP APIs, HDT-in-bundle, delta-log, VC crypto (HACL* wasm)
F* unit regressions & adjacent engines: GeoSPARQL / XPath / TOAN / XForms 169/28/0 of 197
150 pass, 0 fail, 0 skip (of 150) across GeoSPARQL v0, XPath 1.0, the TOAN/Matrix CAS engines, and the XForms model — shipped F* engines now surfaced from the native unit harness or the JS bundle; the aggregate row below reports how many of the 41 native unit files link and pass in this checkout.
GeoSPARQL (geof: topology + WKT)37/0/0 of 37
37 pass, 0 fail, 0 skip (out of 37)
Runner: tests/unit/run-all.sh geosparql_v0_unit · F* RDF.Geo.* — exact-rational WKT geometry + Simple-Features topology + geof: distance/envelope. Hub post21 pins 12 more assertions live against the browser bundle (counted in the npm/hub runtime family).
XPath 1.0 (unit)100/0/0 of 100
100 pass, 0 fail, 0 skip (out of 100)
Runner: tests/unit/run-all.sh xpath_tests · the F* XPath 1.0 evaluator over location-step / predicate / node-set / function-library cases.
TOAN CAS + Math.Matrix engines11/0/0 of 11
11 pass, 0 fail, 0 skip (out of 11)
Runner: node --test npm/factoidal/test/{toan,matrix}.test.js · the F* CAS (summation / product / simplify / diff / subst) + exact-rational matrix engines via the JS bundle. The npm package suite (runtime family above) additionally exercises xslt, mathml, xforms, jsonschema, schematron, hdt, and vc-crypto — each a shipped engine with its own test file.
XForms model (bind/recalc)2/0/0 of 2
2 pass, 0 fail, 0 skip (out of 2)
Runner: node --test npm/factoidal/test/xforms.test.js · the F* XForms bind/recalc model via the JS bundle. The larger native suite (tests/unit/xforms_tests.ml, ~29 bind/recalc cases) can't link in this checkout — its XForms_Bind.cmx is not in the committed artifact set — so the harness fix (#82) is landed but that suite awaits a committed build; no fabricated 29/29 is shown.
F* unit regressions (tests/unit)19/28/0 of 47
19 pass, 28 fail, 0 skip (out of 47)
Runner: tests/unit/run-all.sh · files-passing / files-failing across all 41 native unit files, each relinked against its own committed-.cmx dependency closure. The 28 failing files either need a module with no committed .cmx (Math_* / MathML_* / XForms_Bind engines) or hit a committed-.cmx epoch mismatch (RML_Eval vs SPARQL11_Algebra committed in different builds) — an artifact-staleness gap, not a harness bug.
Quad-store baseline measured 2026-08-25T12:47:29Z · commit abd7ff6
Measured against the committed bin/linux-x86_64/factoidal binary (median of 3 runs, no toolchain build). Store-path queries use the native factoidal import writer -- the pycottas/DuckDB cottas-import writer is measured for cold import time and size only; its stores are rejected on open by this binary, reproduced this run (issue #445 format-compatibility gate).
QLever comparison: unavailable: pip install qlever succeeds (installs the Python control script only); qlever index --system native fails ('qlever-index: command not found' -- the QLever C++ binaries are not shipped by the pip package and are not present in this container); the container-native alternative, --system docker (the package's default), cannot be reached either (docker ps: 'dial unix /var/run/docker.sock: connect: no such file or directory' -- no daemon running). No native or containerized QLever run was possible in this container.
Source: tools/bench-quadstore-baseline.sh, raw data perf-quadstore-baseline.json.
Import (cold wall time, on-disk size)
| Writer | Sidecars | Dataset | Quads | Median seconds | Peak kB | data.cottas bytes | Sidecars bytes |
|---|---|---|---|---|---|---|---|
| pycottas | false | 10k | 10000 | 1.7607s | 156456 | 10193 | 0 |
| pycottas | true | 10k | 10000 | skipped: Error: cottas_ondisk_open returned None for /dev/shm/factoidal-bench-quadstore/stores/pycottas-10k-sctrue/bench-10k/v1/data.cottas | |||
| pycottas | false | 100k | 100000 | 12.2761s | 360856 | 73698 | 0 |
| pycottas | true | 100k | 100000 | skipped: Error: cottas_ondisk_open returned None for /dev/shm/factoidal-bench-quadstore/stores/pycottas-100k-sctrue/bench-100k/v1/data.cottas | |||
| pycottas | false | 1m | 1000000 | 66.9816s | 1855444 | 704704 | 0 |
| pycottas | true | 1m | 1000000 | skipped: Error: cottas_ondisk_open returned None for /dev/shm/factoidal-bench-quadstore/stores/pycottas-1m-sctrue/bench-1m/v1/data.cottas | |||
| pycottas | false | quads100k | 100000 | 14.1451s | 391636 | 73924 | 0 |
| pycottas | true | quads100k | 100000 | skipped: Error: cottas_ondisk_open returned None for /dev/shm/factoidal-bench-quadstore/stores/pycottas-quads100k-sctrue/bench-quads100k/v1/data.cottas | |||
| native | false | 10k | 10000 | 0.7060s | 60892 | 377799 | 0 |
| native | true | 10k | 10000 | 0.7891s | 60892 | 377799 | 646508 |
| native | false | 100k | 100000 | 7.9769s | 424876 | 3898116 | 0 |
| native | true | 100k | 100000 | 9.0953s | 424952 | 3898116 | 6585570 |
| native | false | 1m | 1000000 | 95.7620s | 2098512 | 40120436 | 0 |
| native | true | 1m | 1000000 | skipped: timeout: exceeded 120s cap | |||
| native | false | quads100k | 100000 | 9.7523s | 414060 | 3945718 | 0 |
| native | true | quads100k | 100000 | 9.8475s | 432256 | 3945718 | 6585691 |
| extra | false | ukparliament | 0 | skipped: third_party/data/ukparliament/*.trig not present in this checkout | |||
| extra | false | berlin | 0 | skipped: examples/data/third_party/Berlin.ttl not present in this checkout | |||
Query latency (store vs. in-memory)
| Path | Writer | Sidecars | Dataset | Query | Median seconds | Peak kB |
|---|---|---|---|---|---|---|
| cottas | native | false | 10k | q1 | 0.2167s | 20828 |
| cottas | native | false | 10k | q2 | 0.2764s | 20828 |
| cottas | native | false | 10k | q3 | 0.2018s | 20828 |
| cottas | native | true | 10k | q1 | 0.1940s | 20956 |
| cottas | native | true | 10k | q2 | 0.1766s | 20828 |
| cottas | native | true | 10k | q3 | 0.2235s | 20828 |
| cottas | native | false | 100k | q1 | 1.6748s | 73560 |
| cottas | native | false | 100k | q2 | 1.5835s | 73728 |
| cottas | native | false | 100k | q3 | 1.6088s | 73592 |
| cottas | native | true | 100k | q1 | 1.7552s | 74076 |
| cottas | native | true | 100k | q2 | 1.4968s | 74240 |
| cottas | native | true | 100k | q3 | 1.6720s | 74244 |
| cottas | native | false | 1m | q1 | 4.1374s | 155388 |
| cottas | native | false | 1m | q2 | 4.3689s | 149416 |
| cottas | native | false | 1m | q3 | 5.5351s | 191088 |
| cottas | native | true | 1m | q1 | 3.7874s | 155404 |
| cottas | native | true | 1m | q2 | 4.0709s | 149396 |
| cottas | native | true | 1m | q3 | 5.7100s | 191036 |
| cottas | native | false | quads100k | q1 | 2.0161s | 68188 |
| cottas | native | false | quads100k | q2 | 2.3545s | 72592 |
| cottas | native | false | quads100k | q3 | 2.4418s | 68188 |
| cottas | native | true | quads100k | q1 | 1.9427s | 68700 |
| cottas | native | true | quads100k | q2 | 1.8906s | 73144 |
| cottas | native | true | quads100k | q3 | 2.2492s | 68828 |
| cottas | pycottas | n/a | all | skipped: issue #445 format-compatibility gate rejects pycottas/DuckDB-written stores on this binary (see header comment); not retried per-size | ||
| memory | n/a | n/a | 10k | q1 | 0.2608s | 21852 |
| memory | n/a | n/a | 10k | q2 | 0.2352s | 21852 |
| memory | n/a | n/a | 10k | q3 | 0.2656s | 25948 |
| memory | n/a | n/a | 100k | q1 | 2.4634s | 91276 |
| memory | n/a | n/a | 100k | q2 | 2.6031s | 91352 |
| memory | n/a | n/a | 100k | q3 | 3.1050s | 147024 |
| memory | n/a | n/a | 1m | q1 | 26.7314s | 875624 |
| memory | n/a | n/a | 1m | q2 | 29.8641s | 875624 |
| memory | n/a | n/a | 1m | q3 | 41.8131s | 1435232 |
| memory | n/a | n/a | quads100k | q1 | 2.8032s | 79992 |
| memory | n/a | n/a | quads100k | q2 | 2.8404s | 79992 |
| memory | n/a | n/a | quads100k | q3 | 3.1405s | 128776 |
RDFS / OWL-RL closure benchmark measured 2026-08-02T22:30:47Z · commit 175da4e
factoidal entail --regime RDFS against the committed bin/linux-x86_64/factoidal binary, median of 3 runs, each run capped at 120s. End-to-end: read + parse + closure + serialize. Wall time and CPU time are reported separately because the OWL cap-escape family is budgeted in CPU seconds.
Source: tools/bench-closure.sh, raw data closure-bench.json.
Synthetic shapes (scaling)
| Case | n | Input triples | Output triples | Expansion | Wall s (median) | CPU s (median) |
|---|---|---|---|---|---|---|
| chain | 20 | 21 | 309 | 14.7x | 0.013 | 0.013 |
| chain | 40 | 41 | 979 | 23.9x | 0.032 | 0.032 |
| chain | 80 | 81 | 3519 | 43.4x | 0.119 | 0.118 |
| chain | 160 | 161 | 13399 | 83.2x | 0.646 | 0.643 |
| tree | 128 | 128 | 949 | 7.4x | 0.028 | 0.028 |
| tree | 256 | 256 | 2103 | 8.2x | 0.067 | 0.067 |
| tree | 512 | 512 | 4665 | 9.1x | 0.168 | 0.168 |
| tree | 1024 | 1024 | 10299 | 10.1x | 0.448 | 0.447 |
| diamond | 32 | 57 | 494 | 8.7x | 0.020 | 0.020 |
| diamond | 64 | 121 | 1966 | 16.2x | 0.063 | 0.064 |
| diamond | 128 | 249 | 7982 | 32.1x | 0.310 | 0.309 |
| diamond | 192 | 377 | 18094 | 48.0x | 0.834 | 0.832 |
| wideflat | 500 | 502 | 2044 | 4.1x | 0.056 | 0.056 |
| wideflat | 1000 | 1002 | 4044 | 4.0x | 0.113 | 0.113 |
| wideflat | 2000 | 2002 | 8044 | 4.0x | 0.231 | 0.231 |
| wideflat | 4000 | 4002 | 16044 | 4.0x | 0.490 | 0.488 |
| dense | 250 | 1000 | 2803 | 2.8x | 0.084 | 0.084 |
| dense | 500 | 2000 | 5553 | 2.8x | 0.176 | 0.176 |
| dense | 1000 | 4000 | 11053 | 2.8x | 0.390 | 0.384 |
| dense | 2000 | 8000 | 22053 | 2.8x | 0.933 | 0.898 |
Real vocabularies
| Case | n | Input triples | Output triples | Expansion | Wall s (median) | CPU s (median) |
|---|---|---|---|---|---|---|
| skos | 254 | 422 | 1.7x | 0.025 | 0.026 | |
| foaf | 635 | 863 | 1.4x | 0.039 | 0.039 | |
| dcterms | 700 | 1272 | 1.8x | 0.055 | 0.055 | |
| schemaorg | 17949 | 29275 | 1.6x | 1.977 | 1.973 | |
| qudt | 130404 | 508139 | 3.9x | 87.093 | 86.945 |
Fitted scaling exponents (log-log least squares over the n sweep): chain time n1.892, output n1.816; tree time n1.333, output n1.147; diamond time n2.097, output n2.011; wideflat time n1.042, output n0.991; dense time n1.155, output n0.992.
Machine-readable artifacts
latest.csv— one row per suite (timestamp, commit, branch, category, suite, pass/fail/skip/unsupported)latest.json— same data plus totals, structuredhistory/<timestamp>.csv/.json— timestamped copies, one pair per runner invocationperf-parse-serialize.json— parse/serialize/canonicalize throughput (if present; produced bytools/bench-parse-serialize.sh, not this script)perf-quadstore-baseline.json— quad-store import/query baseline (if present; produced bytools/bench-quadstore-baseline.sh, not this script)
The raw runner logs (including per-test FAIL lines with diffs) are committed under
formal/fstar/ocaml-output/*_results.log — one file per suite, named to
match each suite's .github/test-suites/<suite>.yaml manifest.
Raw per-suite numbers
SPARQL 1.1
631 total: 631 pass, 0 fail, 0 skip, 0 unsupported add pass:8 fail:0 skip:0 unsupported:0 aggregates pass:47 fail:0 skip:0 unsupported:0 basic-update pass:13 fail:0 skip:0 unsupported:0 bind pass:10 fail:0 skip:0 unsupported:0 bindings pass:11 fail:0 skip:0 unsupported:0 cast pass:6 fail:0 skip:0 unsupported:0 clear pass:4 fail:0 skip:0 unsupported:0 construct pass:7 fail:0 skip:0 unsupported:0 copy pass:6 fail:0 skip:0 unsupported:0 csv-tsv-res pass:6 fail:0 skip:0 unsupported:0 delete pass:19 fail:0 skip:0 unsupported:0 delete-data pass:6 fail:0 skip:0 unsupported:0 delete-insert pass:17 fail:0 skip:0 unsupported:0 delete-where pass:6 fail:0 skip:0 unsupported:0 drop pass:4 fail:0 skip:0 unsupported:0 entailment pass:70 fail:0 skip:0 unsupported:0 exists pass:6 fail:0 skip:0 unsupported:0 functions pass:75 fail:0 skip:0 unsupported:0 grouping pass:6 fail:0 skip:0 unsupported:0 http-rdf-update pass:19 fail:0 skip:0 unsupported:0 json-res pass:4 fail:0 skip:0 unsupported:0 move pass:6 fail:0 skip:0 unsupported:0 negation pass:12 fail:0 skip:0 unsupported:0 project-expression pass:7 fail:0 skip:0 unsupported:0 property-path pass:33 fail:0 skip:0 unsupported:0 protocol pass:34 fail:0 skip:0 unsupported:0 service pass:7 fail:0 skip:0 unsupported:0 service-description pass:3 fail:0 skip:0 unsupported:0 subquery pass:14 fail:0 skip:0 unsupported:0 syntax-fed pass:3 fail:0 skip:0 unsupported:0 syntax-query pass:94 fail:0 skip:0 unsupported:0 syntax-update-1 pass:54 fail:0 skip:0 unsupported:0 syntax-update-2 pass:1 fail:0 skip:0 unsupported:0 update-silent pass:13 fail:0 skip:0 unsupported:0
RDF 1.1
1031 total: 1030 pass, 0 fail, 0 skip, 1 unsupported rdf-mt pass:38 fail:0 skip:0 unsupported:1 rdf-n-quads pass:87 fail:0 skip:0 unsupported:0 rdf-n-triples pass:70 fail:0 skip:0 unsupported:0 rdf-trig pass:356 fail:0 skip:0 unsupported:0 rdf-turtle pass:313 fail:0 skip:0 unsupported:0 rdf-xml pass:166 fail:0 skip:0 unsupported:0
Shapes / Rules / Mapping / JSON-LD / VC
shacl-core: 98 pass, 0 fail, 0 skip (of 98) — present=1 shacl-sparql: 22 pass, 0 fail, 0 skip (of 22) — present=1 shex: 1182 pass, 0 fail, 0 skip (of 1182) — present=1 shex-negative-syntax: 0 pass, 0 fail (of 0) — present=0 jsonld-tordf: 467 pass, 0 fail, 0 skip (of 467) — present=1 rml-core: 76 pass, 0 fail, 0 skip (of 76) — present=1 rif-core: 46 pass, 0 fail, 4 skip (of 50) — present=1 vc-stage1: 117 pass, 0 fail, 0 skip (of 117) — present=1
How this page is generated
Source: formal/fstar/generate-report.sh. It shells out to the
w3c_runner binary (extracted from F* specs, compiled via OCaml),
scrapes per-suite counts, and writes index.html, latest.csv,
and latest.json. Run ./generate-report.sh --run in
formal/fstar/ to regenerate. CI re-runs on every push and nightly
at 06:00 UTC.