The architectural proposals in this note were supplied by Dan Brickley in the Factoidal working session dated 2026-08-30. They include the Web-scale crawl/named-graph working-set model; modest warm composition of many prepared graphs; decentralized numeric IDs; reusable merged/resorted derived blocks; the distinction between page graphs and site/user/product/offer views; assertion-plus-evidence deduplication; RDFC-1.0 as a derivation boundary; and identity-neighbourhood-aware partitioning with publisher hints. Codex recorded, organized, and elaborated the proposals in this repository.
This is contemporaneous project provenance, not a legal conclusion about ownership, inventorship, patentability, copyright, or contractual rights.
The target is a Web-scale collection of immutable crawl/extraction snapshots, where a named graph commonly corresponds to a page but must not be treated as the sole useful knowledge-graph boundary. A site can contribute both site-authored and user-attributed material; an e-commerce page can combine site-wide organisation facts, product facts, offer facts, and page-local presentation/extraction observations.
The physical system must therefore support both page provenance and derived, non-page-shaped views such as all offers for a product, all organisation facts for a site, or a selected cross-source authority neighbourhood.
Many independently prepared named graphs and their blocks should be composable within milliseconds without fully decoding every graph. Warm state means manifests, graph/block directories, common-vocabulary dictionary pages, workload statistics/sketches, and recently used mapped pages — not a resident parsed RDF graph per crawl page.
When selected blocks already share compatible numeric term IDs, joins require
no semantic ID reconciliation. They can still cost many manifest lookups,
small reads, and a k-way merge of separately sorted streams. Repeatedly used
graph groups should therefore be eligible for content-addressed derived blocks
whose rows are globally resorted for their workload (for example PSOG,
SPOG, OSPG, or GSPO).
Derived blocks should normally retain graphId as a quad column. A flattened
union is a declared view over those blocks, not an irreversible loss of named
graph identity. This retains GRAPH semantics, source trust, retraction, and
provenance while permitting union-like execution.
Derived mini-KGs should deduplicate compatible assertion content without erasing evidence:
canonical assertion (s, p, o)
<- asserted-in / observed-in -> page graph, snapshot, extractor, voice role
Repeated equivalent site/product assertions can then occupy one execution row with multiple evidence links. Context-sensitive or contradictory claims stay distinct via their assertion/evidence records. This is the right model for site voice versus user voice and for site-wide versus offer/page-local facts.
Canonical RDF term identity remains the semantic ground truth. Compact IDs
are execution keys. Publisher-announced scoped numeric IDs (for common
vocabularies, controlled identifiers, or datasets) are welcome acceleration
certificates with issuer, namespace and version provenance; they are not
unqualified global RDF identity. A practical representation admits
well-known registry IDs, issuer-scoped IDs, and content-derived fallback IDs
behind one TermRef/dictionary contract.
Use the existing/formal RDFC-1.0 work at immutable ingestion and derivation boundaries: canonicalize a declared RDF dataset scope, preserve canonical N-Quads and a digest, and record source/derivation identity. This supports isomorphism-aware comparison, signatures, change detection, and reproducible derived working-set keys.
Do not put RDFC-1.0 on the hot block scan path. In particular, blank-node canonicalization has graph-wide and potentially adversarially costly behaviour. Runtime blocks use established compact IDs and sorted layouts after the canonicalization/assurance step.
Block boundaries must not be a rigid byte, character, or quad count. A payload block may contain five execution quads while its manifest records a larger declared identity/canonicalization neighbourhood used to establish the term references in those quads. This is particularly important for blank nodes and richly described entities: splitting away all identifying context can make canonicalization and entity reconciliation unstable or needlessly expensive.
The required distinction is:
payload quads -- scanned by the physical operator
identity neighbourhood -- bounded contextual evidence used at derivation
assertion/evidence links -- provenance and duplicate observations
Term identity itself remains stable and is not recomputed per block. The neighbourhood yields a checked mapping from source terms/blank nodes to canonical and compact IDs at ingest/derivation time; runtime blocks carry the resulting IDs plus an artifact reference to the evidence that established them.
Publisher partition hints should be accepted as advisory, provenance-bearing input: entity roots, record boundaries, blank-node closure groups, stable Skolem IRIs, and safe co-location groups. The importer must validate bounds and may enlarge a declared neighbourhood or reject a malformed one. This allows a publisher to keep an entity description together without making an untrusted page author control execution layout.
owl:FunctionalProperty and owl:InverseFunctionalProperty statements are
logical evidence, not licence to silently merge arbitrary Web terms. They may
inform a separately named reconciliation/entailment profile with explicit
scope, source trust, time/version, and explanation. Physical partitioning
should retain the relevant declaration and supporting neighbourhood together
where practical, but never detach it and then claim unconditional identity.
An especially common and useful reconciliation case, proposed by Dan
Brickley in the 2026-08-31 Factoidal working session, is a blank-node entity
with a value of a property accepted as an owl:InverseFunctionalProperty
under a named trust/profile boundary. For example, if the trusted profile
accepts un-blablah:countryCode as inverse-functional, occurrences of:
_:x un-blablah:countryCode "FR"
_:y un-blablah:countryCode "FR"
are evidence that _:x and _:y denote the same entity under that profile.
The physical engine can exploit this as an identity-key candidate:
profile / trust / version
+ predicate canonical identity
+ RDF-term-identity of object value
-> derived reconciliation key
This is deliberately not the raw blank-node label, a universally valid
Skolem IRI, or an unconditional global sameAs assertion. Its scope must
record at least the accepted IFP declaration, source/trust policy, RDF-term
identity version, observation/derivation time, conflict policy, and an
artifact/evidence reference. Multiple candidate keys, contradictory claims,
or an untrusted declaration must remain visible rather than forcing a merge.
The important implementation opportunity is earlier than post-hoc matching:
within one declared publisher/voice/subgraph boundary, this key can be the
deterministic basis for physical entity-ID assignment at ingest. Instead of
allocating an arbitrary new local number for every blank node and reconciling
later, the importer can derive a profile-scoped EntityId from the accepted
IFP predicate and object value. Separate crawls, blocks or installations
that apply the same profile obtain the same candidate ID without a central
allocator. That makes its rows directly joinable, sortable, co-locatable and
mergeable as they are published.
The storage model needs two identifiers, rather than overwriting RDF blank node identity:
SourceTermId -- document/graph-scoped RDF term; preserves _:x exactly
ProfileEntityId -- deterministic physical entity handle from trusted IFP key
Rows can retain SourceTermId for exact RDF reconstruction/provenance while
using ProfileEntityId as a derived join/sort key in an identity-aware block
or working-set overlay. The boundary can be a signed publisher dataset, a
declared site/voice, a product-record subgraph, or another explicitly named
identity domain. It must not quietly span unrelated Web sources merely
because their lexical literal happens to match.
If an entity supplies several accepted inverse-functional values, the profile can retain all candidate keys and choose a deterministic primary key only under a documented consistency rule. A collision, disagreement between keys, or later withdrawal of the IFP assertion records an identity conflict rather than changing historical source rows. A downstream asserted identity or Skolemization remains a separate, provenance-bearing derivation.
The next block manifest should therefore allow a compact advisory identity profile reference on a predicate-local artifact, for example:
predicate: un-blablah:countryCode
identity-key-profile: ifp/profile-17@2026-08
object-term-identity: rdf12-term-id-v1
It need not make raw IBK row bytes semantically depend on OWL. An IBK2 block format can contain several predicate segments, but the current Shardborough publisher emits predicate-local artifacts. That makes the manifest-level profile reference natural: it is an optional planning/reconciliation certificate for a known predicate, independently checked before use. A future multi-predicate block can carry the same information per segment or in its containing manifest; a single anonymous "IFP bit" on arbitrary block bytes would be too ambiguous about predicate, scope and authority.
Duplication is permitted at the physical/evidence level when it improves locality: a context quad or entity neighbourhood can be present in several derived blocks. Canonical assertion identity and artifact provenance ensure this does not become duplicated logical content or double-counted query results.
Dan Brickley's 2026-09-01 follow-up extends the same scoped-certificate idea
from identity to predicate access. If a query requests a superproperty such as
dc:description, an admitted schema may establish that exact predicate blocks
for application properties also contribute via rdfs:subPropertyOf. The
physical engine should not have to materialize every unrelated predicate block
to discover that relationship.
The acceleration artifact must be derived from and identify the exact source generation, schema/rules artifact, entailment profile, named-graph scope, and trust policy. It can map a requested superproperty to its eligible subordinate predicate partitions, or name a pre-merged materialized block for that view. It is advisory unless those identities match the query context; otherwise the engine falls back to complete semantic evaluation.
RDF graph set semantics matter here. If the same subject/object pair reaches a
superproperty through two subordinate predicates, it denotes one inferred
superproperty triple. A runtime union or materialized view must deduplicate at
that boundary before SPARQL solution multiplicity is exposed. More complex
property semantics—owl:inverseOf, transitivity, property chains and
owl:sameAs—need separate typed plans rather than being collapsed into the
subproperty map.
This keeps IBK row bytes vocabulary-neutral while making trusted schema
relationships useful for block pruning, sorting, merging and reusable warm
views. The consolidated contract is now in
docs/shardborough-storage-spec.md §8.
graphId and provenance links.factoidal-builds, retaining only provenance, commands, queries and result
manifests here.The first executable manifestation of this direction is SBM0, a deliberately
small Shardborough manifest format in
formal/lean4/L4Factoidal/Storage/ShardManifest.lean. It names an ordered
set of predicate-local IBK2 artifacts, each with a relative key, byte length,
SHA-256 digest and row count, together with source identity, term-registry
version and layout label.
SBM0 has strict Lean encode? and decode? functions. The decoder rejects
bad magic/version, malformed UTF-8 or predicate IRI, truncated fields,
incorrect digest width, trailing bytes and structurally invalid manifests.
This is a host-neutral control plane: an artifact key may later resolve to a
local file/range, PostgreSQL bytea, a TiKV value, browser OPFS, or object
storage. It is not yet a general graph/quad manifest or a Merkle tree; those
belong in the next layout versions once the checked child-artifact opening path
exists.
The next increment landed that opening path and the reference local host
vertical. ShardManifest.openStore? takes an injected ArtifactKey → Option ByteArray reader, checks each child's listed byte length and SHA-256,
then accepts it only if IndexedBlockWireV2.open? accepts its framing,
dictionary, directory and checksum. Its readOps is the established Lean
SPARQL backend capability. l4block-shard-pack now writes manifest.sbm0
as well as its TSV, while l4block-shard-query DIR --query SELECT... is the
first local-file host harness.
Its 2026-08-30 smoke used the 77-triple music Turtle fixture. A parsed,
two-predicate join with a filter and ORDER BY returned the three Radiohead
albums via seven verified child blocks. The eager opener was deliberately a
correctness reference. The local query host now adds a conservative first
lazy-selection step for ordinary parsed SPARQL: when the pattern is a native
BGP/join/union/minus composition and every triple pattern has a constant IRI
predicate, it reads, hash-verifies and opens only the corresponding predicate
artifacts. Its status line reports open-mode=predicate-selective(n) and
artifact-bytes=loaded/total. A deliberately small native FILTER subset
(constants, variables, comparisons, boolean logic and arithmetic) also retains
that path: it depends only on each solution mapping, not the active graph. The
PostgreSQL smoke's FILTER(?band = ex:radiohead) consequently opens only the
two by/title artifacts and still returns the three expected rows. OPTIONAL,
property paths, GRAPH, SERVICE, sub-SELECT, variable predicates, EXISTS and
other graph-dependent expressions deliberately use open-mode=full-manifest,
because those current evaluator paths can materialise the active graph. This
is a sound artifact-level I/O reduction, not yet mmap/range I/O inside a
selected IBK2 artifact. A later lazy/mmap/range reader must preserve the same
acceptance and readOps behavior before replacing it.
The same route was exercised on the bundled life-sciences chromosome.ttl
fixture: 9,227 triples packed to one predicate-local IBK2 child plus SBM0,
occupying 560 KiB, in approximately 25.5 seconds on the development Mac.
The parsed P31 SELECT … ORDER BY returned 9,227 rows from the manifested
artifact. The current packer is correctness-first and its load time is not a
throughput claim; it establishes the next corpus-sized regression point for
streaming and mmap/range work.
The same canonical-object contract has now crossed PostgreSQL bytea in a
local smoke: the manifest and every child block round-trip byte-for-byte, then
the Lean verifier and ordinary parsed SPARQL evaluation run on the retrieved
objects. This demonstrates an interchangeable persistence realization, not
yet PostgreSQL-side execution. TiKV should implement this exact artifact and
reader contract before any coprocessor/pushdown work is attempted.
Harness/ShardManifestSession.lean adds l4block-shard-session, a native
batch-session host for the same SBM0/IBK2 reader contract. It reads the
manifest once, accepts one complete SELECT query per input line, and keeps
only successfully length- and SHA-256-verified OpenBlock values in a local
in-memory cache for the duration of that request batch. A cache hit therefore
uses the immutable already-admitted bytes; an external change to the file
after admission cannot alter that session's cached execution input. A cache
miss repeats the normal safe leaf-name, file read, digest, and IBK2 structural
validation path before it is admitted.
The session remains deliberately bounded and line-oriented rather than using
an unbounded Lean partial loop. It is a host-level warm-working-set building
block: a supervisor can choose request batching/process lifetime, while the
shared parser, physical-planner guard, OpenStore.readOps, and SPARQL
evaluator are unchanged. Each result reports its shard count, selective or
full-manifest mode, cache hits/misses, newly admitted bytes, and cache size.
tools/blockengine-shard-session-smoke.sh establishes the concrete behavior
on the music fixture: a two-predicate parsed join admits two artifacts; a
subsequent native FILTER query over m:by has one cache hit and zero new
artifact bytes; a later variable-predicate query safely expands to all seven
predicate artifacts (two hits, five new admissions). This is an observable
modest warm state, not an assertion that mmap, PostgreSQL workers, TiKV, or
per-segment positioned reads are complete.
The predicate-shard packer now emits both manifest.sbm1 (64 KiB fixed chunk
commitments) and a compatibility manifest.sbm0 from the same IBK2 children.
The native one-shot and bounded-session query hosts prefer SBM1 whenever it is
present and fall back to SBM0 for older packed directories. The current host
continues whole-artifact SHA-256 admission, so this operational change proves
version selection and compatibility only; the subsequent range host will use
the SBM1 root plus ChunkedArtifact proofs to avoid that full read.
The same admission boundary now verifies that an SBM0 entry's declared row
count agrees both with the complete decoded IBK2 block and with the entry's
declared predicate segment. Only then may readOps.estimate use that count as
the exact cost of a predicate-only pattern; subject/object-constrained and
unbound patterns still take the ordinary scan path. This removes repeated
whole-block scans during the common first-stage join ordering decision without
turning unverified manifest metadata into planner truth.