SRI1 is canonical and correct, but its reader currently authenticates and decodes the entire flat sidecar before selecting a subject's row offsets. The gene benchmark therefore made TLI1 and PTD1 page-selective while retaining one large all-at-once SRI1 read. SRI2 is the compatible immutable successor.
It is not a new RDF identity scheme. Its keys remain IBK3-local subject IDs, so it must be bound to the exact target IBK3 digest and every returned offset must still be checked by reading the corresponding IBK3 row.
magic u32le "SRI2"
version u8 2
targetIBKSha256 32 bytes
rowCount u32le
pairCount u32le
pagePairs u32le 256
pageCount u32le
directoryBytes u32le
pageBytes u32le
directory (firstSubjectId, lastSubjectId, offset, length) × pageCount
pages sorted (subjectId, sourceRowOffset) u32 pairs
crc32c u32le over post-version bytes
Pairs retain SRI1's strict ordering: first by subject local ID, then by source row offset. A directory entry names its inclusive subject range. Lookup fetches the directory and fetches every page whose range can contain the requested subject. This deliberately handles a very frequent subject whose posting list spans several pages: it never guesses from a duplicated first-subject boundary and silently omits earlier postings. Normal lookups select one page; the multi-page case remains bounded and deterministic.
The decoded directory requires nondecreasing firstSubjectId and
lastSubjectId endpoints. The reader uses lower/upper-bound searches over
those two endpoint sequences, then scans only the resulting candidate interval.
Thus a normal lookup has logarithmic directory selection plus one page, while
a subject spanning several full pages deliberately receives every such page.
pairCount == rowCount; offsets are a permutation of 0 .. rowCount-1.SBM4 and SRI1 remain readable throughout. New packs publish SBM5; the older format remains a compatibility reader, rather than being silently rewritten.
Steps 1 and the storage-adjacent part of step 2 are now implemented in Lean:
Harness.IndexedBlockV3Materialize.subjectPostingsV2For? reads an explicitly
supplied, Merkle-committed SRI2 artifact by fetching its prefix, directory, and
only candidate pages. It verifies the SRI2 target digest and row count against
the IBK3 entry before returning postings, and returns the read footprints in
the normal query counters.
SBM5 now uses this reader for the parsed two-predicate shared-subject join.
The packer and compactor write predicate-N.ibk3.sri2; activation fully
decodes it, verifies its target digest, row count and complete posting relation
against the paired IBK3 block, then atomically permits CURRENT to select the
generation. Accordingly, the native SBM5 query route requires the collection
root with its CURRENT pointer; a bare unactivated SRI2 directory is rejected
rather than being treated as proof of a complete subject index. The query route announces
ibk3-sri2-tli1-subject-join; its first persistent smoke run returns the
established 290 result rows. This avoids an uncommitted convenience file
silently altering query answers.