Ballyhoo Backlog#
Ballyhoo is the experimental parser / I/O / storage track.
It does not replace the current more compliant parser stack yet. Its role is
to explore faster ingestion, binary/on-disk paths, corpus organization, and
streaming interfaces without forcing everything through the current text-parser
architecture.
These tasks are intentionally tracked in local .md files rather than GitHub
issues for now.
Current Ballyhoo scope#
- streaming/chunk-oriented parsing experiments
- fast ingestion paths for large RDF inputs
- on-disk corpus organization
- binary artifact consumption, especially HDT-like formats
- bridge tooling for format conversion during testing
- Keep
Parser.Ballyhoo.fst separate and clearly marked experimental.
- Add a Ballyhoo-specific benchmark path so its timings are measured directly,
instead of inferring performance through the main parser stack.
- Add a small CLI or test harness for Ballyhoo event-counting and dataset
construction on line-oriented RDF inputs.
- Measure Ballyhoo on:
- small W3C N-Quads samples
- Berlin converted to
.nt
- large N-Quads slices
Corpus / TOC tasks#
- Define a TOC vocabulary for corpus chunks, source files, generated artifacts,
graph IRIs, and provenance.
- Treat the TOC as just another corpus chunk under
Corpus/toc/.
- Standardize the initial filesystem layout:
Corpus/
toc/
berlin-dbpedia-page/
v1/
data.hdt
- Define recognized version-directory conventions such as
v1/, v2/, etc.
- Decide how non-versioned corpora are represented.
- Record graph IRIs and artifact paths in the TOC graph.
HDT / binary ingestion tasks#
- Evaluate a binary ingestion path based on HDT or related formats.
- Keep in mind that standard HDT is graph-oriented, not dataset-oriented.
- For N-Quads corpora, consider:
- partition by graph IRI
- emit one HDT per logical graph
- record the mapping in the TOC graph
- Create a
Parser.BallyhooHDT.fst design note and module skeleton.
- Investigate which HDT metadata and structural checks are suitable for F*
binary parsing techniques.
- Tie HDT work to a SPARQL storage backend rather than treating HDT as only a
parser/import concern.
See
docs/designissues/sparql-store-backend.md.
SPARQL storage tasks#
- Stop assuming that SPARQL evaluation always runs over
list triple.
- Introduce a storage/query boundary under
SPARQL11.Algebra.
- Keep the algebra in F*, but let extracted backends answer indexed
triple-pattern lookups.
- Support both:
- persisted graph stores such as HDT-backed corpus chunks
- ephemeral in-memory indexed graph stores built from parsed RDF
- Rework named-graph access around manifest/TOC metadata, not quad-native HDT
assumptions.
Parallel ingestion tasks#
- Use N-Quads as the primary large-ingest format for parallel parsing.
- Split large N-Quads files by byte range, then adjust chunk boundaries to
newline boundaries.
- Confirm and document the key property: N-Quads lines can be parsed
independently once line boundaries are respected.
- Build a sharding utility that can:
- split N-Quads safely
- route lines by graph IRI
- emit line-oriented per-graph files for later HDT generation
- Keep
tools/rdf_convert.py as the standard local RDF format conversion tool.
- Add utilities around it for corpus import workflows.
- Benchmark and document
rdflib-based conversion on representative files.
- Consider temporary external-tool workflows acceptable for corpus preparation,
while keeping parser logic in F*.
Validation tasks#
- Run parser-quality checks against temporary bridge tooling such as
rdflib
so we know how much to trust conversions used for benchmarking or corpus
preparation.
- Separate these results from Factoidal’s own W3C conformance numbers.
- Record known bridge-tool limitations locally in docs.
Open architectural questions#
- How much of corpus metadata should be RDF immediately, versus bootstrap text
files that later generate RDF?
- Whether the TOC should itself be stored as HDT eventually.
- Whether dataset-preserving persistence should use HDT extensions, per-graph
HDTs, or another format entirely.
- What the clean API boundary should be between:
- corpus lookup
- artifact loading
- RDF graph/dataset construction
- SPARQL execution