Precompiled ShEx schemas
rudof can compile a ShEx schema down to its internal representation
(SchemaIR) once and reuse that artefact across many validations
without paying the parsing and compilation cost every time. This guide
explains what the precompiled cache is, when it is worth using, and
how to drive it from the CLI.
Why precompile a schema?
Every rudof shex-validate invocation does two things before it can
touch a single triple of data:
- Parses the ShEx source (ShExC, ShExJ, Turtle, …) into an AST.
- Compiles the AST into
SchemaIR, the shape of the schema that the validator actually runs against.
For small schemas both steps are essentially free. For larger schemas (hundreds of shapes, many imports, deep inheritance) the compilation step is expensive, and it is exactly the same work every time you validate the same schema. The precompiled cache lets you do that work once and skip it on every subsequent run.
The compile - validate loop
1. Compile the schema
Use the shex command with --compile-to:
rudof shex --schema examples/user.shex --compile-to user.ircache
rudof loads the schema exactly as it would for validation
and writes the resulting SchemaIR to user.ircache. The file is a pure
binary artefact: it starts with the magic bytes RSIR, followed by an on-disk
envelope version, a length-prefixed postcard
header (body format, rudof version, negation-cycle flag), and the
postcard-encoded SchemaIR body. Because the file is not text, do not open it in
editors that may rewrite line endings or normalise encoding.
You can also compile as a side-effect of a validation run:
rudof shex-validate \
--schema examples/user.shex \
--compile-to user.ircache \
--shapemap examples/user.sm \
examples/user.ttl
2. Validate against the cache
Point shex-validate at the cache with --compiled-schema instead
of --schema:
rudof shex-validate \
--compiled-schema user.ircache \
--shapemap examples/user.sm \
examples/user.ttl
3. Inspect a compiled cache with shex
shex also accepts --compiled-schema, mirroring shex-validate. This
is useful to inspect a cache produced elsewhere without needing the
original schema source file:
rudof shex --compiled-schema user.ircache -r internal --statistics true --show-dependencies true
This loads the SchemaIR straight from the cache (skipping parsing,
imports and AST-to-IR compilation) and prints its internal
representation, shape statistics, and dependency graph.
Only -r internal (and statistics/dependencies, which read from the
same SchemaIR) work against a compiled schema. Every other result
format (shexc, shexj, json, jsonld, turtle, plantuml, …)
needs the original parsed schema (AST), which the precompiled cache
does not preserve – only the compiled SchemaIR is. Asking for one of
those formats fails with an error naming the limitation. To convert a
schema to one of those formats, load the original source with
--schema instead of --compiled-schema.
4. The unified alternative: binary as a ShExFormat
Everything above also works without the dedicated --compile-to/
--compiled-schema flags, since binary is a regular ShExFormat value
usable anywhere a ShEx format is accepted – the same -s/-f/-r/-o
flags used to convert between shexc, shexj, turtle, etc. drive
compiling and loading the cache too:
# Compile: same effect as --compile-to.
rudof shex -s examples/user.shex -r binary -o user.ircache
# Load the cache back: same effect as --compiled-schema.
rudof shex -f binary -s user.ircache -r internal --statistics true
# shex-validate accepts -f binary the same way.
rudof shex-validate -s user.ircache -f binary --shapemap examples/user.sm examples/user.ttl
Both spellings produce byte-identical caches and behave identically –
--compile-to/--compiled-schema are dedicated shortcuts for the same
underlying ShExFormat::Binary handling, kept around because they read
more explicitly in scripts. Pick whichever fits your command better; the
“only internal is recoverable afterwards” limitation above applies the
same way regardless of which flags loaded the cache.
Other Considerations
- The cache is bound to a version. Loading a cache produced by an incompatible version fails with an error message.
- Strict-mode compiles are already verified. If the schema was
compiled under
--reader-mode strict, the negation-cycle check ran at compile time and its result is recorded in the header, so loading under strict mode does not repeat the check. - No source-schema flags.
--compiled-schemaconflicts with--schema,--schema-format, and--base-schema. The cache already carries a fully-resolved schema; there is nothing left for those flags to do. - No external-shape resolvers.
--compiled-schemaconflicts with--external-resolver. SomeEXTERNALshape declarations are substituted during AST to IR compilation, so a resolver passed at validation time would arrive too late. If your schema usesEXTERNALshapes, register the resolvers at compile time. - Semantic actions. The cache does not embed semantic-action
implementations. If your schema references a
SemActIRI, the extension that handles that IRI must be registered in the runningrudofprocess when the cache is loaded.
Full end-to-end example
Starting from the examples/
folder in the repo:
# 1. Compile the schema once.
rudof shex --schema examples/user.shex --compile-to user.ircache
# 2. Validate any number of data sets against the same cache.
rudof shex-validate \
--compiled-schema user.ircache \
--shapemap examples/user.sm \
examples/user.ttl
rudof shex-validate \
--compiled-schema user.ircache \
--shapemap examples/user.sm \
examples/user_fail.ttl