Getting started with rudof#
This chapter is a tour of the pieces of rudof that every other chapter relies on: how to install it, how a session keeps state, how errors are reported and how to configure it. The chapters that follow use RDF, SPARQL, ShEx, SHACL and the rest; this one is about the tool itself.
Installing#
rudof is a Rust library. The Python bindings are published on PyPI as
pyrudof as pre-built wheels, so there is no Rust
toolchain to install and nothing to compile. Python 3.10 or newer is required.
%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
The same functionality is also available as a command line tool with binaries for Windows, Linux and macOS, and as a set of Rust crates. The Python API documentation lives at pyrudof.readthedocs.io.
The Rudof session#
The main entry point is the
Rudof class.
from pyrudof import Rudof
rudof = Rudof()
rudof.get_version()
'0.3.23'
A Rudof instance is a stateful session, not a collection of pure functions. Each
read_* method loads something into the session and leaves it there. This is what makes a pipeline cheap: read a graph once, then validate it, query it and serialize it without re-parsing anything.
rudof.read_data("""
prefix : <http://example.org/>
:alice :knows :bob .
""")
print(rudof.serialize_data())
@prefix : <http://example.org/> .
:alice :knows :bob .
Clearing state#
Because state persists, it also has to be cleared. reset_all() empties every slot, and
there is a narrower reset_* method per slot for when you want to keep the rest.
from pyrudof import DataError
rudof.reset_all()
try:
rudof.serialize_data()
except DataError as e:
print(f"the session is empty again: {e}")
the session is empty again: Data error: No data loaded
Note the distinction between the narrow and the broad resets in each validation domain:
reset_shex_schema()drops just the schema, whilereset_shex()also drops the ShapeMap, the validator and the results.reset_shacl_validation()drops the validation results, whilereset_shacl()also drops the shapes graph.
Using Rudof as a context manager#
A Rudof session is also a context manager, and resets itself on exit. This is the
tidiest way to write a self-contained pipeline.
with Rudof() as session:
session.read_data("prefix : <http://example.org/>\n:alice :knows :bob .")
print(session.serialize_data())
# Outside the `with` block the session it created is gone; `rudof` above is untouched.
@prefix : <http://example.org/> .
:alice :knows :bob .
Errors#
Every failure raises an exception derived from
RudofError,
with one subclass per failure domain. You can therefore be as specific or as broad as you
like: catch ShExError to react to a bad schema, or RudofError to react to anything
rudof can go wrong at.
from pyrudof import InputError, RudofError, ShExError
with Rudof() as session:
try:
session.read_shex("this is not a schema")
except ShExError as e:
print(f"the schema itself is at fault: {type(e).__name__}")
try:
session.read_data("/no/such/file.ttl")
except InputError as e:
print(f"the input could not be resolved: {type(e).__name__}")
# The underlying Rust error chain survives as exception notes.
for note in getattr(e, "__notes__", []):
print(f" caused by: {note}")
try:
session.validate_shex() # nothing was loaded
except RudofError as e:
print(f"caught by the base class: {type(e).__name__}")
the schema itself is at fault: ShExError
the input could not be resolved: InputError
caused by: while processing 'input'
caught by the base class: DataError
Because the hierarchy is rooted at a single class, except RudofError is a reliable
catch-all: panics inside the Rust library are converted at the language boundary into
InternalError, which is also a RudofError, so nothing escapes it.
Inputs: content, paths, URLs and endpoints#
Every read_* method accepts three kinds of input, and works out which it was given: the
content itself, a path to a local file, or a URL to fetch. A str that looks like a path
or a URL is resolved as one; anything else is parsed as content. Passing a pathlib.Path
removes the ambiguity, and is the clearer choice when you mean a file.
from pathlib import Path
from tempfile import TemporaryDirectory
with TemporaryDirectory() as tmpdir:
path = Path(tmpdir) / "data.ttl"
path.write_text("prefix : <http://example.org/>\n:alice :knows :bob .")
with Rudof() as session:
session.read_data(path) # from a file
print(f"read from file: {'alice' in session.serialize_data()}")
with Rudof() as session:
session.read_data(path.read_text()) # the same content, inline
print(f"read inline: {'alice' in session.serialize_data()}")
read from file: True
read inline: True
When the input is a file or a URL, the format is inferred from the extension. The format
argument overrides that, and is how you tell rudof what inline content is. With no
extension to go on, read_data assumes Turtle.
RDF data can also come from a remote SPARQL endpoint instead of a local graph, using the
endpoint argument. rudof ships a list of well-known endpoints that can be named rather
than spelled out in full:
with Rudof() as session:
session.read_data("prefix : <http://example.org/>\n:alice :knows :bob .")
for name, url in sorted(session.list_endpoints()):
print(f"{name:18} {url}")
DBpedia https://dbpedia.org/sparql
UniProt https://sparql.uniprot.org/sparql
Wikidata https://query.wikidata.org/sparql
Wikidata_qlever https://qlever.dev/api/wikidata
mardi https://query.portal.mardi4nfdi.de/sparql
togoid https://rdfportal.org/primary/sparql
So read_data(endpoint="wikidata") and
read_data(endpoint="https://query.wikidata.org/sparql") mean the same thing. The
SPARQL and ShEx chapters both query endpoints this way.
Formats are enums, not strings#
Whenever a method takes a format, it takes a member of a dedicated enum:
RDFFormat, ShExFormat, ShaclFormat, ResultDataFormat, and so on. This is what makes
a typo a TypeError at the call site instead of a parse failure deep inside the library.
Every enum has an all() classmethod, which is a quick way to discover what a given
operation actually supports:
from pyrudof import RDFFormat, ResultDataFormat
print("input formats: ", [str(f) for f in RDFFormat.all()])
print("output formats:", [str(f) for f in ResultDataFormat.all()])
input formats: ['Turtle', 'NTriples', 'RdfXml', 'TriG', 'N3', 'NQuads', 'JsonLd', 'Pg']
output formats: ['Turtle', 'NTriples', 'JsonLd', 'RdfXml', 'TriG', 'N3', 'NQuads', 'Compact', 'Json', 'PlantUML', 'Svg', 'Png']
Note that the input and output enums differ. ResultDataFormat includes PlantUML,
Svg and Png, which rudof can write but cannot read. That asymmetry is how the
diagrams in the following chapters are produced.
Enums that wrap a format with a conventional spelling also accept it as a string through
from_str():
print(RDFFormat.from_str("turtle"), RDFFormat.from_str("ntriples"))
Turtle NTriples
The prefix map#
A session carries its own prefix map, used to expand the qualified names you pass to
methods like node_info and read_shapemap. It can be edited directly:
with Rudof() as session:
session.add_prefix("ex", "http://example.org/")
session.add_prefix("foaf", "http://xmlns.com/foaf/0.1/")
print(sorted(alias for alias, _iri in session.prefixes()))
session.rename_prefix("foaf", "f")
session.copy_prefix("ex", "example")
session.remove_prefix("f")
print(sorted(session.prefixes()))
['ex', 'foaf']
[('ex', 'http://example.org/'), ('example', 'http://example.org/')]
Configuration#
Rudof() starts from a default configuration. To customize it, pass a
RudofConfig,
which is usually read from a TOML file. The file is divided into sections, one per
subsystem: [rdf_data] for readers and endpoints, [shex] and [shex_validator] for ShEx,
[tap2shex] for the DCTAP converter, [shex2uml] for diagrams, and so on.
from pyrudof import RudofConfig
config_toml = """
[rdf_data]
base = "http://example.org/"
[shex]
check_well_formed = true
"""
with TemporaryDirectory() as tmpdir:
config_path = Path(tmpdir) / "config.toml"
config_path.write_text(config_toml)
config = RudofConfig.from_path(config_path)
session = Rudof(config)
print(session.get_version())
0.3.23
The configuration of a live session can also be replaced with update_config(config). The
DCTAP chapter puts this to real use: the DCTAP-to-ShEx converter needs a
base IRI and a prefix map, and those come from [tap2shex].