Getting started with rudof#

This chapter is a tour of the pieces of rudof that every other chapter relies on: how to install it, how a session keeps state, how errors are reported and how to configure it. The chapters that follow use RDF, SPARQL, ShEx, SHACL and the rest; this one is about the tool itself.

Installing#

rudof is a Rust library. The Python bindings are published on PyPI as pyrudof as pre-built wheels, so there is no Rust toolchain to install and nothing to compile. Python 3.10 or newer is required.

%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.

The same functionality is also available as a command line tool with binaries for Windows, Linux and macOS, and as a set of Rust crates. The Python API documentation lives at pyrudof.readthedocs.io.

The Rudof session#

The main entry point is the Rudof class.

from pyrudof import Rudof

rudof = Rudof()
rudof.get_version()
'0.3.23'

A Rudof instance is a stateful session, not a collection of pure functions. Each read_* method loads something into the session and leaves it there. This is what makes a pipeline cheap: read a graph once, then validate it, query it and serialize it without re-parsing anything.

rudof.read_data("""
prefix : <http://example.org/>

:alice :knows :bob .
""")

print(rudof.serialize_data())
@prefix : <http://example.org/> .
:alice :knows :bob .

Clearing state#

Because state persists, it also has to be cleared. reset_all() empties every slot, and there is a narrower reset_* method per slot for when you want to keep the rest.

from pyrudof import DataError

rudof.reset_all()

try:
    rudof.serialize_data()
except DataError as e:
    print(f"the session is empty again: {e}")
the session is empty again: Data error: No data loaded

Note the distinction between the narrow and the broad resets in each validation domain:

  • reset_shex_schema() drops just the schema, while reset_shex() also drops the ShapeMap, the validator and the results.

  • reset_shacl_validation() drops the validation results, while reset_shacl() also drops the shapes graph.

Using Rudof as a context manager#

A Rudof session is also a context manager, and resets itself on exit. This is the tidiest way to write a self-contained pipeline.

with Rudof() as session:
    session.read_data("prefix : <http://example.org/>\n:alice :knows :bob .")
    print(session.serialize_data())

# Outside the `with` block the session it created is gone; `rudof` above is untouched.
@prefix : <http://example.org/> .
:alice :knows :bob .

Errors#

Every failure raises an exception derived from RudofError, with one subclass per failure domain. You can therefore be as specific or as broad as you like: catch ShExError to react to a bad schema, or RudofError to react to anything rudof can go wrong at.

from pyrudof import InputError, RudofError, ShExError

with Rudof() as session:
    try:
        session.read_shex("this is not a schema")
    except ShExError as e:
        print(f"the schema itself is at fault: {type(e).__name__}")

    try:
        session.read_data("/no/such/file.ttl")
    except InputError as e:
        print(f"the input could not be resolved: {type(e).__name__}")
        # The underlying Rust error chain survives as exception notes.
        for note in getattr(e, "__notes__", []):
            print(f"  caused by: {note}")

    try:
        session.validate_shex()   # nothing was loaded
    except RudofError as e:
        print(f"caught by the base class: {type(e).__name__}")
the schema itself is at fault: ShExError
the input could not be resolved: InputError
  caused by: while processing 'input'
caught by the base class: DataError

Because the hierarchy is rooted at a single class, except RudofError is a reliable catch-all: panics inside the Rust library are converted at the language boundary into InternalError, which is also a RudofError, so nothing escapes it.

Inputs: content, paths, URLs and endpoints#

Every read_* method accepts three kinds of input, and works out which it was given: the content itself, a path to a local file, or a URL to fetch. A str that looks like a path or a URL is resolved as one; anything else is parsed as content. Passing a pathlib.Path removes the ambiguity, and is the clearer choice when you mean a file.

from pathlib import Path
from tempfile import TemporaryDirectory

with TemporaryDirectory() as tmpdir:
    path = Path(tmpdir) / "data.ttl"
    path.write_text("prefix : <http://example.org/>\n:alice :knows :bob .")

    with Rudof() as session:
        session.read_data(path)           # from a file
        print(f"read from file: {'alice' in session.serialize_data()}")

    with Rudof() as session:
        session.read_data(path.read_text())   # the same content, inline
        print(f"read inline:    {'alice' in session.serialize_data()}")
read from file: True
read inline:    True

When the input is a file or a URL, the format is inferred from the extension. The format argument overrides that, and is how you tell rudof what inline content is. With no extension to go on, read_data assumes Turtle.

RDF data can also come from a remote SPARQL endpoint instead of a local graph, using the endpoint argument. rudof ships a list of well-known endpoints that can be named rather than spelled out in full:

with Rudof() as session:
    session.read_data("prefix : <http://example.org/>\n:alice :knows :bob .")
    for name, url in sorted(session.list_endpoints()):
        print(f"{name:18} {url}")
DBpedia            https://dbpedia.org/sparql
UniProt            https://sparql.uniprot.org/sparql
Wikidata           https://query.wikidata.org/sparql
Wikidata_qlever    https://qlever.dev/api/wikidata
mardi              https://query.portal.mardi4nfdi.de/sparql
togoid             https://rdfportal.org/primary/sparql

So read_data(endpoint="wikidata") and read_data(endpoint="https://query.wikidata.org/sparql") mean the same thing. The SPARQL and ShEx chapters both query endpoints this way.

Formats are enums, not strings#

Whenever a method takes a format, it takes a member of a dedicated enum: RDFFormat, ShExFormat, ShaclFormat, ResultDataFormat, and so on. This is what makes a typo a TypeError at the call site instead of a parse failure deep inside the library.

Every enum has an all() classmethod, which is a quick way to discover what a given operation actually supports:

from pyrudof import RDFFormat, ResultDataFormat

print("input formats: ", [str(f) for f in RDFFormat.all()])
print("output formats:", [str(f) for f in ResultDataFormat.all()])
input formats:  ['Turtle', 'NTriples', 'RdfXml', 'TriG', 'N3', 'NQuads', 'JsonLd', 'Pg']
output formats: ['Turtle', 'NTriples', 'JsonLd', 'RdfXml', 'TriG', 'N3', 'NQuads', 'Compact', 'Json', 'PlantUML', 'Svg', 'Png']

Note that the input and output enums differ. ResultDataFormat includes PlantUML, Svg and Png, which rudof can write but cannot read. That asymmetry is how the diagrams in the following chapters are produced.

Enums that wrap a format with a conventional spelling also accept it as a string through from_str():

print(RDFFormat.from_str("turtle"), RDFFormat.from_str("ntriples"))
Turtle NTriples

The prefix map#

A session carries its own prefix map, used to expand the qualified names you pass to methods like node_info and read_shapemap. It can be edited directly:

with Rudof() as session:
    session.add_prefix("ex", "http://example.org/")
    session.add_prefix("foaf", "http://xmlns.com/foaf/0.1/")
    print(sorted(alias for alias, _iri in session.prefixes()))

    session.rename_prefix("foaf", "f")
    session.copy_prefix("ex", "example")
    session.remove_prefix("f")
    print(sorted(session.prefixes()))
['ex', 'foaf']
[('ex', 'http://example.org/'), ('example', 'http://example.org/')]

Configuration#

Rudof() starts from a default configuration. To customize it, pass a RudofConfig, which is usually read from a TOML file. The file is divided into sections, one per subsystem: [rdf_data] for readers and endpoints, [shex] and [shex_validator] for ShEx, [tap2shex] for the DCTAP converter, [shex2uml] for diagrams, and so on.

from pyrudof import RudofConfig

config_toml = """
[rdf_data]
base = "http://example.org/"

[shex]
check_well_formed = true
"""

with TemporaryDirectory() as tmpdir:
    config_path = Path(tmpdir) / "config.toml"
    config_path.write_text(config_toml)

    config = RudofConfig.from_path(config_path)
    session = Rudof(config)
    print(session.get_version())
0.3.23

The configuration of a live session can also be replaced with update_config(config). The DCTAP chapter puts this to real use: the DCTAP-to-ShEx converter needs a base IRI and a prefix map, and those come from [tap2shex].