Converting from RDF to property graphs#

The two graph models are good at different things, and sooner or later you want data that was published as RDF inside a property-graph database.

rudof does that translation in one step, and this chapter is about what the step actually does: how RDF turns into tables, what the generated DDL looks like, and what validating on the way in protects you from.

Preliminaries: install and configure rudof#

%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pathlib import Path
from tempfile import TemporaryDirectory

from pyrudof import DdlDialect, Rudof

Deriving a schema from the data#

A property-graph database wants its schema declared up front, so the first question is what that schema should be. pg_db_ddl answers it by inspecting RDF data:

  • each class becomes a node table;

  • a predicate whose values are literals becomes a column on that table;

  • a predicate whose values are IRIs becomes a relationship table between two node tables.

DdlDialect lists the dialects it can emit:

DATA = """
@prefix : <http://example.org/> .

:alice a :Person ;
    :name  "Alice" ;
    :knows :bob .
:bob a :Person ;
    :name  "Bob" .
"""
print([str(d) for d in DdlDialect.all()])
['Cypher', 'Gql']
with Rudof() as session:
    print(session.pg_db_ddl(DATA, DdlDialect.Cypher))
CREATE NODE TABLE Person (id STRING, knows STRING, name STRING, PRIMARY KEY(id));
CREATE REL TABLE knows (FROM Person TO Person);

DdlDialect.Gql emits the same structure as standard GQL instead, for a database that speaks that rather than Cypher’s dialect.

Loading, and validating on the way in#

connect_pg_db opens (or creates) the database and load_pg_db fills it. Before it inserts anything, rudof validates the RDF against SHACL shapes (taken from the shapes argument, or from the data itself) so the database never receives anything its DDL could not enforce.

SHAPES = """
@prefix :   <http://example.org/> .
@prefix sh: <http://www.w3.org/ns/shacl#> .
@prefix xsd:<http://www.w3.org/2001/XMLSchema#> .

:PersonShape a sh:NodeShape ;
  sh:targetClass :Person ;
  sh:property [ sh:path :name ; sh:datatype xsd:string ;
                sh:minCount 1 ; sh:maxCount 1 ] .
"""

tmpdir = TemporaryDirectory()
session = Rudof()
session.connect_pg_db(Path(tmpdir.name) / "example.lbug")

print(session.load_pg_db(DATA, shapes=SHAPES))
Loaded 5 triples from RDF data
Loaded SHACL shapes (2 shapes)
SHACL validation PASSED ✓
Creating schema (1 node table(s), 1 relationship table(s)) and loading data...
  Created node table: Person
  Created relationship table: knows (Person → Person)
  Inserted 2 node(s)
  Inserted 1 relationship(s)
  ✓ Load complete!

That check is doing real work, because the two schemas disagree about what is optional. RDF is happy to let :carol have no :name; the generated node table has a name column and no idea that it may be missing. With the shapes in place the mismatch surfaces as a refusal rather than as a row full of nulls (the load is aborted, so nothing partial ends up in the database):

INCOMPLETE = DATA + """
:carol a :Person ;
    :knows :alice .
"""

strict = Rudof()
strict.connect_pg_db(Path(tmpdir.name) / "strict.lbug")

try:
    strict.load_pg_db(INCOMPLETE, shapes=SHAPES)
except Exception as e:
    print(f"{type(e).__name__}: {e}")

strict.reset_pg_db_connection()
PgDbError: Property Graph database error: Data does not conform to SHACL shapes; aborting load (1 violation(s))

Two practical notes on that. load_pg_db accumulates into the session’s data, so a second call sees the triples the first one read, which is why the failing load above uses its own Rudof() rather than reusing session. And skip_validation=True turns the check off entirely, which is reasonable for data you already trust and is what the Cypher chapter uses to keep its output short.

connect_pg_db also takes in_memory=True for a transient database. That connection cannot be reused across calls, though, since it does not outlive the call that created it, so a path is what you want whenever you intend to load and then query.

Checking the result#

The data is now queryable with Cypher:

result = session.query_cypher("""
    MATCH (a:Person)-[:knows]->(b:Person)
    RETURN a.name AS who, b.name AS knows
""")

print(result["columns"])
for row in result["rows"]:
    print(" ", row)
['who', 'knows']
  ['Alice', 'Bob']
session.reset_pg_db_connection()
tmpdir.cleanup()

References#

  • GQL - the ISO standard graph query language.

  • openCypher - the Cypher dialect used above.

  • LadybugDB - the embedded database rudof uses.