PG schemas#

A PG schema plays the role that ShEx and SHACL play for RDF: it says what a well-formed node or edge looks like. Unlike those two, it is usually written before the data, a property-graph database wants its schema declared up front.

This chapter writes one, reads some property-graph data, and validates the second against the first. The property graphs chapter introduces the model and the data syntax used below.

Preliminaries: install and configure rudof#

%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pyrudof import PgSchemaFormat, RDFFormat, Rudof

rudof = Rudof()

Writing a schema#

The syntax rudof reads, PgSchemaC, is close to the GQL CREATE … TYPE statements graph databases use:

print([str(f) for f in PgSchemaFormat.all()])
['PgSchemaC']
SCHEMA = """
CREATE NODE TYPE ( AdultStudentType: Student {
    name: STRING ,
    age: INTEGER CHECK > 18
})
"""

rudof.read_pgschema(SCHEMA, PgSchemaFormat.PgSchemaC)

Reading that: a node type named AdultStudentType, whose nodes carry the label Student and a record with a name string and an age integer that must be over 18. CHECK is the constraint mechanism, the counterpart of ShEx’s value constraints.

serialize_pgschema shows how rudof understood it, which is how you check that a constraint parsed the way you meant:

print(rudof.serialize_pgschema())
Property Graph Schema:
Node Types:
  (0): Content(Label(Student), name: String(1) , age: (Integer(1) ∩ Condition((> 18))))
Edge Types:

Property graph data#

Property-graph data is one of the formats read_data accepts, as RDFFormat.Pg. A node is written as an identifier, a set of labels in braces, and a record in brackets:

PG_DATA = """
(n1      {"Student"}["name": "Alice", "age": 23])
(n2      {"Student"}["name": "Bob",   "age": 12])
(n3      {"Student"}["name": "Carol", "age": 31])
"""

rudof.read_data(PG_DATA, RDFFormat.Pg)

Typemaps#

Just as ShEx needs a shape map to say which node to check against which shape, PG schema validation needs a typemap: a list of node: type pairs.

rudof.read_typemap("""
n1: AdultStudentType,
n2: AdultStudentType,
n3: AdultStudentType
""")

Validating#

report = rudof.validate_pgschema()

print("conforms:", report.conforms)
print("checked: ", len(report))
conforms: False
checked:  3

validate_pgschema returns a PgSchemaValidationReport, shaped like the ShEx and SHACL reports: conforms for the verdict, iteration for every node that was checked, violations for the failures.

for entry in report:
    print(f"{entry.node_id} as {entry.type_name}: {entry.conforms}")
n1 as AdultStudentType: True
n2 as AdultStudentType: False
n3 as AdultStudentType: True

n2 is twelve years old and the type requires age > 18, so it is the single violation. The details field explains exactly which part of the record failed against which part of the type:

for entry in report.violations:
    print(f"{entry.node_id}:")
    print(entry.details)
n2:
Record does not conform to type content:
 Record:
{age: 12, name: Bob}
 Types:
{age: (Integer(1) ∩ Condition((> 18))), name: String(1)}

As with the other validators, there is a rendering for humans:

from pyrudof import ResultPgSchemaValidationFormat

print([str(f) for f in ResultPgSchemaValidationFormat.all()])
print(rudof.serialize_pgschema_validation_results(ResultPgSchemaValidationFormat.Compact))
['Compact', 'Details', 'Json', 'Csv']
Result - valid?: false: 
n1:AdultStudentType: true  Details:   Labels match Student and record: {age: 23, name: Alice} conforms to {age: (Integer(1) ∩ Condition((> 18))), name: String(1)}
n2:AdultStudentType: false  Details:   Record does not conform to type content:
 Record:
{age: 12, name: Bob}
 Types:
{age: (Integer(1) ∩ Condition((> 18))), name: String(1)}
n3:AdultStudentType: true  Details:   Labels match Student and record: {age: 31, name: Carol} conforms to {age: (Integer(1) ∩ Condition((> 18))), name: String(1)}

References#