PG schemas#
A PG schema plays the role that ShEx and SHACL play for RDF: it says what a well-formed node or edge looks like. Unlike those two, it is usually written before the data, a property-graph database wants its schema declared up front.
This chapter writes one, reads some property-graph data, and validates the second against the first. The property graphs chapter introduces the model and the data syntax used below.
Preliminaries: install and configure rudof#
%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pyrudof import PgSchemaFormat, RDFFormat, Rudof
rudof = Rudof()
Writing a schema#
The syntax rudof reads, PgSchemaC, is close to the GQL CREATE … TYPE statements graph
databases use:
print([str(f) for f in PgSchemaFormat.all()])
['PgSchemaC']
SCHEMA = """
CREATE NODE TYPE ( AdultStudentType: Student {
name: STRING ,
age: INTEGER CHECK > 18
})
"""
rudof.read_pgschema(SCHEMA, PgSchemaFormat.PgSchemaC)
Reading that: a node type named AdultStudentType, whose nodes carry the label Student
and a record with a name string and an age integer that must be over 18. CHECK is the
constraint mechanism, the counterpart of ShEx’s value constraints.
serialize_pgschema shows how rudof understood it, which is how you check that a
constraint parsed the way you meant:
print(rudof.serialize_pgschema())
Property Graph Schema:
Node Types:
(0): Content(Label(Student), name: String(1) , age: (Integer(1) ∩ Condition((> 18))))
Edge Types:
Property graph data#
Property-graph data is one of the formats read_data accepts, as RDFFormat.Pg. A node is
written as an identifier, a set of labels in braces, and a record in brackets:
PG_DATA = """
(n1 {"Student"}["name": "Alice", "age": 23])
(n2 {"Student"}["name": "Bob", "age": 12])
(n3 {"Student"}["name": "Carol", "age": 31])
"""
rudof.read_data(PG_DATA, RDFFormat.Pg)
Typemaps#
Just as ShEx needs a shape map to say which node to check against which
shape, PG schema validation needs a typemap: a list of node: type pairs.
rudof.read_typemap("""
n1: AdultStudentType,
n2: AdultStudentType,
n3: AdultStudentType
""")
Validating#
report = rudof.validate_pgschema()
print("conforms:", report.conforms)
print("checked: ", len(report))
conforms: False
checked: 3
validate_pgschema returns a
PgSchemaValidationReport,
shaped like the ShEx and SHACL reports: conforms for the verdict, iteration for every
node that was checked, violations for the failures.
for entry in report:
print(f"{entry.node_id} as {entry.type_name}: {entry.conforms}")
n1 as AdultStudentType: True
n2 as AdultStudentType: False
n3 as AdultStudentType: True
n2 is twelve years old and the type requires age > 18, so it is the single violation.
The details field explains exactly which part of the record failed against which part of
the type:
for entry in report.violations:
print(f"{entry.node_id}:")
print(entry.details)
n2:
Record does not conform to type content:
Record:
{age: 12, name: Bob}
Types:
{age: (Integer(1) ∩ Condition((> 18))), name: String(1)}
As with the other validators, there is a rendering for humans:
from pyrudof import ResultPgSchemaValidationFormat
print([str(f) for f in ResultPgSchemaValidationFormat.all()])
print(rudof.serialize_pgschema_validation_results(ResultPgSchemaValidationFormat.Compact))
['Compact', 'Details', 'Json', 'Csv']
Result - valid?: false:
n1:AdultStudentType: true Details: Labels match Student and record: {age: 23, name: Alice} conforms to {age: (Integer(1) ∩ Condition((> 18))), name: String(1)}
n2:AdultStudentType: false Details: Record does not conform to type content:
Record:
{age: 12, name: Bob}
Types:
{age: (Integer(1) ∩ Condition((> 18))), name: String(1)}
n3:AdultStudentType: true Details: Labels match Student and record: {age: 31, name: Carol} conforms to {age: (Integer(1) ∩ Condition((> 18))), name: String(1)}
References#
PG-Schema: Schemas for Property Graphs - the paper the language comes from.