Property graphs#
RDF is one graph data model; the property graph is the other. Where RDF describes everything with triples, a property graph has nodes and edges that each carry labels and a record of key/value properties. It is the model behind Neo4j, Kuzu and the GQL standard.
rudof works with both. This chapter introduces the model and the syntax; the rest of this part covers what you can do with it.
Preliminaries: install and configure rudof#
%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pyrudof import RDFFormat, ResultDataFormat, Rudof
rudof = Rudof()
Nodes, edges, labels and records#
Property-graph data is one of the formats read_data accepts, as RDFFormat.Pg. There are
exactly two kinds of line.
A node is an identifier, a set of labels in braces, and a record in brackets:
(n1 {Person} [name: "Alice", age: 23])
An edge joins two nodes, and carries a label and a record of its own:
(n1) - (e1 {knows} [since: 2020]) -> (n2)
That second line is the whole difference from RDF in one example. n1 knows n2 is a single
RDF triple with nowhere to put since: 2020; here the edge is a thing with an identity
(e1) and a record, and the annotation has an obvious home.
PG_DATA = """
(n1 {Person} [name: "Alice", age: 23])
(n2 {Person, Student} [name: "Bob"])
(n3 {Course} [name: "Algebra"])
(n1) - (e1 {knows} [since: 2020]) -> (n2)
(n2) - (e2 {enrolledIn} [start: 2024, end: 2025]) -> (n3)
"""
rudof.read_data(PG_DATA, RDFFormat.Pg)
Note n2, which carries two labels: a node is a Person and a Student at once,
without either being a class in a hierarchy. Labels are tags, not types.
serialize_data writes the graph back out. The RDF serializations do not apply to
property-graph data (asking for Turtle gets the property-graph syntax back) but
ResultDataFormat.Json gives the structural view, which is the one that shows what rudof
actually built:
print(rudof.serialize_data(ResultDataFormat.Json))
{
"edges": [
{
"id": "e1",
"labels": [
"knows"
],
"properties": {
"since": [
2020
]
},
"source": "n1",
"target": "n2"
},
{
"id": "e2",
"labels": [
"enrolledIn"
],
"properties": {
"end": [
2025
],
"start": [
2024
]
},
"source": "n2",
"target": "n3"
}
],
"nodes": [
{
"id": "n1",
"labels": [
"Person"
],
"properties": {
"age": [
23
],
"name": [
"Alice"
]
}
},
{
"id": "n2",
"labels": [
"Person",
"Student"
],
"properties": {
"name": [
"Bob"
]
}
},
{
"id": "n3",
"labels": [
"Course"
],
"properties": {
"name": [
"Algebra"
]
}
}
]
}
Every node and every edge has an id, a list of labels, and a properties map whose
values are lists. The lists matter: a property may repeat, which is how a property graph
expresses what RDF expresses by simply having two triples.
References#
GQL - the ISO standard graph query language.
YARS-PG - the serialization format the syntax above follows.
openCypher - the query language of the next chapter.