Property graphs#

RDF is one graph data model; the property graph is the other. Where RDF describes everything with triples, a property graph has nodes and edges that each carry labels and a record of key/value properties. It is the model behind Neo4j, Kuzu and the GQL standard.

rudof works with both. This chapter introduces the model and the syntax; the rest of this part covers what you can do with it.

Preliminaries: install and configure rudof#

%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pyrudof import RDFFormat, ResultDataFormat, Rudof

rudof = Rudof()

Nodes, edges, labels and records#

Property-graph data is one of the formats read_data accepts, as RDFFormat.Pg. There are exactly two kinds of line.

A node is an identifier, a set of labels in braces, and a record in brackets:

(n1 {Person} [name: "Alice", age: 23])

An edge joins two nodes, and carries a label and a record of its own:

(n1) - (e1 {knows} [since: 2020]) -> (n2)

That second line is the whole difference from RDF in one example. n1 knows n2 is a single RDF triple with nowhere to put since: 2020; here the edge is a thing with an identity (e1) and a record, and the annotation has an obvious home.

PG_DATA = """
(n1 {Person} [name: "Alice", age: 23])
(n2 {Person, Student} [name: "Bob"])
(n3 {Course} [name: "Algebra"])

(n1) - (e1 {knows}       [since: 2020])            -> (n2)
(n2) - (e2 {enrolledIn}  [start: 2024, end: 2025]) -> (n3)
"""

rudof.read_data(PG_DATA, RDFFormat.Pg)

Note n2, which carries two labels: a node is a Person and a Student at once, without either being a class in a hierarchy. Labels are tags, not types.

serialize_data writes the graph back out. The RDF serializations do not apply to property-graph data (asking for Turtle gets the property-graph syntax back) but ResultDataFormat.Json gives the structural view, which is the one that shows what rudof actually built:

print(rudof.serialize_data(ResultDataFormat.Json))
{
  "edges": [
    {
      "id": "e1",
      "labels": [
        "knows"
      ],
      "properties": {
        "since": [
          2020
        ]
      },
      "source": "n1",
      "target": "n2"
    },
    {
      "id": "e2",
      "labels": [
        "enrolledIn"
      ],
      "properties": {
        "end": [
          2025
        ],
        "start": [
          2024
        ]
      },
      "source": "n2",
      "target": "n3"
    }
  ],
  "nodes": [
    {
      "id": "n1",
      "labels": [
        "Person"
      ],
      "properties": {
        "age": [
          23
        ],
        "name": [
          "Alice"
        ]
      }
    },
    {
      "id": "n2",
      "labels": [
        "Person",
        "Student"
      ],
      "properties": {
        "name": [
          "Bob"
        ]
      }
    },
    {
      "id": "n3",
      "labels": [
        "Course"
      ],
      "properties": {
        "name": [
          "Algebra"
        ]
      }
    }
  ]
}

Every node and every edge has an id, a list of labels, and a properties map whose values are lists. The lists matter: a property may repeat, which is how a property graph expresses what RDF expresses by simply having two triples.

References#

  • GQL - the ISO standard graph query language.

  • YARS-PG - the serialization format the syntax above follows.

  • openCypher - the query language of the next chapter.