Cypher#

Cypher is the query language most property-graph databases speak, and the ancestor of the ISO GQL standard. Its central idea is that a query looks like the path it matches: (a)-[:knows]->(b) is a node, an edge and another node, drawn in ASCII.

This chapter is about the queries. If you have not met the model the queries run over, the property graphs chapter introduces it first. rudof embeds LadybugDB, so everything here runs in-process with no database to install.

Preliminaries: install and configure rudof#

%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pathlib import Path
from tempfile import TemporaryDirectory

from pyrudof import Rudof

A graph to query#

rudof’s way into the database is from RDF, in one call: load_pg_db derives a property-graph schema from RDF data and inserts the data under it. The conversion chapter is about what that step does; here it is setup, and the rest of this chapter is queries.

DATA = """
@prefix : <http://example.org/> .

:alice a :Person ; :name "Alice" ; :age 34 ; :knows :bob, :carol .
:bob   a :Person ; :name "Bob"   ; :age 27 ; :knows :carol .
:carol a :Person ; :name "Carol" ; :age 41 .
"""

connect_pg_db opens (or creates) a database at a path, and load_pg_db fills it. skip_validation=True skips the SHACL check on the way in, which keeps the output short here; the next chapter explains what leaving it off buys you.

tmpdir = TemporaryDirectory()
session = Rudof()
session.connect_pg_db(Path(tmpdir.name) / "example.lbug")

print(session.load_pg_db(DATA, skip_validation=True))
Loaded 12 triples from RDF data
SHACL validation skipped (--skip-validation): the database DDL enforces conformance
Creating schema (1 node table(s), 1 relationship table(s)) and loading data...
  Created node table: Person
  Created relationship table: knows (Person → Person)
  Inserted 3 node(s)
  Inserted 3 relationship(s)
  ✓ Load complete!

MATCH … RETURN#

The basic shape of a Cypher query is MATCH a pattern, RETURN something from it. (n:Person) means a node labelled Person, bound to the variable n; n.name reads a property out of its record.

result = session.query_cypher("""
    MATCH (n:Person)
    RETURN n.name AS name, n.age AS age
    ORDER BY n.name
""")

print(result["columns"])
for row in result["rows"]:
    print(" ", row)
['name', 'age']
  ['Alice', '34']
  ['Bob', '27']
  ['Carol', '41']

query_cypher returns a dictionary: columns names them, rows carries the data, and two timing fields report what the database spent.

print({k: v for k, v in result.items() if k.endswith("_ms")})
{'compiling_time_ms': 2.798, 'execution_time_ms': 0.893}

WHERE filters the matches, exactly as in SPARQL:

result = session.query_cypher("""
    MATCH (n:Person)
    WHERE n.age > 30
    RETURN n.name AS name
    ORDER BY name
""")

for row in result["rows"]:
    print(" ", row)
  ['Alice']
  ['Carol']

Traversals#

This is where the model earns its keep. An edge is written between the two nodes it joins, with its label in square brackets.

result = session.query_cypher("""
    MATCH (a:Person)-[:knows]->(b:Person)
    RETURN a.name AS who, b.name AS knows
    ORDER BY who, knows
""")

print(result["columns"])
for row in result["rows"]:
    print(" ", row)
['who', 'knows']
  ['Alice', 'Bob']
  ['Alice', 'Carol']
  ['Bob', 'Carol']

Aggregation works over the matches, so counting someone’s acquaintances is one line:

result = session.query_cypher("""
    MATCH (a:Person)-[:knows]->(b)
    RETURN a.name AS who, count(b) AS acquaintances
    ORDER BY who
""")

for row in result["rows"]:
    print(" ", row)
  ['Alice', 2]
  ['Bob', 1]

A *1..2 on the edge makes it a variable-length path: follow knows between one and two times.

result = session.query_cypher("""
    MATCH (a:Person)-[:knows*1..2]->(b:Person)
    RETURN a.name AS who, b.name AS reaches
    ORDER BY who, reaches
""")

for row in result["rows"]:
    print(" ", row)
  ['Alice', 'Bob']
  ['Alice', 'Carol']
  ['Alice', 'Carol']
  ['Bob', 'Carol']

Alice reaches Carol twice: directly, and through Bob. Cypher returns one row per matching path, not per pair of endpoints. Add DISTINCT if what you want is the set of people reached.

Finally, close the connection:

session.reset_pg_db_connection()
tmpdir.cleanup()

References#

  • openCypher - the specification of the dialect used above.

  • GQL - the ISO standard it fed into.

  • LadybugDB - the embedded database rudof uses.