SPARQL#

This chapter is a short introduction to SPARQL, the RDF query language, using rudof.

Preliminaries#

%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pyrudof import Rudof

rudof = Rudof()

Querying a local graph#

SPARQL is a query language for RDF, so we need some RDF first. rudof can query a graph it holds in memory or a remote SPARQL endpoint; we start local.

rudof.read_data("""
prefix :    <http://example.org/>
prefix xsd: <http://www.w3.org/2001/XMLSchema#>

:alice a :Person ;
 :name      "Alice"                ;
 :birthDate "2005-03-01"^^xsd:date ;
 :worksFor  :acme                   .
:bob a :Person   ;
 :name      "Robert Smith"         ;
 :birthDate "2003-01-02"^^xsd:date ;
 :worksFor  :acme  .
:acme a :Company ;
 :name "Acme Inc." .
""")

A SPARQL query is built around a graph pattern: a set of triple patterns in which some positions are replaced by variables, written ?name. Running the query means finding every way of substituting values for those variables so that the resulting triples are all in the graph. Each such substitution is one solution, and a SELECT query returns them as a table.

Running a query in rudof takes two steps: read_query loads it into the session, run_query executes it against the data that is loaded.

rudof.read_query("""
PREFIX : <http://example.org/>

SELECT ?person ?name ?date ?company WHERE {
  ?person a          :Person ;
          :name      ?name   ;
          :birthDate ?date   ;
          :worksFor  ?c      .
  ?c      :name      ?company .
}
""")

results = rudof.run_query()

run_query returns a QueryResults object rather than a pre-formatted string, so the bindings are reachable as data. For a SELECT query it behaves like a table: variables names the columns, len() counts the rows, and iterating yields one dictionary per solution.

print("variables:", results.variables)
print("rows:     ", len(results))

for row in results:
    print(row)
variables: ['?person', '?name', '?date', '?company']
rows:      2
{'?person': '<http://example.org/bob>', '?name': '"Robert Smith"', '?date': '"2003-01-02"^^<http://www.w3.org/2001/XMLSchema#date>', '?company': '"Acme Inc."'}
{'?person': '<http://example.org/alice>', '?name': '"Alice"', '?date': '"2005-03-01"^^<http://www.w3.org/2001/XMLSchema#date>', '?company': '"Acme Inc."'}

rows gives the same data positionally, as a list of lists aligned with variables, which is the convenient shape when you want to build a table:

print(f"{'person':<32} {'name':<16} {'company'}")
for person, name, _date, company in results.rows:
    print(f"{person:<32} {name:<16} {company}")
person                           name             company
<http://example.org/bob>         "Robert Smith"   "Acme Inc."
<http://example.org/alice>       "Alice"          "Acme Inc."

Values arrive in their RDF surface syntax: <http://example.org/alice> for an IRI, "Alice" for a plain literal, "2005-03-01"^^xsd:date for a typed one. to_python() returns the whole result set as plain Python lists and dictionaries, which is handy for feeding it straight into json.dumps or a dataframe.

import json

print(json.dumps(results.to_python()[:1], indent=2))
[
  {
    "?person": "<http://example.org/bob>",
    "?name": "\"Robert Smith\"",
    "?date": "\"2003-01-02\"^^<http://www.w3.org/2001/XMLSchema#date>",
    "?company": "\"Acme Inc.\""
  }
]

Serializing results#

If what you want is a rendered table rather than the data, serialize_query_results formats the current results. QueryResultFormat lists the available formats: the SPARQL-standard result serializations plus a few convenience ones:

from pyrudof import QueryResultFormat

print([str(f) for f in QueryResultFormat.all()])
['Internal', 'Turtle', 'NTriples', 'JsonLd', 'Json', 'RdfXml', 'Csv', 'Markdown', 'TriG', 'N3', 'NQuads']
print(rudof.serialize_query_results(QueryResultFormat.Csv))
?person,?name,?date,?company
:bob,"""Robert Smith""","""2003-01-02""^^xsd:date","""Acme Inc."""
:alice,"""Alice""","""2005-03-01""^^xsd:date","""Acme Inc."""
print(rudof.serialize_query_results())
╭───┬─────────┬────────────────┬────────────────────────┬─────────────╮
│   │ ?person │ ?name          │ ?date                  │ ?company    │
├───┼─────────┼────────────────┼────────────────────────┼─────────────┤
│ 1 │ :bob    │ "Robert Smith" │ "2003-01-02"^^xsd:date │ "Acme Inc." │
├───┼─────────┼────────────────┼────────────────────────┼─────────────┤
│ 2 │ :alice  │ "Alice"        │ "2005-03-01"^^xsd:date │ "Acme Inc." │
╰───┴─────────┴────────────────┴────────────────────────┴─────────────╯

Querying a SPARQL endpoint#

The same queries can run against a remote SPARQL endpoint. Pointing the session at one is a matter of calling read_data with an endpoint argument instead of with data; everything downstream (queries, node_info, even ShEx validation) then works against the endpoint.

Endpoints are identified by an IRI, but rudof also knows a handful of popular ones by name:

rudof.reset_all()
rudof.read_data("prefix : <http://example.org/>\n:a :b :c .")

for name, url in sorted(rudof.list_endpoints()):
    print(f"{name:18} {url}")
DBpedia            https://dbpedia.org/sparql
UniProt            https://sparql.uniprot.org/sparql
Wikidata           https://query.wikidata.org/sparql
Wikidata_qlever    https://qlever.dev/api/wikidata
mardi              https://query.portal.mardi4nfdi.de/sparql
togoid             https://rdfportal.org/primary/sparql

So endpoint="wikidata" is shorthand for endpoint="https://query.wikidata.org/sparql". The query below asks Wikidata for people (wdt:P31 wd:Q5) born in Oviedo (wdt:P19 wd:Q14317), together with their occupations.

rudof.reset_all()
rudof.read_data(endpoint="wikidata")

rudof.read_query("""
PREFIX wd:   <http://www.wikidata.org/entity/>
PREFIX wdt:  <http://www.wikidata.org/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>

SELECT ?person ?occupation WHERE {
  ?p wdt:P31  wd:Q5      ;
     wdt:P106 ?o         ;
     rdfs:label ?person  ;
     wdt:P19  wd:Q14317  .
  ?o rdfs:label ?occupation
  FILTER (lang(?person) = "en" && lang(?occupation) = "en")
}
LIMIT 10
""")

results = rudof.run_query()
print(f"{len(results)} solutions")

for row in results:
    print(f"{row['?person']:<42} {row['?occupation']}")
10 solutions
"Joaquín Fernández Prida"@en               "diplomat"@en
"Julio Somoano Rodríguez"@en               "writer"@en
"Leticia Sánchez Ruiz"@en                  "writer"@en
"Luis López-Doriga"@en                     "politician"@en
"Luis López-Doriga"@en                     "Catholic priest"@en
"Julio Somoano Rodríguez"@en               "teacher"@en
"Lalo Azcona"@en                           "businessperson"@en
"Juan Uría Ríu"@en                         "historian"@en
"Luis López-Doriga"@en                     "theologian"@en
"Juan Uría Ríu"@en                         "ethnologist"@en

The four query forms#

SELECT is only one of SPARQL’s four query forms, and QueryResults reports which one it is holding. This matters because the three shapes carry different things, and asking for the wrong one is an error rather than a silently wrong answer.

Form

Returns

Read it with

SELECT

a table of bindings

variables, rows, iteration, len()

ASK

a single boolean

.boolean

CONSTRUCT / DESCRIBE

an RDF graph

.graph

ASK#

An ASK query does not return bindings; it answers does this pattern match at all?

rudof.reset_all()
rudof.read_data("""
prefix :    <http://example.org/>
prefix xsd: <http://www.w3.org/2001/XMLSchema#>

:alice a :Person ; :name "Alice"        ; :worksFor :acme .
:bob   a :Person ; :name "Robert Smith" ; :worksFor :acme .
:acme  a :Company ; :name "Acme Inc." .
""")

rudof.read_query("""
PREFIX : <http://example.org/>
ASK { ?person :name "Alice" . }
""")

results = rudof.run_query()
print("is_boolean:", results.is_boolean)
print("answer:    ", results.boolean)
is_boolean: True
answer:     True

Because an ASK result has no rows, len() on one is a TypeError pointing you at .boolean, rather than a meaningless number. bool(results) is always defined, though, and means did this query return anything?

try:
    len(results)
except TypeError as e:
    print(f"TypeError: {e}")

print("bool(results):", bool(results))
TypeError: an ASK result has no length; use `.boolean` for the answer, or `bool(results)` to test whether the query returned anything
bool(results): True

CONSTRUCT#

A CONSTRUCT query builds a new RDF graph out of the solutions, using a template. It is how you can transform one shape of data into another; here, turning Wikidata’s representation into our own vocabulary.

rudof.reset_all()
rudof.read_data(endpoint="wikidata")

rudof.read_query("""
PREFIX wd:   <http://www.wikidata.org/entity/>
PREFIX wdt:  <http://www.wikidata.org/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
PREFIX :     <http://example.org/>

CONSTRUCT {
   ?p a          :Person     ;
      :name      ?person     ;
      :occupation ?occupation
} WHERE {
  ?p wdt:P31  wd:Q5      ;
     wdt:P106 ?o         ;
     rdfs:label ?person  ;
     wdt:P19  wd:Q14317  .
  ?o rdfs:label ?occupation
  FILTER (lang(?person) = "en" && lang(?occupation) = "en")
}
LIMIT 10
""")

results = rudof.run_query()
print("is_graph:", results.is_graph)
is_graph: True

The constructed graph is on the result object as .graph, but note it is a string, not a graph object. run_query serialized the triples to Turtle the moment the query ran, and QueryResults only carries that text:

print("type:", type(results.graph).__name__)
print("\n".join(results.graph.splitlines()[-8:]))
type: str

wd:Q5481143 a <http://example.org/Person> ;
	<http://example.org/name> "Javier Almuzara"@en ;
	<http://example.org/occupation> "writer"@en .

wd:Q5359004 a <http://example.org/Person> ;
	<http://example.org/name> "Elena Fernández"@en ;
	<http://example.org/occupation> "actor"@en .

serialize_query_results prints that same text. For a SELECT its argument chooses the serialization, but a CONSTRUCT result was already serialized when the query ran, so here the argument makes no difference. The output is Turtle whatever you pass.

Two things follow from the result being text. The constructed triples are not loaded into the session: rudof is still pointing at the Wikidata endpoint, and nothing in the data it holds was added to or replaced. And the result has no triple count and does not iterate, because nothing parsed it back. To work with the triples, read the string in as data (here into the same session), which drops the endpoint and leaves the constructed graph as the data rudof holds:

rudof.reset_all()
rudof.read_data(results.graph)

rudof.read_query("""
PREFIX : <http://example.org/>

SELECT ?name ?occupation WHERE {
  ?p :name ?name ; :occupation ?occupation .
}
LIMIT 5
""")

for row in rudof.run_query():
    print(f"{row['?name']:<42} {row['?occupation']}")
"David Cabarcos"@en                        "association football player"@en
"Alberto Rionda"@en                        "guitarist"@en
"Alberto Rionda"@en                        "composer"@en
"Carlos Tartiere de las Alas Pumariño"@en  "businessperson"@en
"Mercedes Neuschäfer-Carlón"@en            "writer"@en

Federated queries#

SPARQL’s SERVICE keyword lets a query delegate part of its pattern to another endpoint, so a single query can join data held by different publishers. The example below holds a small schedule locally (the kind of thing no public dataset knows about) and identifies each speaker by their Wikidata item. The SERVICE block asks Wikidata who those items are; the join on ?speaker puts the two halves back together.

rudof.reset_all()
rudof.read_data("""
prefix :   <http://example.org/>
prefix wd: <http://www.wikidata.org/entity/>

:talk1 :slot "09:00" ; :speaker wd:Q11934954 .
:talk2 :slot "10:00" ; :speaker wd:Q5199463  .
:talk3 :slot "11:00" ; :speaker wd:Q12387207 .
""")

rudof.read_query("""
PREFIX :     <http://example.org/>
PREFIX wd:   <http://www.wikidata.org/entity/>
PREFIX wdt:  <http://www.wikidata.org/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>

SELECT ?slot ?name ?occupation WHERE {
  ?talk :slot ?slot ; :speaker ?speaker .

  SERVICE <https://query.wikidata.org/sparql> {
    VALUES ?speaker { wd:Q11934954 wd:Q5199463 wd:Q12387207 }
    ?speaker rdfs:label ?name ;
             wdt:P106   ?o   .
    ?o rdfs:label ?occupation
    FILTER (lang(?name) = "en" && lang(?occupation) = "en")
  }
}
ORDER BY ?slot ?occupation
""")

for slot, name, occupation in rudof.run_query().rows:
    print(f"{slot:<9} {name:<34} {occupation}")
"09:00"   "Manuel González Llana"@en         "journalist"@en
"09:00"   "Manuel González Llana"@en         "politician"@en
"09:00"   "Manuel González Llana"@en         "writer"@en
"10:00"   "Alberto Rionda"@en                "composer"@en
"10:00"   "Alberto Rionda"@en                "guitarist"@en
"11:00"   "Dorotea Bárcena"@en               "actor"@en
"11:00"   "Dorotea Bárcena"@en               "theatre director"@en

References#

Learning SPARQL#

Finding endpoints#