SPARQL#
This chapter is a short introduction to SPARQL, the RDF query language, using rudof.
Preliminaries#
%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pyrudof import Rudof
rudof = Rudof()
Querying a local graph#
SPARQL is a query language for RDF, so we need some RDF first. rudof can query a graph it holds in memory or a remote SPARQL endpoint; we start local.
rudof.read_data("""
prefix : <http://example.org/>
prefix xsd: <http://www.w3.org/2001/XMLSchema#>
:alice a :Person ;
:name "Alice" ;
:birthDate "2005-03-01"^^xsd:date ;
:worksFor :acme .
:bob a :Person ;
:name "Robert Smith" ;
:birthDate "2003-01-02"^^xsd:date ;
:worksFor :acme .
:acme a :Company ;
:name "Acme Inc." .
""")
A SPARQL query is built around a graph pattern: a set of triple patterns in which
some positions are replaced by variables, written ?name. Running the query means finding
every way of substituting values for those variables so that the resulting triples are all
in the graph. Each such substitution is one solution, and a SELECT query returns them
as a table.
Running a query in rudof takes two steps: read_query loads it into the session,
run_query executes it against the data that is loaded.
rudof.read_query("""
PREFIX : <http://example.org/>
SELECT ?person ?name ?date ?company WHERE {
?person a :Person ;
:name ?name ;
:birthDate ?date ;
:worksFor ?c .
?c :name ?company .
}
""")
results = rudof.run_query()
run_query returns a
QueryResults
object rather than a pre-formatted string, so the bindings are reachable as data. For a
SELECT query it behaves like a table: variables names the columns, len() counts the
rows, and iterating yields one dictionary per solution.
print("variables:", results.variables)
print("rows: ", len(results))
for row in results:
print(row)
variables: ['?person', '?name', '?date', '?company']
rows: 2
{'?person': '<http://example.org/bob>', '?name': '"Robert Smith"', '?date': '"2003-01-02"^^<http://www.w3.org/2001/XMLSchema#date>', '?company': '"Acme Inc."'}
{'?person': '<http://example.org/alice>', '?name': '"Alice"', '?date': '"2005-03-01"^^<http://www.w3.org/2001/XMLSchema#date>', '?company': '"Acme Inc."'}
rows gives the same data positionally, as a list of lists aligned with variables, which
is the convenient shape when you want to build a table:
print(f"{'person':<32} {'name':<16} {'company'}")
for person, name, _date, company in results.rows:
print(f"{person:<32} {name:<16} {company}")
person name company
<http://example.org/bob> "Robert Smith" "Acme Inc."
<http://example.org/alice> "Alice" "Acme Inc."
Values arrive in their RDF surface syntax: <http://example.org/alice> for an IRI,
"Alice" for a plain literal, "2005-03-01"^^xsd:date for a typed one. to_python() returns the
whole result set as plain Python lists and dictionaries, which is handy for feeding it
straight into json.dumps or a dataframe.
import json
print(json.dumps(results.to_python()[:1], indent=2))
[
{
"?person": "<http://example.org/bob>",
"?name": "\"Robert Smith\"",
"?date": "\"2003-01-02\"^^<http://www.w3.org/2001/XMLSchema#date>",
"?company": "\"Acme Inc.\""
}
]
Serializing results#
If what you want is a rendered table rather than the data, serialize_query_results
formats the current results. QueryResultFormat lists the available formats: the
SPARQL-standard result serializations plus a few convenience ones:
from pyrudof import QueryResultFormat
print([str(f) for f in QueryResultFormat.all()])
['Internal', 'Turtle', 'NTriples', 'JsonLd', 'Json', 'RdfXml', 'Csv', 'Markdown', 'TriG', 'N3', 'NQuads']
print(rudof.serialize_query_results(QueryResultFormat.Csv))
?person,?name,?date,?company
:bob,"""Robert Smith""","""2003-01-02""^^xsd:date","""Acme Inc."""
:alice,"""Alice""","""2005-03-01""^^xsd:date","""Acme Inc."""
print(rudof.serialize_query_results())
╭───┬─────────┬────────────────┬────────────────────────┬─────────────╮
│ │ ?person │ ?name │ ?date │ ?company │
├───┼─────────┼────────────────┼────────────────────────┼─────────────┤
│ 1 │ :bob │ "Robert Smith" │ "2003-01-02"^^xsd:date │ "Acme Inc." │
├───┼─────────┼────────────────┼────────────────────────┼─────────────┤
│ 2 │ :alice │ "Alice" │ "2005-03-01"^^xsd:date │ "Acme Inc." │
╰───┴─────────┴────────────────┴────────────────────────┴─────────────╯
Querying a SPARQL endpoint#
The same queries can run against a remote SPARQL
endpoint. Pointing the session at one is a
matter of calling read_data with an endpoint argument instead of with data; everything
downstream (queries, node_info, even ShEx validation) then works against the endpoint.
Endpoints are identified by an IRI, but rudof also knows a handful of popular ones by name:
rudof.reset_all()
rudof.read_data("prefix : <http://example.org/>\n:a :b :c .")
for name, url in sorted(rudof.list_endpoints()):
print(f"{name:18} {url}")
DBpedia https://dbpedia.org/sparql
UniProt https://sparql.uniprot.org/sparql
Wikidata https://query.wikidata.org/sparql
Wikidata_qlever https://qlever.dev/api/wikidata
mardi https://query.portal.mardi4nfdi.de/sparql
togoid https://rdfportal.org/primary/sparql
So endpoint="wikidata" is shorthand for
endpoint="https://query.wikidata.org/sparql". The query below asks
Wikidata for people (wdt:P31 wd:Q5) born in Oviedo
(wdt:P19 wd:Q14317), together with their occupations.
rudof.reset_all()
rudof.read_data(endpoint="wikidata")
rudof.read_query("""
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
SELECT ?person ?occupation WHERE {
?p wdt:P31 wd:Q5 ;
wdt:P106 ?o ;
rdfs:label ?person ;
wdt:P19 wd:Q14317 .
?o rdfs:label ?occupation
FILTER (lang(?person) = "en" && lang(?occupation) = "en")
}
LIMIT 10
""")
results = rudof.run_query()
print(f"{len(results)} solutions")
for row in results:
print(f"{row['?person']:<42} {row['?occupation']}")
10 solutions
"Joaquín Fernández Prida"@en "diplomat"@en
"Julio Somoano Rodríguez"@en "writer"@en
"Leticia Sánchez Ruiz"@en "writer"@en
"Luis López-Doriga"@en "politician"@en
"Luis López-Doriga"@en "Catholic priest"@en
"Julio Somoano Rodríguez"@en "teacher"@en
"Lalo Azcona"@en "businessperson"@en
"Juan Uría Ríu"@en "historian"@en
"Luis López-Doriga"@en "theologian"@en
"Juan Uría Ríu"@en "ethnologist"@en
The four query forms#
SELECT is only one of SPARQL’s four query forms, and QueryResults reports which one it
is holding. This matters because the three shapes carry different things, and asking for
the wrong one is an error rather than a silently wrong answer.
Form |
Returns |
Read it with |
|---|---|---|
|
a table of bindings |
|
|
a single boolean |
|
|
an RDF graph |
|
ASK#
An ASK query does not return bindings; it answers does this pattern match at all?
rudof.reset_all()
rudof.read_data("""
prefix : <http://example.org/>
prefix xsd: <http://www.w3.org/2001/XMLSchema#>
:alice a :Person ; :name "Alice" ; :worksFor :acme .
:bob a :Person ; :name "Robert Smith" ; :worksFor :acme .
:acme a :Company ; :name "Acme Inc." .
""")
rudof.read_query("""
PREFIX : <http://example.org/>
ASK { ?person :name "Alice" . }
""")
results = rudof.run_query()
print("is_boolean:", results.is_boolean)
print("answer: ", results.boolean)
is_boolean: True
answer: True
Because an ASK result has no rows, len() on one is a TypeError pointing you at
.boolean, rather than a meaningless number. bool(results) is always defined, though,
and means did this query return anything?
try:
len(results)
except TypeError as e:
print(f"TypeError: {e}")
print("bool(results):", bool(results))
TypeError: an ASK result has no length; use `.boolean` for the answer, or `bool(results)` to test whether the query returned anything
bool(results): True
CONSTRUCT#
A CONSTRUCT query builds a new RDF graph out of the solutions, using a template. It is
how you can transform one shape of data into another; here, turning Wikidata’s
representation into our own vocabulary.
rudof.reset_all()
rudof.read_data(endpoint="wikidata")
rudof.read_query("""
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
PREFIX : <http://example.org/>
CONSTRUCT {
?p a :Person ;
:name ?person ;
:occupation ?occupation
} WHERE {
?p wdt:P31 wd:Q5 ;
wdt:P106 ?o ;
rdfs:label ?person ;
wdt:P19 wd:Q14317 .
?o rdfs:label ?occupation
FILTER (lang(?person) = "en" && lang(?occupation) = "en")
}
LIMIT 10
""")
results = rudof.run_query()
print("is_graph:", results.is_graph)
is_graph: True
The constructed graph is on the result object as .graph, but note it is a string, not
a graph object. run_query serialized the triples to Turtle the moment the query ran, and
QueryResults only carries that text:
print("type:", type(results.graph).__name__)
print("\n".join(results.graph.splitlines()[-8:]))
type: str
wd:Q5481143 a <http://example.org/Person> ;
<http://example.org/name> "Javier Almuzara"@en ;
<http://example.org/occupation> "writer"@en .
wd:Q5359004 a <http://example.org/Person> ;
<http://example.org/name> "Elena Fernández"@en ;
<http://example.org/occupation> "actor"@en .
serialize_query_results prints that same text. For a SELECT its argument chooses the
serialization, but a CONSTRUCT result was already serialized when the query ran, so here
the argument makes no difference. The output is Turtle whatever you pass.
Two things follow from the result being text. The constructed triples are not loaded
into the session: rudof is still pointing at the Wikidata endpoint, and nothing in the
data it holds was added to or replaced. And the result has no triple count and does not
iterate, because nothing parsed it back. To work with the triples, read the string in as
data (here into the same session), which drops the endpoint and leaves the constructed
graph as the data rudof holds:
rudof.reset_all()
rudof.read_data(results.graph)
rudof.read_query("""
PREFIX : <http://example.org/>
SELECT ?name ?occupation WHERE {
?p :name ?name ; :occupation ?occupation .
}
LIMIT 5
""")
for row in rudof.run_query():
print(f"{row['?name']:<42} {row['?occupation']}")
"David Cabarcos"@en "association football player"@en
"Alberto Rionda"@en "guitarist"@en
"Alberto Rionda"@en "composer"@en
"Carlos Tartiere de las Alas Pumariño"@en "businessperson"@en
"Mercedes Neuschäfer-Carlón"@en "writer"@en
Federated queries#
SPARQL’s SERVICE keyword lets a query delegate part of its pattern to another endpoint,
so a single query can join data held by different publishers. The example below holds a
small schedule locally (the kind of thing no public dataset knows about) and identifies
each speaker by their Wikidata item. The SERVICE block asks Wikidata who those items
are; the join on ?speaker puts the two halves back together.
rudof.reset_all()
rudof.read_data("""
prefix : <http://example.org/>
prefix wd: <http://www.wikidata.org/entity/>
:talk1 :slot "09:00" ; :speaker wd:Q11934954 .
:talk2 :slot "10:00" ; :speaker wd:Q5199463 .
:talk3 :slot "11:00" ; :speaker wd:Q12387207 .
""")
rudof.read_query("""
PREFIX : <http://example.org/>
PREFIX wd: <http://www.wikidata.org/entity/>
PREFIX wdt: <http://www.wikidata.org/prop/direct/>
PREFIX rdfs: <http://www.w3.org/2000/01/rdf-schema#>
SELECT ?slot ?name ?occupation WHERE {
?talk :slot ?slot ; :speaker ?speaker .
SERVICE <https://query.wikidata.org/sparql> {
VALUES ?speaker { wd:Q11934954 wd:Q5199463 wd:Q12387207 }
?speaker rdfs:label ?name ;
wdt:P106 ?o .
?o rdfs:label ?occupation
FILTER (lang(?name) = "en" && lang(?occupation) = "en")
}
}
ORDER BY ?slot ?occupation
""")
for slot, name, occupation in rudof.run_query().rows:
print(f"{slot:<9} {name:<34} {occupation}")
"09:00" "Manuel González Llana"@en "journalist"@en
"09:00" "Manuel González Llana"@en "politician"@en
"09:00" "Manuel González Llana"@en "writer"@en
"10:00" "Alberto Rionda"@en "composer"@en
"10:00" "Alberto Rionda"@en "guitarist"@en
"11:00" "Dorotea Bárcena"@en "actor"@en
"11:00" "Dorotea Bárcena"@en "theatre director"@en
References#
Learning SPARQL#
SPARQL 1.1 Query Language - the specification.
Finding endpoints#
SPARQL endpoints collected in Wikidata.
SPARQL endpoints collected at W3C.