Service descriptions#
Before you can write a useful SPARQL query against an unfamiliar endpoint, you need to know something about it: which named graphs it holds, how big they are, which classes and properties actually occur, which SPARQL features it supports. A service description is the endpoint’s own answer to those questions, published as RDF.
This chapter shows how rudof reads one.
Preliminaries: install and configure rudof#
%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pyrudof import RDFFormat, ReaderMode, Rudof, ServiceDescriptionFormat
rudof = Rudof()
The vocabulary#
A service description is a graph rooted at an sd:Service node. The pieces that matter
most in practice:
Term |
What it says |
|---|---|
|
the URL queries are sent to |
|
the dataset queried when no |
|
the named graphs, and what is in each |
|
optional capabilities, such as |
|
|
Descriptions are usually combined with the VoID vocabulary, which is what carries the statistics: triple counts, distinct subjects, and the class and property partitions that tell you what a graph is actually made of.
Here is a minimal description, read from a string:
rudof.read_service_description("""
@prefix sd: <http://www.w3.org/ns/sparql-service-description#> .
@prefix : <http://example.org/> .
:svc a sd:Service ;
sd:endpoint <http://example.org/sparql> ;
sd:feature sd:BasicFederatedQuery ;
sd:supportedLanguage sd:SPARQL11Query ;
sd:defaultDataset [ a sd:Dataset ] .
""", RDFFormat.Turtle, None, ReaderMode.Lax)
print(rudof.serialize_service_description(ServiceDescriptionFormat.Internal))
Service
endpoint: http://example.org/sparql
supportedLanguage: [SPARQL11Query]
feature: [BasicFederatedQuery]
result_format: []
default_dataset: Dataset: f2245dc47f8b54573533e70e958e7cc1
named_graphs: []
availableGraphs:
Reading a real endpoint#
The same call accepts an endpoint URL instead of RDF content, in which case rudof fetches the description from the endpoint itself. UniProt publishes a particularly rich one.
rudof.reset_all()
service = "https://sparql.uniprot.org/sparql"
rudof.read_service_description(service, RDFFormat.Turtle, None, ReaderMode.Strict)
ServiceDescriptionFormat lists the renderings rudof can produce:
print([str(f) for f in ServiceDescriptionFormat.all()])
['Internal', 'Mie', 'Json']
Json is the structured one:
service_json = rudof.serialize_service_description(ServiceDescriptionFormat.Json)
print(service_json[:1500], "...")
{
"title": "UniProt",
"endpoint": "https://sparql.uniprot.org/sparql",
"default_dataset": {
"id": {
"Iri": "https://sparql.uniprot.org/sparql#sparql-default-dataset"
},
"default_graph": {
"id": {
"Iri": "https://sparql.uniprot.org/.well-known/void#sparql-default-graph"
},
"triples": 244341423708
},
"named_graphs": [
{
"id": {
"Iri": "https://sparql.rhea-db.org/rhea"
},
"name": "https://sparql.rhea-db.org/rhea",
"graphs": [
{
"id": {
"Iri": "https://sparql.uniprot.org/.well-known/void#_graph_rhea!a873d337"
},
"triples": 2053334,
"classes": 3,
"property_partition": [
{
"id": {
"Iri": "https://sparql.uniprot.org/.well-known/void#rhea!c74e2b73!type"
},
"property": "http://www.w3.org/1999/02/22-rdf-syntax-ns#type",
"triples": 233426
},
{
"id": {
"Iri": "https://sparql.uniprot.org/.well-known/void#rhea!819e57ec!contains6"
},
"property": "http://rdf.rhea-db.org/contains6",
"triples": 137
},
{
"id": {
"Iri": "https://sparql.uniprot.org/.well-known/void#rhea!25568ad2!isChemicallyBalanced"
},
"property": ...
Note what is in there: the endpoint’s title, its default dataset with a triple count, and its named graphs. Those numbers come from the VoID statistics the endpoint publishes, and they are the difference between guessing at a query and knowing which graph to restrict it to.
import json
description = json.loads(service_json)
print("title: ", description.get("title"))
print("endpoint:", description.get("endpoint"))
print("features:", description.get("feature"))
default_dataset = description.get("default_dataset") or {}
default_graph = default_dataset.get("default_graph") or {}
print(f"default graph: {default_graph.get('triples'):,} triples")
title: UniProt
endpoint: https://sparql.uniprot.org/sparql
features: ['BasicFederatedQuery', 'UnionDefaultGraph']
default graph: 244,341,423,708 triples
Each named graph carries one or more VoID descriptions, and those are where the counts live. Sorting the graphs by size tells you immediately where the bulk of the data is:
named = default_dataset.get("named_graphs") or []
def graph_triples(entry):
return sum(g.get("triples") or 0 for g in entry.get("graphs") or [])
print(f"{len(named)} named graphs, the largest few:\n")
for entry in sorted(named, key=graph_triples, reverse=True)[:8]:
print(f" {graph_triples(entry):>16,} {entry.get('name')}")
21 named graphs, the largest few:
198,735,586,260 http://sparql.uniprot.org/uniparc
37,329,393,214 http://sparql.uniprot.org/uniprot
5,375,029,833 http://sparql.uniprot.org/uniref
2,504,600,781 http://sparql.uniprot.org/obsolete
244,587,074 http://sparql.uniprot.org/citationmapping
70,047,936 http://sparql.uniprot.org/taxonomy
44,980,518 http://sparql.uniprot.org/proteomes
30,730,282 http://sparql.uniprot.org/citations
Better still, a VoID property partition counts the triples per predicate. That is the single most useful thing to know before writing a query against an unfamiliar graph: it tells you which properties are actually populated, rather than which ones the ontology allows.
biggest = max(named, key=graph_triples)
partitions = [
partition
for graph in biggest.get("graphs") or []
for partition in graph.get("property_partition") or []
]
print(f"most used properties in {biggest.get('name')}:\n")
for partition in sorted(partitions, key=lambda p: p.get("triples") or 0, reverse=True)[:10]:
print(f" {partition.get('triples'):>14,} {partition.get('property')}")
most used properties in http://sparql.uniprot.org/uniparc:
40,586,252,222 http://www.w3.org/1999/02/22-rdf-syntax-ns#type
14,986,819,511 http://biohackathon.org/resource/faldo#position
14,986,819,511 http://biohackathon.org/resource/faldo#reference
9,629,969,851 http://biohackathon.org/resource/faldo#begin
9,629,969,851 http://biohackathon.org/resource/faldo#end
9,629,969,851 http://purl.uniprot.org/core/signatureSequenceMatch
8,900,439,952 http://purl.uniprot.org/core/organism
8,785,485,184 http://purl.uniprot.org/core/version
8,785,485,037 http://purl.uniprot.org/core/sequenceFor
8,636,750,419 http://purl.uniprot.org/core/created
Conversion to MIE#
rudof can also emit the description in MIE format, a compact summary of what a dataset contains that is designed to be handed to an agent rather than read whole.
mie_str = rudof.serialize_service_description(ServiceDescriptionFormat.Mie)
print(mie_str[:1500], "...")
{
"schema_info": {
"title": "UniProt",
"endpoint": "https://sparql.uniprot.org/sparql",
"graphs": [
"http://sparql.uniprot.org/uniprot",
"http://sparql.uniprot.org/locations",
"http://sparql.uniprot.org/chebi",
"http://sparql.uniprot.org/keywords",
"http://sparql.uniprot.org/taxonomy",
"http://sparql.uniprot.org/journal",
"https://sparql.rhea-db.org/rhea",
"http://sparql.uniprot.org/uniref",
"http://sparql.uniprot.org/uniparc",
"http://sparql.uniprot.org/enzymes",
"http://sparql.uniprot.org/go",
"http://sparql.uniprot.org/tissues",
"http://sparql.uniprot.org/pathways",
"http://sparql.uniprot.org/citationmapping",
"http://sparql.uniprot.org/diseases",
"http://purl.uniprot.org/core",
"http://sparql.uniprot.org/citations",
"http://sparql.uniprot.org/obsolete",
"http://sparql.uniprot.org/database",
"http://sparql.uniprot.org/proteomes",
"https://sparql.uniprot.org/.well-known/sparql-examples"
]
},
"prefixes": {
"pav": "http://purl.org/pav/",
"": "http://www.w3.org/ns/sparql-service-description#",
"dcterms": "http://purl.org/dc/terms/",
"rdf": "http://www.w3.org/1999/02/22-rdf-syntax-ns#",
"void": "http://rdfs.org/ns/void#",
"formats": "http://www.w3.org/ns/formats/",
"void_ext": "http://ldf.fi/void-ext#",
"xsd": "http://www.w3.org/2001/XMLSchema#"
},
"shape_expressions": {},
"sample_rdf_entries": ...
References#
SPARQL 1.1 Service Description - the specification.
VoID - the vocabulary carrying the dataset statistics.
SPARQL - querying the endpoints described here.