Service descriptions#

Before you can write a useful SPARQL query against an unfamiliar endpoint, you need to know something about it: which named graphs it holds, how big they are, which classes and properties actually occur, which SPARQL features it supports. A service description is the endpoint’s own answer to those questions, published as RDF.

This chapter shows how rudof reads one.

Preliminaries: install and configure rudof#

%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pyrudof import RDFFormat, ReaderMode, Rudof, ServiceDescriptionFormat

rudof = Rudof()

The vocabulary#

A service description is a graph rooted at an sd:Service node. The pieces that matter most in practice:

Term

What it says

sd:endpoint

the URL queries are sent to

sd:defaultDataset

the dataset queried when no GRAPH is named

sd:namedGraph / sd:graph

the named graphs, and what is in each

sd:feature

optional capabilities, such as sd:BasicFederatedQuery

sd:supportedLanguage

sd:SPARQL11Query, sd:SPARQL11Update, …

Descriptions are usually combined with the VoID vocabulary, which is what carries the statistics: triple counts, distinct subjects, and the class and property partitions that tell you what a graph is actually made of.

Here is a minimal description, read from a string:

rudof.read_service_description("""
@prefix sd: <http://www.w3.org/ns/sparql-service-description#> .
@prefix :   <http://example.org/> .

:svc a sd:Service ;
  sd:endpoint <http://example.org/sparql> ;
  sd:feature sd:BasicFederatedQuery ;
  sd:supportedLanguage sd:SPARQL11Query ;
  sd:defaultDataset [ a sd:Dataset ] .
""", RDFFormat.Turtle, None, ReaderMode.Lax)

print(rudof.serialize_service_description(ServiceDescriptionFormat.Internal))
Service
 endpoint: http://example.org/sparql
  supportedLanguage: [SPARQL11Query]
  feature: [BasicFederatedQuery]
  result_format: []
  default_dataset: Dataset: f2245dc47f8b54573533e70e958e7cc1
 named_graphs: []

  availableGraphs: 

Reading a real endpoint#

The same call accepts an endpoint URL instead of RDF content, in which case rudof fetches the description from the endpoint itself. UniProt publishes a particularly rich one.

rudof.reset_all()

service = "https://sparql.uniprot.org/sparql"
rudof.read_service_description(service, RDFFormat.Turtle, None, ReaderMode.Strict)

ServiceDescriptionFormat lists the renderings rudof can produce:

print([str(f) for f in ServiceDescriptionFormat.all()])
['Internal', 'Mie', 'Json']

Json is the structured one:

service_json = rudof.serialize_service_description(ServiceDescriptionFormat.Json)
print(service_json[:1500], "...")
{
  "title": "UniProt",
  "endpoint": "https://sparql.uniprot.org/sparql",
  "default_dataset": {
    "id": {
      "Iri": "https://sparql.uniprot.org/sparql#sparql-default-dataset"
    },
    "default_graph": {
      "id": {
        "Iri": "https://sparql.uniprot.org/.well-known/void#sparql-default-graph"
      },
      "triples": 244341423708
    },
    "named_graphs": [
      {
        "id": {
          "Iri": "https://sparql.rhea-db.org/rhea"
        },
        "name": "https://sparql.rhea-db.org/rhea",
        "graphs": [
          {
            "id": {
              "Iri": "https://sparql.uniprot.org/.well-known/void#_graph_rhea!a873d337"
            },
            "triples": 2053334,
            "classes": 3,
            "property_partition": [
              {
                "id": {
                  "Iri": "https://sparql.uniprot.org/.well-known/void#rhea!c74e2b73!type"
                },
                "property": "http://www.w3.org/1999/02/22-rdf-syntax-ns#type",
                "triples": 233426
              },
              {
                "id": {
                  "Iri": "https://sparql.uniprot.org/.well-known/void#rhea!819e57ec!contains6"
                },
                "property": "http://rdf.rhea-db.org/contains6",
                "triples": 137
              },
              {
                "id": {
                  "Iri": "https://sparql.uniprot.org/.well-known/void#rhea!25568ad2!isChemicallyBalanced"
                },
                "property":  ...

Note what is in there: the endpoint’s title, its default dataset with a triple count, and its named graphs. Those numbers come from the VoID statistics the endpoint publishes, and they are the difference between guessing at a query and knowing which graph to restrict it to.

import json

description = json.loads(service_json)

print("title:   ", description.get("title"))
print("endpoint:", description.get("endpoint"))
print("features:", description.get("feature"))

default_dataset = description.get("default_dataset") or {}
default_graph = default_dataset.get("default_graph") or {}
print(f"default graph: {default_graph.get('triples'):,} triples")
title:    UniProt
endpoint: https://sparql.uniprot.org/sparql
features: ['BasicFederatedQuery', 'UnionDefaultGraph']
default graph: 244,341,423,708 triples

Each named graph carries one or more VoID descriptions, and those are where the counts live. Sorting the graphs by size tells you immediately where the bulk of the data is:

named = default_dataset.get("named_graphs") or []


def graph_triples(entry):
    return sum(g.get("triples") or 0 for g in entry.get("graphs") or [])


print(f"{len(named)} named graphs, the largest few:\n")
for entry in sorted(named, key=graph_triples, reverse=True)[:8]:
    print(f"  {graph_triples(entry):>16,}  {entry.get('name')}")
21 named graphs, the largest few:

   198,735,586,260  http://sparql.uniprot.org/uniparc
    37,329,393,214  http://sparql.uniprot.org/uniprot
     5,375,029,833  http://sparql.uniprot.org/uniref
     2,504,600,781  http://sparql.uniprot.org/obsolete
       244,587,074  http://sparql.uniprot.org/citationmapping
        70,047,936  http://sparql.uniprot.org/taxonomy
        44,980,518  http://sparql.uniprot.org/proteomes
        30,730,282  http://sparql.uniprot.org/citations

Better still, a VoID property partition counts the triples per predicate. That is the single most useful thing to know before writing a query against an unfamiliar graph: it tells you which properties are actually populated, rather than which ones the ontology allows.

biggest = max(named, key=graph_triples)
partitions = [
    partition
    for graph in biggest.get("graphs") or []
    for partition in graph.get("property_partition") or []
]

print(f"most used properties in {biggest.get('name')}:\n")
for partition in sorted(partitions, key=lambda p: p.get("triples") or 0, reverse=True)[:10]:
    print(f"  {partition.get('triples'):>14,}  {partition.get('property')}")
most used properties in http://sparql.uniprot.org/uniparc:

  40,586,252,222  http://www.w3.org/1999/02/22-rdf-syntax-ns#type
  14,986,819,511  http://biohackathon.org/resource/faldo#position
  14,986,819,511  http://biohackathon.org/resource/faldo#reference
   9,629,969,851  http://biohackathon.org/resource/faldo#begin
   9,629,969,851  http://biohackathon.org/resource/faldo#end
   9,629,969,851  http://purl.uniprot.org/core/signatureSequenceMatch
   8,900,439,952  http://purl.uniprot.org/core/organism
   8,785,485,184  http://purl.uniprot.org/core/version
   8,785,485,037  http://purl.uniprot.org/core/sequenceFor
   8,636,750,419  http://purl.uniprot.org/core/created

Conversion to MIE#

rudof can also emit the description in MIE format, a compact summary of what a dataset contains that is designed to be handed to an agent rather than read whole.

mie_str = rudof.serialize_service_description(ServiceDescriptionFormat.Mie)
print(mie_str[:1500], "...")
{
  "schema_info": {
    "title": "UniProt",
    "endpoint": "https://sparql.uniprot.org/sparql",
    "graphs": [
      "http://sparql.uniprot.org/uniprot",
      "http://sparql.uniprot.org/locations",
      "http://sparql.uniprot.org/chebi",
      "http://sparql.uniprot.org/keywords",
      "http://sparql.uniprot.org/taxonomy",
      "http://sparql.uniprot.org/journal",
      "https://sparql.rhea-db.org/rhea",
      "http://sparql.uniprot.org/uniref",
      "http://sparql.uniprot.org/uniparc",
      "http://sparql.uniprot.org/enzymes",
      "http://sparql.uniprot.org/go",
      "http://sparql.uniprot.org/tissues",
      "http://sparql.uniprot.org/pathways",
      "http://sparql.uniprot.org/citationmapping",
      "http://sparql.uniprot.org/diseases",
      "http://purl.uniprot.org/core",
      "http://sparql.uniprot.org/citations",
      "http://sparql.uniprot.org/obsolete",
      "http://sparql.uniprot.org/database",
      "http://sparql.uniprot.org/proteomes",
      "https://sparql.uniprot.org/.well-known/sparql-examples"
    ]
  },
  "prefixes": {
    "pav": "http://purl.org/pav/",
    "": "http://www.w3.org/ns/sparql-service-description#",
    "dcterms": "http://purl.org/dc/terms/",
    "rdf": "http://www.w3.org/1999/02/22-rdf-syntax-ns#",
    "void": "http://rdfs.org/ns/void#",
    "formats": "http://www.w3.org/ns/formats/",
    "void_ext": "http://ldf.fi/void-ext#",
    "xsd": "http://www.w3.org/2001/XMLSchema#"
  },
  "shape_expressions": {},
  "sample_rdf_entries":  ...

References#