Generating RDF data from shapes#

Every chapter so far has used shapes to check data. This one runs them backwards: given a ShEx or SHACL schema, rudof can generate synthetic RDF that conforms to it.

That is more useful than it first sounds. You need realistic data to benchmark a triple store, to exercise a pipeline before the real data exists, to reproduce a bug at scale, or to publish an example dataset without publishing anyone’s actual records. Writing that data by hand is tedious and, worse, tends to be unrepresentativ. A schema you already have is a much better description of what the data should look like.

Preliminaries#

%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pathlib import Path
from tempfile import TemporaryDirectory

from pyrudof import DataGenerator, GeneratorConfig, OutputFormat

workdir = TemporaryDirectory()
WORK = Path(workdir.name)

A schema to generate from#

We start from a small ShEx schema: people, with an optional birth date, enrolled in courses.

SCHEMA = """prefix :    <http://example.org/>
prefix xsd: <http://www.w3.org/2001/XMLSchema#>

:Person {
    :name       xsd:string  ;
    :birthdate  xsd:date  ? ;
    :enrolledIn @:Course  *
}

:Course {
    :name xsd:string
}
"""

schema_path = WORK / "schema.shex"
schema_path.write_text(SCHEMA)

print(SCHEMA)
prefix :    <http://example.org/>
prefix xsd: <http://www.w3.org/2001/XMLSchema#>

:Person {
    :name       xsd:string  ;
    :birthdate  xsd:date  ? ;
    :enrolledIn @:Course  *
}

:Course {
    :name xsd:string
}

The shape of the API#

Generation has two objects:

  • GeneratorConfig: every knob, how many entities, which strategies, where and in what format to write, how to parallelize.

  • DataGenerator: built from a config, and then run against a schema.

run(path) is the one-call form: it detects whether the schema is ShEx or SHACL from the file, loads it, and generates.

config = GeneratorConfig()
config.set_entity_count(10)
config.set_output_path(WORK / "basic.ttl")
config.set_output_format(OutputFormat.Turtle)

DataGenerator(config).run(schema_path)

print((WORK / "basic.ttl").read_text()[:700], "...")
<http://example.org/Course-3> a <http://example.org/Course> ;
	<http://example.org/name> "Gamma589" .
<http://example.org/Person-4> <http://example.org/birthdate> "1999-09-14"^^<http://www.w3.org/2001/XMLSchema#date> ;
	a <http://example.org/Person> ;
	<http://example.org/name> "Delta717" ;
	<http://example.org/enrolledIn> <http://example.org/Course-4> .
<http://example.org/Course-2> a <http://example.org/Course> ;
	<http://example.org/name> "Alpha510" .
<http://example.org/Course-5> a <http://example.org/Course> ;
	<http://example.org/name> "Epsilon879" .
<http://example.org/Person-3> a <http://example.org/Person> ;
	<http://example.org/name> "Epsilon572" .
<http://example.org/Course-1> a < ...

Note what the generator did with the schema: it created instances of both shapes, gave every :Person a :name because the schema requires one, gave some of them a :birthdate because it is optional, and wired :enrolledIn to actual :Course instances rather than to invented IRIs. Shape references are followed, which is what keeps the generated graph internally consistent.

Inspecting a configuration#

show() prints the complete resolved configuration, including every default you did not set. It is the quickest way to find out what a knob is currently doing.

print(config.show())
GeneratorConfig { generation: GenerationConfig { entity_count: 10, seed: None, entity_distribution: Equal, cardinality_strategy: Balanced, schema_format: None, property_fill_probability: 1.0, ignore_min_cardinality: false, max_properties_per_instance: 0, property_selection_strategy: All, property_count_variance: 0.0, excluded_properties: [], type_overrides: {}, unbounded_property_values: 1, unbounded_reference_values: 1 }, field_generators: FieldGeneratorConfig { default: DefaultFieldConfig { locale: "en", quality: Medium }, datatypes: {}, properties: {} }, output: OutputConfig { path: "/tmp/tmpv8sohblm/basic.ttl", format: Turtle, compress: false, write_stats: true, parallel_writing: false, parallel_file_count: 0 }, parallel: ParallelConfig { worker_threads: None, batch_size: 100, parallel_shapes: true, parallel_fields: true } }

Statistics#

set_write_stats(True) makes the generator write a .stats.json file next to its output, describing what it produced.

stats_config = GeneratorConfig()
stats_config.set_entity_count(20)
stats_config.set_output_path(WORK / "stats.ttl")
stats_config.set_output_format(OutputFormat.Turtle)
stats_config.set_write_stats(True)

DataGenerator(stats_config).run(schema_path)
import json

stats = json.loads((WORK / "stats.stats.json").read_text())
print(json.dumps(stats, indent=2))
{
  "total_triples": 50,
  "total_subjects": 20,
  "total_predicates": 4,
  "total_objects": 32,
  "generation_time": "0ms",
  "shape_counts": {
    "http://example.org/Person": 10,
    "http://example.org/Course": 10
  },
  "conformance_metrics": {
    "total_generated_triples": 50,
    "valid_triples": 50,
    "triple_validity_percentage": 100.0,
    "original_schema_constraints": 8,
    "represented_constraints_in_unified": 8,
    "shape_translation_loss_percentage": 0.0
  }
}

Two parts are worth pointing out. shape_counts shows how the entity budget was split between the shapes, and conformance_metrics reports how much of the schema the generator was able to honour. shape_translation_loss_percentage above zero means some constraint in your schema has no generation strategy behind it, and the output will not exercise it.

Cardinality strategies#

The schema says :enrolledIn @:Course * (zero or more). How many is a choice, and CardinalityStrategy makes it:

Strategy

Behaviour

Minimum

the fewest values the schema allows

Maximum

the most the schema allows

Random

a random value in range

Balanced

a moderate value between the two (the default)

Minimum gives you the sparsest graph that still conforms. Maximum gives you the densest.

from pyrudof import CardinalityStrategy

CARDINALITY_SCHEMA = """
PREFIX :    <http://example.org/>
PREFIX xsd: <http://www.w3.org/2001/XMLSchema#>

:User {
    :name      xsd:string    ;
    :hasFriend @:User {2,5}  ;
    :hasPost   @:Post {0,10}
}

:Post {
    :title xsd:string
}
"""

cardinality_schema = WORK / "cardinality.shex"
cardinality_schema.write_text(CARDINALITY_SCHEMA)
print(f"wrote {cardinality_schema.name}")
wrote cardinality.shex
print(f"{'strategy':<10} {'triples':>8} {'hasFriend':>10} {'hasPost':>8}")

for strategy in CardinalityStrategy.all():
    name = str(strategy).lower()
    output = WORK / f"card_{name}.ttl"

    config = GeneratorConfig()
    config.set_entity_count(10)
    config.set_cardinality_strategy(strategy)
    config.set_output_path(output)
    config.set_output_format(OutputFormat.NTriples)
    config.set_write_stats(True)

    DataGenerator(config).run(cardinality_schema)

    data = output.read_text()
    triples = json.loads((WORK / f"card_{name}.stats.json").read_text())["total_triples"]
    print(f"{name:<10} {triples:>8} {data.count('hasFriend'):>10} {data.count('hasPost'):>8}")
strategy    triples  hasFriend  hasPost
minimum          30         10        0
maximum          70         25       25
random           55         19       16
balanced         46         16       10

The schema said each :User must have between 2 and 5 friends and may have up to 10 posts, and the counts follow: Minimum sticks to 2 friends and no posts, Maximum pushes both to their ceiling, and Balanced and Random land in between.

Reproducibility#

An unseeded generator produces different data on every run, which is fine for a one-off sample and useless for a regression test. set_seed fixes the random source.

def generate(name, seed, sequential):
    config = GeneratorConfig()
    config.set_entity_count(6)
    config.set_seed(seed)
    config.set_output_path(WORK / name)
    config.set_output_format(OutputFormat.NTriples)

    if sequential:
        config.set_parallel_shapes(False)
        config.set_parallel_fields(False)
        config.set_worker_threads(1)

    DataGenerator(config).run(schema_path)
    return sorted((WORK / name).read_text().splitlines())


same_seed_a = generate("seed_a.nt", 42, sequential=True)
same_seed_b = generate("seed_b.nt", 42, sequential=True)
other_seed = generate("seed_c.nt", 7, sequential=True)

print("seed 42 twice ->", same_seed_a == same_seed_b)
print("seed 42 vs 7  ->", same_seed_a == other_seed)
seed 42 twice -> True
seed 42 vs 7  -> False
print("\n".join(same_seed_a[:6]))
<http://example.org/Course-1> <http://example.org/name> "Alpha148" .
<http://example.org/Course-1> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
<http://example.org/Course-2> <http://example.org/name> "Epsilon658" .
<http://example.org/Course-2> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
<http://example.org/Course-3> <http://example.org/name> "Delta587" .
<http://example.org/Course-3> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .

Field generators: quality and locale#

The values that fill each literal come from a field generator chosen per datatype. Two settings steer them: set_data_quality, which trades generation speed against how plausible the values look, and set_locale, which picks the language the generated words are drawn from.

from pyrudof import DataQuality

print([str(q) for q in DataQuality.all()])
['Low', 'Medium', 'High']
config = GeneratorConfig()
config.set_entity_count(4)
config.set_seed(1)
config.set_data_quality(DataQuality.High)
config.set_locale("en")
config.set_output_path(WORK / "quality.ttl")
config.set_output_format(OutputFormat.NTriples)

DataGenerator(config).run(schema_path)

print((WORK / "quality.ttl").read_text()[:400])
<http://example.org/Course-2> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
<http://example.org/Course-2> <http://example.org/name> "Beta695" .
<http://example.org/Person-2> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Person> .
<http://example.org/Person-2> <http://example.org/name> "Epsilon985" .
<http://example.org/Person-2> <http://exa

Both settings live under [field_generators.default] in a saved configuration, and the same file can override them per datatype or per property. The generated :birthdate values above are already datatype-aware: they are valid xsd:date literals, not arbitrary text.

Output formats#

print([str(f) for f in OutputFormat.all()])
['Turtle', 'NTriples']
for fmt, suffix in [(OutputFormat.Turtle, "ttl"), (OutputFormat.NTriples, "nt")]:
    config = GeneratorConfig()
    config.set_entity_count(3)
    config.set_seed(3)
    config.set_output_path(WORK / f"formats.{suffix}")
    config.set_output_format(fmt)

    DataGenerator(config).run(schema_path)

    print(f"===== {fmt} =====")
    print((WORK / f"formats.{suffix}").read_text()[:300])
===== Turtle =====
<http://example.org/Course-2> a <http://example.org/Course> ;
	<http://example.org/name> "Beta413" .
<http://example.org/Person-1> a <http://example.org/Person> ;
	<http://example.org/name> "Delta660" .
<http://example.org/Course-1> a <http://example.org/Course> ;
	<http://example.org/name> "Epsilon
===== NTriples =====
<http://example.org/Course-2> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
<http://example.org/Course-2> <http://example.org/name> "Beta413" .
<http://example.org/Course-1> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
<http://exam

set_compress(True) gzips the output, which matters once a dataset stops being a toy.

Generating from SHACL#

The generator reads SHACL shapes as well as ShEx. SchemaFormat names the two:

from pyrudof import SchemaFormat

print([str(f) for f in SchemaFormat.all()])
['ShEx', 'Shacl']
SHACL_SCHEMA = """
@prefix sh:  <http://www.w3.org/ns/shacl#> .
@prefix :    <http://example.org/> .
@prefix xsd: <http://www.w3.org/2001/XMLSchema#> .

:PersonShape a sh:NodeShape ;
    sh:targetClass :Person ;
    sh:property [
        sh:path :name ;
        sh:datatype xsd:string ;
        sh:minCount 1 ; sh:maxCount 1 ;
    ] ;
    sh:property [
        sh:path :age ;
        sh:datatype xsd:integer ;
        sh:minCount 1 ; sh:maxCount 1 ;
    ] .
"""

shacl_path = WORK / "schema.shacl.ttl"
shacl_path.write_text(SHACL_SCHEMA)
print(f"wrote {shacl_path.name}")
wrote schema.shacl.ttl

Loading explicitly#

run() auto-detects the schema language, which is usually what you want.

config = GeneratorConfig()
config.set_entity_count(5)
config.set_output_path(WORK / "from_shacl.ttl")
config.set_output_format(OutputFormat.Turtle)

generator = DataGenerator(config)
generator.load_shacl_schema(shacl_path)   # explicit: this is SHACL
generator.generate()

print((WORK / "from_shacl.ttl").read_text()[:400])
<http://example.org/PersonShape-4> <http://example.org/age> 6772 ;
	<http://example.org/name> "Beta941" ;
	a <http://example.org/Person> .
<http://example.org/PersonShape-3> <http://example.org/age> 3874 ;
	<http://example.org/name> "Delta564" ;
	a <http://example.org/Person> .
<http://example.org/PersonShape-1> <http://example.org/age> 3976 ;
	<http://example.org/name> "Beta920" ;
	a <http://exam

The three loading routes are:

Call

Behaviour

load_shex_schema(path) / load_shacl_schema(path)

explicit, then generate()

load_schema_auto(path)

detect the language, then generate()

run_with_format(path, SchemaFormat.ShEx)

load with a stated format and generate

run(path)

detect, load and generate in one call

config = GeneratorConfig()
config.set_entity_count(5)
config.set_output_path(WORK / "with_format.ttl")
config.set_output_format(OutputFormat.Turtle)

DataGenerator(config).run_with_format(schema_path, SchemaFormat.ShEx)

print(f"generated: {(WORK / 'with_format.ttl').stat().st_size} bytes")
generated: 663 bytes

Configuration files#

A configuration built in code can be written to TOML and read back, which is how a generation setup gets version-controlled alongside the schema it belongs to.

config = GeneratorConfig()
config.set_entity_count(100)
config.set_seed(1234)
config.set_output_path(WORK / "from_config.ttl")
config.set_output_format(OutputFormat.Turtle)
config.set_cardinality_strategy(CardinalityStrategy.Balanced)
config.set_write_stats(True)

config_file = WORK / "generator.toml"
config.to_toml_file(config_file)

print(config_file.read_text())
[generation]
entity_count = 100
seed = 1234
entity_distribution = "Equal"
cardinality_strategy = "Balanced"
property_fill_probability = 1.0
ignore_min_cardinality = false
max_properties_per_instance = 0
property_selection_strategy = "All"
property_count_variance = 0.0
excluded_properties = []
unbounded_property_values = 1
unbounded_reference_values = 1

[generation.type_overrides]

[field_generators.default]
locale = "en"
quality = "Medium"

[field_generators.datatypes]

[field_generators.properties]

[output]
path = "/tmp/tmpv8sohblm/from_config.ttl"
format = "Turtle"
compress = false
write_stats = true
parallel_writing = false
parallel_file_count = 0

[parallel]
batch_size = 100
parallel_shapes = true
parallel_fields = true

GeneratorConfig.from_toml_file reads it back, and the result is an ordinary config that can still be adjusted. from_json_file does the same for JSON.

loaded = GeneratorConfig.from_toml_file(config_file)
print("entity count from file:", loaded.get_entity_count())
print("seed from file:        ", loaded.get_seed())

loaded.set_entity_count(25)
loaded.set_output_path(WORK / "from_config_modified.ttl")
print("entity count after override:", loaded.get_entity_count())
entity count from file: 100
seed from file:         1234
entity count after override: 25

validate() checks a configuration for internal consistency before you spend time generating against it:

loaded.validate()
print("configuration is valid")
configuration is valid

Scaling up#

For large datasets, the generator can write several files in parallel instead of one:

config = GeneratorConfig()
config.set_entity_count(200)
config.set_output_path(WORK / "parallel.ttl")
config.set_output_format(OutputFormat.Turtle)
config.set_parallel_writing(True)
config.set_parallel_file_count(3)
config.set_batch_size(50)

DataGenerator(config).run(schema_path)

for path in sorted(WORK.glob("parallel*")):
    print(f"  {path.name}: {path.stat().st_size} bytes")
  parallel.manifest.txt: 226 bytes
  parallel.stats.json: 491 bytes
  parallel_part_001.ttl: 9605 bytes
  parallel_part_002.ttl: 9533 bytes
  parallel_part_003.ttl: 9247 bytes

set_worker_threads(n) bounds the thread pool, and set_batch_size(n) controls how many entities are held before a flush.

workdir.cleanup()

References#