Generating RDF data from shapes#
Every chapter so far has used shapes to check data. This one runs them backwards: given a ShEx or SHACL schema, rudof can generate synthetic RDF that conforms to it.
That is more useful than it first sounds. You need realistic data to benchmark a triple store, to exercise a pipeline before the real data exists, to reproduce a bug at scale, or to publish an example dataset without publishing anyone’s actual records. Writing that data by hand is tedious and, worse, tends to be unrepresentativ. A schema you already have is a much better description of what the data should look like.
Preliminaries#
%pip install -q "pyrudof>=0.3.22"
Note: you may need to restart the kernel to use updated packages.
from pathlib import Path
from tempfile import TemporaryDirectory
from pyrudof import DataGenerator, GeneratorConfig, OutputFormat
workdir = TemporaryDirectory()
WORK = Path(workdir.name)
A schema to generate from#
We start from a small ShEx schema: people, with an optional birth date, enrolled in courses.
SCHEMA = """prefix : <http://example.org/>
prefix xsd: <http://www.w3.org/2001/XMLSchema#>
:Person {
:name xsd:string ;
:birthdate xsd:date ? ;
:enrolledIn @:Course *
}
:Course {
:name xsd:string
}
"""
schema_path = WORK / "schema.shex"
schema_path.write_text(SCHEMA)
print(SCHEMA)
prefix : <http://example.org/>
prefix xsd: <http://www.w3.org/2001/XMLSchema#>
:Person {
:name xsd:string ;
:birthdate xsd:date ? ;
:enrolledIn @:Course *
}
:Course {
:name xsd:string
}
The shape of the API#
Generation has two objects:
GeneratorConfig: every knob, how many entities, which strategies, where and in what format to write, how to parallelize.DataGenerator: built from a config, and then run against a schema.
run(path) is the one-call form: it detects whether the schema is ShEx or SHACL from the
file, loads it, and generates.
config = GeneratorConfig()
config.set_entity_count(10)
config.set_output_path(WORK / "basic.ttl")
config.set_output_format(OutputFormat.Turtle)
DataGenerator(config).run(schema_path)
print((WORK / "basic.ttl").read_text()[:700], "...")
<http://example.org/Course-3> a <http://example.org/Course> ;
<http://example.org/name> "Gamma589" .
<http://example.org/Person-4> <http://example.org/birthdate> "1999-09-14"^^<http://www.w3.org/2001/XMLSchema#date> ;
a <http://example.org/Person> ;
<http://example.org/name> "Delta717" ;
<http://example.org/enrolledIn> <http://example.org/Course-4> .
<http://example.org/Course-2> a <http://example.org/Course> ;
<http://example.org/name> "Alpha510" .
<http://example.org/Course-5> a <http://example.org/Course> ;
<http://example.org/name> "Epsilon879" .
<http://example.org/Person-3> a <http://example.org/Person> ;
<http://example.org/name> "Epsilon572" .
<http://example.org/Course-1> a < ...
Note what the generator did with the schema: it created instances of both shapes, gave
every :Person a :name because the schema requires one, gave some of them a :birthdate
because it is optional, and wired :enrolledIn to actual :Course instances rather than to
invented IRIs. Shape references are followed, which is what keeps the generated graph
internally consistent.
Inspecting a configuration#
show() prints the complete resolved configuration, including every default you did not
set. It is the quickest way to find out what a knob is currently doing.
print(config.show())
GeneratorConfig { generation: GenerationConfig { entity_count: 10, seed: None, entity_distribution: Equal, cardinality_strategy: Balanced, schema_format: None, property_fill_probability: 1.0, ignore_min_cardinality: false, max_properties_per_instance: 0, property_selection_strategy: All, property_count_variance: 0.0, excluded_properties: [], type_overrides: {}, unbounded_property_values: 1, unbounded_reference_values: 1 }, field_generators: FieldGeneratorConfig { default: DefaultFieldConfig { locale: "en", quality: Medium }, datatypes: {}, properties: {} }, output: OutputConfig { path: "/tmp/tmpv8sohblm/basic.ttl", format: Turtle, compress: false, write_stats: true, parallel_writing: false, parallel_file_count: 0 }, parallel: ParallelConfig { worker_threads: None, batch_size: 100, parallel_shapes: true, parallel_fields: true } }
Statistics#
set_write_stats(True) makes the generator write a .stats.json file next to its output,
describing what it produced.
stats_config = GeneratorConfig()
stats_config.set_entity_count(20)
stats_config.set_output_path(WORK / "stats.ttl")
stats_config.set_output_format(OutputFormat.Turtle)
stats_config.set_write_stats(True)
DataGenerator(stats_config).run(schema_path)
import json
stats = json.loads((WORK / "stats.stats.json").read_text())
print(json.dumps(stats, indent=2))
{
"total_triples": 50,
"total_subjects": 20,
"total_predicates": 4,
"total_objects": 32,
"generation_time": "0ms",
"shape_counts": {
"http://example.org/Person": 10,
"http://example.org/Course": 10
},
"conformance_metrics": {
"total_generated_triples": 50,
"valid_triples": 50,
"triple_validity_percentage": 100.0,
"original_schema_constraints": 8,
"represented_constraints_in_unified": 8,
"shape_translation_loss_percentage": 0.0
}
}
Two parts are worth pointing out. shape_counts shows how the entity budget was split
between the shapes, and conformance_metrics reports how much of the schema the generator
was able to honour. shape_translation_loss_percentage above zero means some constraint in
your schema has no generation strategy behind it, and the output will not exercise it.
Cardinality strategies#
The schema says :enrolledIn @:Course * (zero or more). How many is a choice, and
CardinalityStrategy makes it:
Strategy |
Behaviour |
|---|---|
|
the fewest values the schema allows |
|
the most the schema allows |
|
a random value in range |
|
a moderate value between the two (the default) |
Minimum gives you the sparsest graph that still conforms. Maximum gives you the densest.
from pyrudof import CardinalityStrategy
CARDINALITY_SCHEMA = """
PREFIX : <http://example.org/>
PREFIX xsd: <http://www.w3.org/2001/XMLSchema#>
:User {
:name xsd:string ;
:hasFriend @:User {2,5} ;
:hasPost @:Post {0,10}
}
:Post {
:title xsd:string
}
"""
cardinality_schema = WORK / "cardinality.shex"
cardinality_schema.write_text(CARDINALITY_SCHEMA)
print(f"wrote {cardinality_schema.name}")
wrote cardinality.shex
print(f"{'strategy':<10} {'triples':>8} {'hasFriend':>10} {'hasPost':>8}")
for strategy in CardinalityStrategy.all():
name = str(strategy).lower()
output = WORK / f"card_{name}.ttl"
config = GeneratorConfig()
config.set_entity_count(10)
config.set_cardinality_strategy(strategy)
config.set_output_path(output)
config.set_output_format(OutputFormat.NTriples)
config.set_write_stats(True)
DataGenerator(config).run(cardinality_schema)
data = output.read_text()
triples = json.loads((WORK / f"card_{name}.stats.json").read_text())["total_triples"]
print(f"{name:<10} {triples:>8} {data.count('hasFriend'):>10} {data.count('hasPost'):>8}")
strategy triples hasFriend hasPost
minimum 30 10 0
maximum 70 25 25
random 55 19 16
balanced 46 16 10
The schema said each :User must have between 2 and 5 friends and may have up to 10 posts,
and the counts follow: Minimum sticks to 2 friends and no posts, Maximum pushes both to
their ceiling, and Balanced and Random land in between.
Reproducibility#
An unseeded generator produces different data on every run, which is fine for a one-off
sample and useless for a regression test. set_seed fixes the random source.
def generate(name, seed, sequential):
config = GeneratorConfig()
config.set_entity_count(6)
config.set_seed(seed)
config.set_output_path(WORK / name)
config.set_output_format(OutputFormat.NTriples)
if sequential:
config.set_parallel_shapes(False)
config.set_parallel_fields(False)
config.set_worker_threads(1)
DataGenerator(config).run(schema_path)
return sorted((WORK / name).read_text().splitlines())
same_seed_a = generate("seed_a.nt", 42, sequential=True)
same_seed_b = generate("seed_b.nt", 42, sequential=True)
other_seed = generate("seed_c.nt", 7, sequential=True)
print("seed 42 twice ->", same_seed_a == same_seed_b)
print("seed 42 vs 7 ->", same_seed_a == other_seed)
seed 42 twice -> True
seed 42 vs 7 -> False
print("\n".join(same_seed_a[:6]))
<http://example.org/Course-1> <http://example.org/name> "Alpha148" .
<http://example.org/Course-1> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
<http://example.org/Course-2> <http://example.org/name> "Epsilon658" .
<http://example.org/Course-2> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
<http://example.org/Course-3> <http://example.org/name> "Delta587" .
<http://example.org/Course-3> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
Field generators: quality and locale#
The values that fill each literal come from a field generator chosen per datatype.
Two settings steer them: set_data_quality, which trades generation speed against how
plausible the values look, and set_locale, which picks the language the generated words
are drawn from.
from pyrudof import DataQuality
print([str(q) for q in DataQuality.all()])
['Low', 'Medium', 'High']
config = GeneratorConfig()
config.set_entity_count(4)
config.set_seed(1)
config.set_data_quality(DataQuality.High)
config.set_locale("en")
config.set_output_path(WORK / "quality.ttl")
config.set_output_format(OutputFormat.NTriples)
DataGenerator(config).run(schema_path)
print((WORK / "quality.ttl").read_text()[:400])
<http://example.org/Course-2> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
<http://example.org/Course-2> <http://example.org/name> "Beta695" .
<http://example.org/Person-2> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Person> .
<http://example.org/Person-2> <http://example.org/name> "Epsilon985" .
<http://example.org/Person-2> <http://exa
Both settings live under [field_generators.default] in a saved configuration, and the
same file can override them per datatype or per property. The generated :birthdate values above
are already datatype-aware: they are valid xsd:date literals, not arbitrary text.
Output formats#
print([str(f) for f in OutputFormat.all()])
['Turtle', 'NTriples']
for fmt, suffix in [(OutputFormat.Turtle, "ttl"), (OutputFormat.NTriples, "nt")]:
config = GeneratorConfig()
config.set_entity_count(3)
config.set_seed(3)
config.set_output_path(WORK / f"formats.{suffix}")
config.set_output_format(fmt)
DataGenerator(config).run(schema_path)
print(f"===== {fmt} =====")
print((WORK / f"formats.{suffix}").read_text()[:300])
===== Turtle =====
<http://example.org/Course-2> a <http://example.org/Course> ;
<http://example.org/name> "Beta413" .
<http://example.org/Person-1> a <http://example.org/Person> ;
<http://example.org/name> "Delta660" .
<http://example.org/Course-1> a <http://example.org/Course> ;
<http://example.org/name> "Epsilon
===== NTriples =====
<http://example.org/Course-2> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
<http://example.org/Course-2> <http://example.org/name> "Beta413" .
<http://example.org/Course-1> <http://www.w3.org/1999/02/22-rdf-syntax-ns#type> <http://example.org/Course> .
<http://exam
set_compress(True) gzips the output, which matters once a dataset stops being a toy.
Generating from SHACL#
The generator reads SHACL shapes as well as ShEx. SchemaFormat names the two:
from pyrudof import SchemaFormat
print([str(f) for f in SchemaFormat.all()])
['ShEx', 'Shacl']
SHACL_SCHEMA = """
@prefix sh: <http://www.w3.org/ns/shacl#> .
@prefix : <http://example.org/> .
@prefix xsd: <http://www.w3.org/2001/XMLSchema#> .
:PersonShape a sh:NodeShape ;
sh:targetClass :Person ;
sh:property [
sh:path :name ;
sh:datatype xsd:string ;
sh:minCount 1 ; sh:maxCount 1 ;
] ;
sh:property [
sh:path :age ;
sh:datatype xsd:integer ;
sh:minCount 1 ; sh:maxCount 1 ;
] .
"""
shacl_path = WORK / "schema.shacl.ttl"
shacl_path.write_text(SHACL_SCHEMA)
print(f"wrote {shacl_path.name}")
wrote schema.shacl.ttl
Loading explicitly#
run() auto-detects the schema language, which is usually what you want.
config = GeneratorConfig()
config.set_entity_count(5)
config.set_output_path(WORK / "from_shacl.ttl")
config.set_output_format(OutputFormat.Turtle)
generator = DataGenerator(config)
generator.load_shacl_schema(shacl_path) # explicit: this is SHACL
generator.generate()
print((WORK / "from_shacl.ttl").read_text()[:400])
<http://example.org/PersonShape-4> <http://example.org/age> 6772 ;
<http://example.org/name> "Beta941" ;
a <http://example.org/Person> .
<http://example.org/PersonShape-3> <http://example.org/age> 3874 ;
<http://example.org/name> "Delta564" ;
a <http://example.org/Person> .
<http://example.org/PersonShape-1> <http://example.org/age> 3976 ;
<http://example.org/name> "Beta920" ;
a <http://exam
The three loading routes are:
Call |
Behaviour |
|---|---|
|
explicit, then |
|
detect the language, then |
|
load with a stated format and generate |
|
detect, load and generate in one call |
config = GeneratorConfig()
config.set_entity_count(5)
config.set_output_path(WORK / "with_format.ttl")
config.set_output_format(OutputFormat.Turtle)
DataGenerator(config).run_with_format(schema_path, SchemaFormat.ShEx)
print(f"generated: {(WORK / 'with_format.ttl').stat().st_size} bytes")
generated: 663 bytes
Configuration files#
A configuration built in code can be written to TOML and read back, which is how a generation setup gets version-controlled alongside the schema it belongs to.
config = GeneratorConfig()
config.set_entity_count(100)
config.set_seed(1234)
config.set_output_path(WORK / "from_config.ttl")
config.set_output_format(OutputFormat.Turtle)
config.set_cardinality_strategy(CardinalityStrategy.Balanced)
config.set_write_stats(True)
config_file = WORK / "generator.toml"
config.to_toml_file(config_file)
print(config_file.read_text())
[generation]
entity_count = 100
seed = 1234
entity_distribution = "Equal"
cardinality_strategy = "Balanced"
property_fill_probability = 1.0
ignore_min_cardinality = false
max_properties_per_instance = 0
property_selection_strategy = "All"
property_count_variance = 0.0
excluded_properties = []
unbounded_property_values = 1
unbounded_reference_values = 1
[generation.type_overrides]
[field_generators.default]
locale = "en"
quality = "Medium"
[field_generators.datatypes]
[field_generators.properties]
[output]
path = "/tmp/tmpv8sohblm/from_config.ttl"
format = "Turtle"
compress = false
write_stats = true
parallel_writing = false
parallel_file_count = 0
[parallel]
batch_size = 100
parallel_shapes = true
parallel_fields = true
GeneratorConfig.from_toml_file reads it back, and the result is an ordinary config that
can still be adjusted. from_json_file does the same for JSON.
loaded = GeneratorConfig.from_toml_file(config_file)
print("entity count from file:", loaded.get_entity_count())
print("seed from file: ", loaded.get_seed())
loaded.set_entity_count(25)
loaded.set_output_path(WORK / "from_config_modified.ttl")
print("entity count after override:", loaded.get_entity_count())
entity count from file: 100
seed from file: 1234
entity count after override: 25
validate() checks a configuration for internal consistency before you spend time
generating against it:
loaded.validate()
print("configuration is valid")
configuration is valid
Scaling up#
For large datasets, the generator can write several files in parallel instead of one:
config = GeneratorConfig()
config.set_entity_count(200)
config.set_output_path(WORK / "parallel.ttl")
config.set_output_format(OutputFormat.Turtle)
config.set_parallel_writing(True)
config.set_parallel_file_count(3)
config.set_batch_size(50)
DataGenerator(config).run(schema_path)
for path in sorted(WORK.glob("parallel*")):
print(f" {path.name}: {path.stat().st_size} bytes")
parallel.manifest.txt: 226 bytes
parallel.stats.json: 491 bytes
parallel_part_001.ttl: 9605 bytes
parallel_part_002.ttl: 9533 bytes
parallel_part_003.ttl: 9247 bytes
set_worker_threads(n) bounds the thread pool, and set_batch_size(n) controls how many
entities are held before a flush.
workdir.cleanup()
References#
rudof_generatedocumentation - the full configuration reference.