rete · Publish & share Triple-store interop

Triple-store interop

A .rete file is not a silo. N-Quads is the interchange: rete export streams a lossless dump that every triple store bulk-loads, and rete build ingests every store's native export. This page gives the verified recipes in both directions for Oxigraph, GraphDB, and Jena/Fuseki — plus the option that skips migration entirely.

Fidelity is not taken on faith: the project's regression suite runs differentially against Oxigraph — the same data loaded into both engines must answer every query identically, on every CI run.

rete → any triple store

rete export streams the dataset with constant memory — it never loads the graph. The default format is N-Quads, lossless: default graph + named graphs, quoted triples included — written as RDF 1.2 triple terms <<( s p o )>>, which is what current parsers read (see Quoted triples below).

rete export data.rete --compress zstd > dump.nq.zst

--compress compresses the dump as it is written, streaming, so the memory bound is unaffected. zstd is the default recommendation over gzip on measured grounds: about 21% smaller output than gzip at the same nominal level, compressed 2.4x faster and decompressed 1.5x faster (numbers and method in the CLI reference). --compress gzip is there for consumers that only take gzip; piping through an external gzip still works too.

TriG is the compact lossless alternative. It is Turtle syntax wrapped in GRAPH <g> { … } blocks, so it keeps every named graph exactly as N-Quads does, while writing each subject and each namespace once instead of once per statement:

rete export data.rete --format trig > dump.trig

Both nq and trig stream, so peak memory follows --memory-budget-mb rather than the size of the graph, and both are byte-identical at any budget. TriG is the smaller of the two — substantially so on data with many statements per subject — and every store on this page loads it. Use N-Quads when the consumer is a line-oriented pipeline (split, grep, a Spark reader); use TriG when the consumer is an RDF parser and the file has to travel.

Quoted triples: two surfaces, one graph

A quoted triple — a statement standing inside another statement — has two spellings in the wild, and rete export writes either:

rete export data.rete --format nq                                  # <<( s p o )>>  RDF 1.2
rete export data.rete --format nq --quoted-triple-syntax rdf-star  # <<s p o>>      RDF-star

RDF 1.2 is the default, and the reason is the rest of this page. Oxigraph 0.5.x is built on oxrdf 0.3 / oxttl 0.2 — the RDF 1.2 generation — and its N-Quads reader rejects the RDF-star surface outright:

Error: Parser error at line 1 between columns 57 and 59:
  The object of a triple must be an IRI, a blank node or a literal

A load is atomic, so that is not "the quoted lines were skipped": it is the whole file, the same all-or-nothing failure an invalid IRI causes. In Turtle and TriG the failure is quieter and worse — an RDF 1.2 parser reads << s p o >> as a reifier, expanding it to a blank node plus an rdf:reifies statement, so the load succeeds and the graph is not the one you exported.

rete stores the RDF-star surface and its N-Quads reader takes both, so a rete → nq → rete round trip is the identity either way. Use --quoted-triple-syntax rdf-star for a consumer on the older stack (Jena's RDF-star mode, GraphDB's, anything on oxttl 0.1).

The same flag reads. rete build and rete validate take --quoted-triple-syntax with the same two values, so Turtle and TriG go both ways too:

rete build claims.ttl -o claims.rete                               # << s p o >> is a quoted triple
rete build claims.ttl -o claims.rete --quoted-triple-syntax rdf12  # << s p o >> is a REIFIER

Those two commands read the same file into different graphs, and both are right — RDF-star and RDF 1.2 assign that syntax different meanings, and nothing in the bytes says which. rdf12 additionally reads RDF 1.2's <<( s p o )>> triple terms, {| … |} annotations and "…"@lang--dir literals.

The input default is rdf-star, not rdf12, and the asymmetry with the export default is deliberate. An RDF-star file read as RDF 1.2 parses and gives you the wrong graph; an RDF 1.2 file read as RDF-star cannot parse at all, because <<( is RDF 1.2's alone — and rete names the flag in the error:

Error: claims.ttl: turtle: Parser error at line 2 column 13: ( is not a valid RDF quoted triple subject: (
hint: this input contains `<<(`, the RDF 1.2 triple-term syntax — which is what `rete export`
writes by default. Reading it takes `--quoted-triple-syntax rdf12`; the default, `rdf-star`,
reads `<< s p o >>` as a quoted triple instead.

Only one of the two mistakes is recoverable, so the default is the one that makes the other one loud. The price is that rete's own default TriG/Turtle dump needs --quoted-triple-syntax rdf12 to come back in; --format nq needs nothing, and neither does an --quoted-triple-syntax rdf-star dump.

rete as a translator between the two worlds

With a surface flag on both ends, rete build and rete export compose into a converter — the two standards' formats, through one graph model that is a superset of both:

# RDF-star Turtle  ->  RDF 1.2 TriG
rete build legacy-star.ttl -o tmp.rete --no-pyramid
rete export tmp.rete --format trig > modern-rdf12.trig

# RDF 1.2 TriG  ->  RDF-star N-Quads (for a consumer still on the older stack)
rete build modern-rdf12.trig -o tmp.rete --no-pyramid --quoted-triple-syntax rdf12
rete export tmp.rete --format nq --quoted-triple-syntax rdf-star > legacy-star.nq

Both directions are lossless for the terms involved, because both surfaces land the same stored token. What the first one does not do is turn RDF-star quoted triples into RDF 1.2 reification: << s p o >> in the input is a term, so it comes out as the term <<( s p o )>>, not as _:r rdf:reifies …. If you want reification, write it — rdf12 on input reads << s p o >>, {| … |} and an explicit rdf:reifies as exactly that, and they all become ordinary statements rete stores and SPARQL queries like any other.

Two shapes have no RDF 1.2 spelling and are refused by name on export rather than mangled: a quoted triple in subject position, and one nested in another quoted triple's subject (ttSubject ::= iri | BlankNode). Export those with --quoted-triple-syntax rdf-star.

Two things the surface cannot paper over:

  • RDF 1.2 places a triple term in object position only. A quoted triple in subject position has no RDF 1.2 spelling, so the export refuses it by name rather than writing something no parser accepts.
  • --format jsonld and --format hdt have no term kind for one at all and refuse a file that contains any.

The RDF 1.2 Turtle/TriG reader is the rdf12-turtle Cargo feature: on by default for every native build, and on in the browser engine too. In the playground's builder it is the Quoted triples: RDF 1.2 choice (RDF-star stays the default); from JavaScript it is build_with_card_syntax(text, format, cardJson, "rdf12"). oxttl 0.2 brings oxrdf 0.3, which labels anonymous blank nodes with random ids drawn through getrandom 0.3; the browser engine takes those from crypto.getRandomValues, the same source its SPARQL RAND/UUID already used. It costs the browser engine 195,624 bytes (+5.6%) and the Asyncify build 594,215 (+5.3%). scripts/check_getrandom03.sh fails the build if anything but that blank-node path comes to depend on getrandom 0.3 in a wasm build.

tests/interop/oxigraph.sh runs both surfaces against the real store, including the negative case — the rejection above is asserted, not remembered — and round-trips a dump in each surface back through rete build with the matching --quoted-triple-syntax.

The single-graph formats

Turtle and JSON-LD carry no graph term, so they serialize one graph. Which one follows a fixed ladder, and the choice is always reported on stderr:

situationwhat is exported
--graph <iri>that graph; an error naming the real graphs if it does not exist
--graph ''the default graph, explicitly
no --graph, default graph has contentthe default graph, with a note that named graphs were left out
no --graph, default graph empty, one named graphthat graph
no --graph, default graph empty, several named graphsan error listing them

Graphs are never silently merged and a non-empty selection is never silently dropped — a quads file that cannot be written as Turtle says so instead of producing a plausible-looking partial dump.

rete export data.rete --format ttl                        > default-graph.ttl
rete export data.rete --format ttl --graph http://g/1     > one-graph.ttl
rete export data.rete --format jsonld                     > default-graph.jsonld

ttl streams like nq and trig. jsonld does not — expanded JSON-LD is a single JSON array, so it is built in memory and is not the format for a file that does not fit in RAM.

HDT: queryable without decompressing

rete export data.rete --format hdt > data.hdt

Everything else on this page is text that a store has to parse before it can answer anything. HDT is the exception: a reader memory-maps the file and answers triple patterns against the mapped bytes. Opening the 1.39 GB reference file costs 50.7 MB of RSS and 0.31 s regardless of its size, and hdtSearch, rdflib-hdt and the Rust hdt crate all read it.

It is worth being equally plain about what it costs:

graphstriples only — one graph, chosen by the same ladder Turtle uses
streamingno — the graph is built in memory, so there is a size ceiling
ceilingwhichever binds first: the memory estimate, or 2^32 object ids
compressionrefused; it would remove the in-place property

Both limits are enforced before any work starts, from counts in the file header, and the refusal names the real limit, this file's numbers and the way out. The object-id cap exists because the reference implementation truncates object ids to 32 bits when it builds its index — a file above it would load and then answer queries wrongly, so rete does not produce one.

It is also not the smallest, and that is worth being blunt about. On a 1.5M-triple graph HDT is 14,217,647 bytes against trig.zst's 4,296,486 — 3.3x larger, and larger than the source .rete as well. On 88M triples it is 1,378,298,174 against 462,563,780, the same 3x. HDT is not the recommended archival or transfer format; --format trig --compress zstd wins on size and has no ceiling.

What HDT buys is the other axis. trig.zst must be decompressed and parsed in full before it answers anything; HDT answers a triple pattern against the mapped file in ~50 MB of RSS. Choose it when something will query the file, not when it will be stored or moved.

Above the ceiling, --format trig --compress zstd is the compact lossless option and has no such limit.

Prefix compression

ttl and trig abbreviate IRIs to QNames, which is where most of their size advantage over N-Quads comes from: the namespace is the repeated part of an IRI, and writing it once in an @prefix line removes it from every term that uses it.

The bindings come from two places. Well-known vocabularies (rdf, rdfs, owl, xsd, dct, skos, foaf, prov, schema, sh, void, the SPAR family, Wikibase, and the common scholarly identifier namespaces) keep their conventional names. On top of that, the exporter reads a bounded sample of the data — a hundred thousand statements, whatever the file's size — and gives a name to the namespaces that are actually frequent in this file, derived from their own last path segment. That second group is usually where the bytes are: a dataset's own entity namespace typically outnumbers every standard vocabulary in it by an order of magnitude.

An IRI is only abbreviated when the part after the namespace is a legal Turtle PN_LOCAL without backslash escaping; anything else is written in full. The result is a file that is smaller but never a file that will not parse.

--no-prefixes turns all of this off and writes every IRI in full — for a consumer that cannot resolve QNames, or to diff two dumps term by term.

Compression and the two wins together

Prefix compression and a general-purpose codec attack different redundancy, and they compose — but not additively, which is worth knowing before choosing. Turtle/TriG remove structural repetition (a subject and its namespaces written once); zstd removes byte repetition. On a 4M-quad dump:

rawzstd -3
nq626,336,88132,787,635
trig259,715,10323,308,451

TriG is 41% of N-Quads raw, but 71% of it after compression — the codec had already found most of what TriG removes structurally. TriG still wins on both, and it wins before decompression too, which is the case that matters when something has to parse the file rather than just store it.

(That table is at zstd -3, so the two formats are compared without codec time dominating. --compress zstd defaults to level 6, which is smaller than both columns shown.)

So: --format trig --compress zstd is the smallest lossless option, and --format nq is the one every line-oriented tool can split and grep.

export reads a local file. For a remote .rete, download it first (it is one GET) — harvesting through paginated CONSTRUCT works but is far slower than a dump.

Load into Oxigraph

Oxigraph's bulk loader is parallel and takes gzip directly:

# CLI (or the same subcommands via the docker image oxigraph/oxigraph):
oxigraph load --location ./store --file dump.nq.gz

# then serve it:
oxigraph serve --location ./store --bind 0.0.0.0:7878

Round-trip check (run verbatim in Docker for this page) — count in rete, count in Oxigraph, same answer, named graph intact:

$ rete sparql people.rete "SELECT (COUNT(*) AS ?n) WHERE { ?s ?p ?o }"
?n="6"^^<http://www.w3.org/2001/XMLSchema#integer>

$ curl -s "http://localhost:7878/query" \
    --data-urlencode "query=SELECT (COUNT(*) AS ?n) WHERE { ?s ?p ?o }" \
    -H "Accept: application/sparql-results+json"
{"head":{"vars":["n"]},"results":{"bindings":[{"n":{"type":"literal","value":"6","datatype":"http://www.w3.org/2001/XMLSchema#integer"}}]}}

When the dump does not load: invalid IRIs

rete's N-Triples reader is deliberately tolerant — it stores whatever sits between < and the next >. Oxigraph's is not, and this is where the two disagree. Measured against Oxigraph 0.5.9 in Docker, on a 9-quad graph of which 5 statements carry an IRI outside the N-Triples IRIREF grammar:

$ oxigraph load --location ./store --file dump.nq
Some files like Wikidata dumps contain invalid IRIs or language tags.
If you want to load them anyway use the `--lenient` option.
Error while loading file dump.nq: Parser error at line 1 between columns 1 and 42:
  Invalid IRI percent encoding '%x-'

$ oxigraph dump --location ./store --format nq | wc -l
0

Two things to take from that.

A load is atomic — one bad IRI costs the whole file. Nine quads went in and zero came out, including the four that were perfectly good. This is why the same defect cost openaire-2021-datasource an entire ~102,000-line chunk to a single IRI with no scheme: the loader rejects the chunk, not the line.

oxigraph load exits 0 even when it rejects the file. The message above goes to stderr and the process returns success, so a bulk-load script that only checks $? will report a clean import of an empty store. Check the quad count, not the exit code. (tests/interop/oxigraph.sh does exactly that.)

rete build now counts such IRIs as it ingests them, so the problem is visible where it enters rather than at the far end of someone else's loader:

warning: 5 statement(s) carry an invalid IRI (5 IRI occurrence(s)).
               2  '[' or ']' outside an IP-literal host
               1  more than one '#'
               1  '%' not followed by two hex digits
               1  a character the IRIREF grammar excludes …

For the dump itself, --sanitize-iris percent-encodes the repairable classes:

rete export data.rete --format nq --sanitize-iris > dump.nq   # now loads

The same 9-quad graph then loads and Oxigraph holds 9 quads — the count the dump carried. But this is not free, and it is opt-in for that reason:

  • the escaped IRIs are different IRIs, so the dump no longer joins against the graph it came from, and rete → Oxigraph → rete is no longer the identity;
  • an IRI with no scheme cannot be repaired at all. Escaping does nothing for it — resolving it needs a base IRI the .rete never recorded — so it is counted, named on stderr, and written verbatim. The dump is then still rejected (No scheme found in an absolute IRI, again with an empty store). Fix those at the source; nothing downstream can.

rete build --strict refuses such input outright, if you would rather find out at build time.

rete decides whether an IRI is valid with oxiri, the crate Oxigraph's own N-Triples reader validates with — so the two agree by construction rather than by our keeping a list of known-bad shapes in step. An IRI Oxigraph will refuse is one rete already counted, including shapes rete has no repair for; those are reported as unrepairable and block a sanitized export instead of passing as clean. Full rules and the classes: CLI → Invalid IRIs.

Load into GraphDB

For big files use the server-side bulk importer (importrdf), pointing at a repository config:

importrdf preload --force -c repo-config.ttl dump.nq.gz

For a running instance, the REST route (also how the workbench imports):

# create a repository, then:
curl -X POST "http://localhost:7200/repositories/myrepo/statements" \
  -H "Content-Type: application/n-quads" \
  --data-binary @dump.nq

Load into Jena / Fuseki

tdb2.tdbloader --loc ./tdb dump.nq.gz        # bulk load
fuseki-server --loc ./tdb /ds                # serve

Any triple store → rete

rete build ingests N-Triples, N-Quads, Turtle, and RDF/XML — so the reverse direction is each store's native dump piped into a build:

From Oxigraph

oxigraph dump --location ./store --file dump.nq --format nq
rete build dump.nq -o out.rete

Verified in Docker for this page, and re-verified against Oxigraph 0.5.9 by tests/interop/oxigraph.sh: dump the Oxigraph store, rebuild the .rete, and the row-level answers match the original query for query — including the named graph, which survives the full rete → Oxigraph → rete cycle. On a 6-quad dataset (3 default + 3 in one named graph) the cycle came back quad for quad identical.

That is the result on clean data, which is what this page was originally written from. It is not the result on data with an invalid IRI — see below.

One spelling does change even on clean data: an N-Triples UCHAR escape. <http://example.org/uchar/caf\u00E9> goes in and <http://example.org/uchar/café> comes back, because Oxigraph resolves the escape while rete stores the token verbatim. Same IRI, different dictionary key — worth knowing if you compare dumps byte for byte, or if a graph contains both spellings (they are one term after the round-trip and two before it).

From GraphDB

# The statements endpoint IS the dump (workbench "Export" does the same):
curl -H "Accept: application/n-quads" \
  "http://localhost:7200/repositories/myrepo/statements?infer=false" > dump.nq
rete build dump.nq -o out.rete

infer=false exports only asserted triples. If you want GraphDB's materialized inferences frozen into the file, drop it — but consider shipping the ontology instead and letting rete's OWL 2 QL reasoning answer entailments at query time.

From Jena

tdb2.tdbdump --loc ./tdb > dump.nq
rete build dump.nq -o out.rete

Scale notes

  • Hundreds of millions of triples are routine builds (data.bnf.fr: 716 M; one 726 M-triple file). If RAM is the constraint, rete build --memory-budget-mb runs the chunked external build — same byte-identical file, bounded memory. It takes the dump in the form it ships: N-Triples, N-Quads, Turtle or TriG, gzipped or not, decompressed while streaming. So a public dump.ttl.gz needs no conversion pass and no room for the expanded copy, which is usually the larger of the two costs — SemOpenAlex measures 146.8 N-Quads bytes per triple, so its 8.5 GiB author dump would land as ~400 GB of .nt before a single triple were indexed.
  • If the dump keeps its data in named graphs — TriG exports, Wikibase and GraphDB dumps — consider --collapse-graphs. In SPARQL the default graph is not the union of the named ones, so without it ?s ?p ?o answers nothing and the pyramid comes out empty. It is a modelling choice, not a build constraint: --memory-budget-mb builds named graphs directly.
  • Add --text-index at build time if you want full-text search over the migrated data, and a Dataset Card so the file explains itself.

The no-migration option: federation

If the goal is only that another engine can query .rete data, skip the dump entirely — any SPARQL 1.1 engine with SERVICE support can federate against a rete endpoint, live and lazy:

SELECT ?law ?title WHERE {
  SERVICE <https://katospiegel-rete.hf.space/sparql/boe> {
    ?law <http://data.europa.eu/eli/ontology#title> ?title .
  }
  # … joined with whatever lives in the local store …
}
LIMIT 10

The gateway turns any published .rete URL into a standard endpoint (/sparql/<full-url> — see Hosting), so this works for files nobody registered anywhere. rete serve does the same for a local file.

Comunica needs no adapter at all (verified):

$ npx -y -p @comunica/query-sparql comunica-sparql \
    "sparql@https://katospiegel-rete.hf.space/sparql/boe" \
    "SELECT ?title WHERE { <https://www.boe.es/eli/es/c/1978/12/27/(1)> <http://data.europa.eu/eli/ontology#title> ?title }"
[{"title":"\"Constitución Española.\""}]

For native (non-endpoint) integration, the npm client ships an RDF/JS ReteSource — see the JavaScript client. Migrate when you need writes, store-specific features (GraphDB's Lucene connectors, say), or co-location with data already living there; federate when you just need the answers.