# RDFLib: a pure Python RDF graph library, from Turtle parsing to SPARQL 1.1

> RDFLib is a pure Python package for parsing, serializing and querying RDF, and it is the default choice for Python projects that need to handle linked data. Its command line tools, store plugins and SPARQL 1.1 support are the parts worth checking before you commit to it.

**RDFLib/rdflib** — RDFLib is a Python library for working with RDF, a simple yet powerful language for representing information.

- Repository: https://github.com/RDFLib/rdflib
- Website: https://rdflib.readthedocs.org
- Stars: 2,521 · Forks: 611
- Language: Python
- License: BSD-3-Clause
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/rdflib-rdflib

## What RDFLib solves, and who ends up using it

RDF is a data model where every statement is a subject, predicate, object triple, and the components are URIs or literals. That model shows up in places where data has to be merged across organisations: library catalogues, government open data, schema.org markup, knowledge graphs published as linked data. The friction is that RDF arrives in several serializations (RDF/XML, Turtle, N3, NTriples, N-Quads, TriX, Trig, JSON-LD, HexTuples), and most languages do not ship a parser for any of them. RDFLib's README states that it contains parsers and serializers for all of those formats, a Graph interface backed by pluggable stores, and a SPARQL 1.1 implementation covering queries and update statements.

The audience is therefore fairly specific. If you are writing Python and you have an .ttl file, a JSON-LD document or a SPARQL endpoint response to deal with, RDFLib is the shortest path. It is also the substrate other Python RDF tooling builds on: the README lists sparqlwrapper, pyLODE, pyrdfa3, pymicrodata, pySHACL and OWL-RL as sibling packages in the RDFLib family, which tells you the ecosystem expects an rdflib Graph as its input type. If you are doing numeric graph analysis rather than semantic data interchange, this is not the library you want, and the README does not pretend otherwise.

## How the Graph, stores and parsers fit together

The central object is Graph, which the README describes as a Python collection of subject, predicate, object triples. Parsing is a method on that object: g.parse(source) loads a document, and the format is inferred from the source or passed explicitly. Iteration yields triples directly, and g.value(subject, predicate) returns a single object matching a triple pattern, or an arbitrary one when several exist. That last detail matters more than it looks: value() is a convenience, not a uniqueness guarantee.

Storage is separated from the API. The README lists store implementations for in-memory, persistent on disk via Berkeley DB, and remote SPARQL endpoints, and notes that additional stores can be supplied as plugins. So the same Graph code can run against memory during tests and against a remote endpoint in production, provided you accept the performance characteristics of a remote store. The Berkeley DB store is behind an optional dependency, which is a signal that the default in-memory store is the one most users run.

SPARQL is implemented in-process for local graphs, with a function extension mechanism for adding custom SPARQL functions. The repository also ships command line entry points declared in pyproject.toml, including rdfpipe for format conversion, csv2rdf, rdf2dot, rdfs2dot, rdfgraphisomorphism and sparqlquery. Those tools make RDFLib usable from a shell without writing Python at all.

## Installing RDFLib and running a first query

The README gives pip as the installation route for the stable release. A plain install pulls only the core dependencies; features such as the Berkeley DB store, networkx integration, HTML parsing, lxml and orjson are extras, and the README shows them listed together in one extras group.

```bash
pip install rdflib
```

If you need any of the optional features, install the corresponding extras instead. The README's example names them explicitly, so you can copy the subset you want rather than all of them.

```bash
pip install rdflib[berkeleydb,networkx,html,lxml,orjson]
```

For a first real use, the README's getting started example creates a Graph and parses a remote document. Parsing a URL fetches the resource, so expect a network round trip and a format guess from the response. Iterating the graph then prints each triple.

```python
from rdflib import Graph
g = Graph()
g.parse('http://dbpedia.org/resource/Semantic_Web')

for s, p, o in g:
    print(s, p, o)
```

The README also shows the namespace module, which is where most real code starts. Importing RDFS and XSD gives you the standard vocabularies without typing full IRIs, and g.value returns the object of a triple pattern.

```python
from rdflib import Graph, URIRef, Literal
from rdflib.namespace import RDFS, XSD

g = Graph()
semweb = URIRef('http://dbpedia.org/resource/Semantic_Web')
type = g.value(semweb, RDFS.label)
```

If you prefer the shell, pyproject.toml declares rdfpipe as a console script, and the repository contains an examples/ directory with runnable files such as sparql_query_example.py, sparql_update_example.py, datasets.py and jsonld_serialization.py. Those are the fastest way to see the intended usage of each feature without reading the full documentation site.

## The in-memory store is the real limit

The default store keeps the entire graph in the process. Nothing in the README suggests otherwise, and the persistent option it names is Berkeley DB rather than a server. That means the practical ceiling is your machine's memory, and a graph that grows past it will fail rather than degrade. For a few hundred thousand triples this is usually fine. For a production knowledge graph with millions of statements, it is not, and the README's answer is to point the Store at a remote SPARQL endpoint instead.

That answer has its own cost. A remote endpoint store changes every traversal into HTTP, so patterns that are cheap in memory become latency-bound. There is no query planner in RDFLib that will rescue a badly written SPARQL query against a remote store; the endpoint's planner does that work, and you lose control over it.

A second limitation is the Python version floor. The pyproject.toml in the repository declares requires-python >= 3.10 and classifiers for 3.10 through 3.14. If you are pinned to an older interpreter, this is a hard stop, and the README's version history shows older release lines existed for older Pythons. Third, the README does not document rollback or migration behaviour between releases, so upgrading across a major version is something you have to verify against the changelog yourself.

## RDFLib compared with NetworkX and with a triple store

The most common confusion is rdflib vs networkx. NetworkX is a graph analytics library: nodes and edges, algorithms, centrality, shortest paths. RDFLib is a data interchange library: IRIs, literals with datatypes, namespaces, and serialization formats. They overlap in the word graph and almost nowhere else. RDFLib has a networkx extra, which tells you the intended relationship: you parse and validate RDF with RDFLib, then hand the structure to NetworkX if you need algorithms. Choosing NetworkX for RDF data means writing your own parser and losing the datatype model.

The other comparison is against a dedicated triple store such as Fuseki or a SPARQL endpoint. A store is a server with indexing, transactions and concurrent access. RDFLib is a library inside your process. The repository does ship a with-fuseki.sh script and an RDF4J client extra in 7.5.0, which suggests the maintainers expect RDFLib to sit next to a store rather than replace it. If several services need to query the same graph concurrently, a library is the wrong shape.

For ontology reasoning specifically, the README points at OWL-RL as a separate package rather than a feature of the core, so do not expect RDFS or OWL inference to run automatically when you load a graph.

## Maintenance, licensing and the cost of upgrading

The repository is not archived, and the last push was on 2026-09-23, five days before this writing. Release 7.6.0 is dated 2026-02-13, 7.5.0 is dated 2025-11-28, and 7.4.0 is dated 2025-10-30, so there is a regular cadence roughly every two to three months. The README describes main as the backwards-compatible development branch and next as the next major release branch, which is a useful convention: if you track main you are opting into fixes and features that are intended to be compatible, and next is where breakage lives.

The licence is BSD-3-Clause, declared both in the README's badge set and in pyproject.toml as license = { text = "BSD-3-Clause" }. That is a permissive licence, which generally means you can use it in closed products provided you keep the copyright notice and disclaimer. This is not legal advice, and the exact obligations are in the LICENSE file at the repository root.

The upgrade cost is mostly dependency-related. The core dependencies are isodate (only for Python below 3.11) and pyparsing bounded to >=3.1.0,<4. Optional extras carry their own upper bounds: berkeleydb <19.0.0, networkx <4, lxml <6.0, orjson <4, httpx <0.29.0 for the rdf4j and graphdb extras. Those upper bounds mean a major release of any of those libraries will not be picked up until RDFLib widens the range. Pinning your own lockfile is the practical response.

## Conclusion

Adopt RDFLib when your data is already RDF and your processing fits in a Python process, or when you need to convert between Turtle, JSON-LD, N-Quads and RDF/XML without running a database. Do not adopt it as a general graph analytics engine, and do not expect the in-memory store to hold a dataset that does not fit in RAM. Before you build on it, check the requires-python floor in pyproject.toml, decide which optional extras you actually need, and confirm whether your target store is one of the ones the README lists.

## FAQ

### What is RDFLib in Python?

RDFLib is a pure Python package for working with RDF. It provides parsers and serializers for RDF/XML, N3, NTriples, N-Quads, Turtle, TriX, Trig, JSON-LD and HexTuples, a Graph interface backed by pluggable stores, and a SPARQL 1.1 implementation.

### How do I install RDFLib?

The README gives pip install rdflib for the stable release. Optional features are installed through extras, for example pip install rdflib[berkeleydb,networkx,html,lxml,orjson].

### How do I use RDFLib to load and read a graph?

Create a Graph, call g.parse() on a file or URL, then iterate the graph to get subject, predicate, object triples. The README's example parses a DBPedia URL and prints each triple.

### Is RDFLib an alternative to NetworkX?

They solve different problems. RDFLib handles RDF data, namespaces and serialization formats, while NetworkX is a graph analytics library. RDFLib ships a networkx extra, which points to using the two together rather than choosing one over the other.

### What is RDFLib?

RDFLib is a pure Python library for working with RDF, the subject, predicate, object data model. It covers parsers and serializers for the common RDF formats, a Graph interface over pluggable stores, and a SPARQL 1.1 implementation.

## Sources

- [License: BSD-3-Clause](https://github.com/RDFLib/rdflib/blob/main/LICENSE)
- [Project website](https://rdflib.readthedocs.org)
- [RDFLib/rdflib on GitHub](https://github.com/RDFLib/rdflib)
- [README](https://github.com/RDFLib/rdflib/blob/main/README.md)
- [Releases](https://github.com/RDFLib/rdflib/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/rdflib-rdflib
