# schemaorg/schemaorg: the repository behind the vocabulary, not a library you install

> Schema.org's repository holds the schemas, examples and publishing software behind the vocabulary. It is aimed at contributors and tool builders, not at developers looking for a runtime dependency, and the README is explicit that the site itself lives elsewhere.

**schemaorg/schemaorg** — Schema.org - schemas and supporting software

- Repository: https://github.com/schemaorg/schemaorg
- Website: https://schema.org/
- Stars: 6,262 · Forks: 964
- Language: HTML
- License: Apache-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/schemaorg-schemaorg

## What schemaorg/schemaorg actually is, and who ends up needing it

The repository is not a client library. The README states that it contains "all the schemas, examples and software used to publish Schema.org", and then points elsewhere for the site itself. That single sentence defines the audience. If you are writing JSON-LD for a product page, this repository is not on your critical path. If you are maintaining a vocabulary, generating documentation, or building tooling that consumes the raw definitions, it is the source of truth.

The README also notes that much of the supporting software is imported from a sub-module named sdopythonapp. That matters for anyone planning to clone and build: part of what you would be editing does not live in the main tree. The top-level layout confirms the split, with data/, docs/, software/, templates/ and versions.json sitting alongside the README and LICENSE.

Governance is handled through GitHub issues and the W3C Schema.org Community Group. Issue #1 is described as the entry point for release planning and for per-release milestones. Substantive changes are recorded in the release notes, and the README says a formal release happens roughly every month after review by the steering group and the wider community.

## How the vocabulary is versioned and released

Releases are dated artifacts, not a rolling stream. The recent list runs v29.4-release in December 2025, v30.0 in March 2026, and v30.1 in September 2026. Those gaps are wider than the monthly cadence described in the README, which suggests the monthly line refers to the general rhythm rather than a guarantee.

For consumers, the practical consequence is that a term you see on the website may not yet be in the release you pinned. The README points to a staging site where draft release notes appear before a release is final, and to the published release history for what has actually shipped. If your pipeline validates structured data against a fixed vocabulary, that staging page is the place to look before a version bump surprises you.

The versions.json file at the top level is the machine-readable counterpart to those human-readable notes. The README does not document its schema, so treat it as something to inspect rather than something to rely on without reading.

## Cloning the repository and finding a term in the schemas

There are no install instructions in the README, no package name, and no published artifact for the schemas themselves. What the README gives you is a pointer to the website and to the W3C community group. So the honest first step is a clone, not an install, and then a look at what is actually in the tree.

The README documents the sub-module by name: much of the supporting software is imported from sdopythonapp. It does not give a command for initializing it, and the repository files supplied here contain no such command, so the only thing that can be stated with confidence is which directory to look at. The top-level entries are .codespellrc, .gitattributes, .github/, .gitignore, LICENSE, README.md, data/, docs/, software/, templates/ and versions.json.

Once cloned, the schemas themselves are the useful part. The definitions live under data/, and the generated documentation and examples under docs/. A first real use is to open those directories and read the term you intend to publish, rather than trusting a blog post about it. The README does not describe the internal format of the files in data/, so expect to read them rather than to be handed a schema for them.

## The contribution bar is deliberately high, and that is a limitation

The README is unusually direct about what will not get merged. Large-scale reorganizations motivated by elegance, "proper modeling", ontological purity or conceptual unification are described as highly unlikely to be taken on. The project states it has traded global consistency for incremental evolution and a tolerance for a style that would be out of place in a formal ontology.

That is a real constraint on contributors. If your proposal is a cleaner hierarchy, it will probably be rejected even if it is technically better. The README also says Schema.org does not attempt to capture the full detail of web content, and that the project will often choose not to add detail in order to keep the vocabulary simple for publishers. So a well-argued request to model something more precisely can be turned down on usability grounds.

The bar for new vocabulary is evidence of a consumer. The README says additions are most likely when there is evidence that some preferably large-scale consuming application will make use of the data, and that search engines in general are not sufficient justification. Smaller, backwards-compatible changes are easier to land. If your change requires consumers to update, expect a longer conversation.

## How schemaorg/schemaorg differs from libraries that wrap it

The related searches around this project mix the vocabulary with client libraries built on top of it, which is where most confusion starts. A Python or JavaScript library that emits Schema.org markup is a different artifact from this repository. The library implements a consumer or producer view of the vocabulary; this repository defines the vocabulary and publishes the site.

The difference shows up in what you get when something changes. If a library lags a release, you upgrade the library. If a term is wrong or missing in the vocabulary itself, the fix has to go through the issue tracker and the community group, and it will arrive in a dated release rather than in your next dependency bump.

Compared with a formal ontology project, the trade-off runs the other way. Schema.org accepts weaker modeling in exchange for markup that non-specialists can write without understanding JSON-LD or RDF/S. The README is candid that logically equivalent structures can produce many more errors from publishers unfamiliar with the underlying formal concepts. If you need strict logical consistency, this is the wrong vocabulary, and the README itself suggests the Ontolog community for that kind of proposal.

## Licence, maintenance and what an upgrade costs you

The repository is Apache-2.0. The README does not resolve the split between document and software licensing in the repository itself; it points to a FAQ entry on the Schema.org site for Creative Commons and open source licensing of documents and software. Read that entry before redistributing the schemas or the generated documentation, because the two categories are treated separately there. This is a pointer, not legal advice.

Maintenance is current: the last push was on 2026-09-22, and the most recent release, v30.1, was published on 2026-09-15. The repository is not archived. The cadence described in the README, a formal release roughly every month after review, is the number to plan against.

Upgrade cost depends on how you consume the vocabulary. If you pin a version and validate against it, a bump means re-checking the terms you use against the release notes at schema.org/docs/releases.html. If you build the site or documentation from source, remember that part of the software comes from the sdopythonapp sub-module, so a clone that does not include it is incomplete. The README does not document a rollback procedure for a published release.

## Extending Schema.org without forking it

The README makes a point that is easy to miss: Schema.org is not a closed system. It uses JSON-LD, Microdata and RDFa specifically to allow independent extension, and it cites GS1's vocabulary as an example of terms defined elsewhere that can be mixed in alongside Schema.org's own.

This is the practical answer for teams whose domain is not covered. Rather than pushing a new term through the release process, you can publish your own vocabulary and reference it from the same markup. The README notes that other initiatives such as Wikidata and GS1 have defined many terms that can be combined with Schema.org's.

Alignment with external standards is also deliberate. The README describes influence from MARC, BibFrame and FRBR in bibliographic and cultural heritage contexts, from Good Relations and GS1 in e-commerce, from IPTC's rNews plus fact-checking and Trust Project work in news, and from the BBC, the European Broadcasting Union, the Music ontology and MusicBrainz in TV and music. The stated preference is to blend existing designs rather than produce an isolated pure model, accepting a less elegant global result.

## Conclusion

Adopt this repository only if you intend to contribute vocabulary changes, regenerate the site, or read the schemas directly; application developers who just need structured data should use the published vocabulary at schema.org and a consumer library instead. Before opening a pull request, read the prioritization rules in the README: simple fixes and backwards-compatible changes to existing schemas, examples and documentation come before new vocabulary, and additions are more likely when some consuming application is expected to use the data. Also check the data/ and software/ directories and the sdopythonapp sub-module before proposing a change, because part of what you would be editing lives outside the main tree.

## FAQ

### What is schema.org used for?

It is a vocabulary for describing things on the web, published with examples and software in this repository. The README states that the repository holds all the schemas, examples and software used to publish Schema.org, and that the site itself is a separate destination.

### How do I use schemaorg/schemaorg?

Clone the repository and work with the definitions under data/ and the generated material under docs/; the README gives no install step or package name. The README points to the website for the published vocabulary and to the W3C Schema.org Community Group if you want to participate in changes.

### Is schemaorg/schemaorg the same as JSON-LD?

No. Schema.org uses web standards such as JSON-LD, Microdata and RDFa to allow independent extension, so JSON-LD is one serialization the vocabulary is expressed in, not the vocabulary itself.

## Sources

- [License: Apache-2.0](https://github.com/schemaorg/schemaorg/blob/main/LICENSE)
- [Project website](https://schema.org/)
- [README](https://github.com/schemaorg/schemaorg/blob/main/README.md)
- [Releases](https://github.com/schemaorg/schemaorg/releases)
- [schemaorg/schemaorg on GitHub](https://github.com/schemaorg/schemaorg)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/schemaorg-schemaorg
