SchemaSpy: Generating Database Documentation and ER Diagrams from a Live Schema
Database documentation built easy
At a glance
- What is it?
- SchemaSpy reads database metadata and writes a static HTML report of tables, relationships and design anomalies. It is a JAR or a Docker image, not a service, and the README leaves several operational questions open.
- Who is it for?
- Adopt SchemaSpy if you need a static, shareable snapshot of a schema and can run a JAR or a Docker image against a read-only replica. Do not adopt it if you need continuous, queryable documentation or a hosted UI, because the report is regenerated per run and the project publishes no server mode.
- Can I use it commercially?
- Yes, with conditions. LGPL-3.0 is a weak copyleft licence: you can use it inside commercial and closed-source software, but if you distribute changes to its own files, you must publish those changes under the same licence.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly HTML, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What SchemaSpy Produces That a Manual ER Diagram Does Not
The README states the problem plainly: entity-relationship diagrams are the preferred way to document a database, but drawing them by hand is slow and error-prone, and the diagrams rarely stay current once drawn. SchemaSpy takes the opposite route. It reads structural metadata from a live database and writes an HTML report, so the diagram is a byproduct of a run rather than an artifact someone maintains.
The intended audience is database administrators and developers who need to visualize, navigate and understand a data model. The README also names a less obvious audience: anyone who wants to share schema information without exposing rows. Because the tool reads only structural information, the README says it works just as well on an empty database replica, which makes the output safe to hand to a third party for analysis.
A third use case is design review. The README says SchemaSpy incorporates knowledge about best practices in database design and can locate anomalies such as missing indexes, implied relationships and orphan tables. That turns the report into something closer to a lint pass over the schema than a picture of it.
How the Metadata Read and HTML Report Generation Fit Together
SchemaSpy is a standalone application without a GUI. That single sentence in the README explains most of the architecture: there is no server to keep running and no web UI to log into. You invoke a process, it connects to the database over JDBC, and it writes files.
The connection details are supplied on the command line, and the database dialect is selected with a type flag rather than inferred. The README's PostgreSQL example passes -t pgsql11 alongside -db, -host, -port, -u, -p and -o. The JDBC driver is not bundled into the bare-bone JAR; it is handed in separately with -dp, which is why the quick start downloads a driver before running anything. The README states that over a dozen databases are supported out of the box and that the list is printed by -dbhelp. Anything outside that list can be plugged in as long as a JDBC driver exists, which the README links to a configuration page for.
The Docker image resolves the driver question differently. Its Dockerfile builds a drivers stage that curls four JDBC drivers (MySQL, MariaDB, PostgreSQL and jTDS) into /drivers_inc, and the final image sets SCHEMASPY_DRIVERS=/drivers and SCHEMASPY_OUTPUT=/output. The image is based on eclipse-temurin:17.0.9_9-jre-jammy and installs graphviz, which is the piece that renders diagrams. The entrypoint is /usr/local/bin/schemaspy, and the container runs as a non-root java user with /output owned by that user.
That layout has a practical consequence worth stating: the Docker image covers four drivers, not the full dozen-plus. If your engine is not among them, you are back to supplying a driver yourself.
Installing SchemaSpy and Running a First Report Against PostgreSQL
The README gives two distribution channels: a JAR file from the releases page and a Docker image. The curl example below downloads a released JAR. Note that the README's snippet still names 6.2.4 and tells you to replace it with the latest version; the most recent release listed for the project is v7.0.2, so the filename and URL path will differ from the example.
# replace '6.2.4' with latest version
curl -L https://github.com/schemaspy/schemaspy/releases/download/v6.2.4/schemaspy-6.2.4.jar \
--output ~/Downloads/schemaspy.jarAfter that, the quick start assumes PostgreSQL 11 or later and downloads the JDBC driver separately. The driver is not optional for the bare-bone JAR.
curl -L https://jdbc.postgresql.org/download/postgresql-42.5.4.jar \
--output ~/Downloads/jdbc-driver.jarWith both files present, the run itself is one command. The -o flag names the output directory, and the README says the result is browsable at DIRECTORY/index.html.
java -jar ~/Downloads/schemaspy.jar \
-t pgsql11 \
-dp ~/Downloads/jdbc-driver.jar \
-db DATABASE \
-host SERVER \
-port 5432 \
-u USER \
-p PASSWORD \
-o DIRECTORYReplace DATABASE, SERVER, USER, PASSWORD and DIRECTORY with your own values; the port 5432 shown is the PostgreSQL default. If you are not on PostgreSQL, run the tool with -dbhelp to print the supported database types before guessing at a -t value. For Maven users, the README lists two artifacts under the same GAV, org.schemaspy:schemaspy:<version>, with the fat JAR distinguished by the app classifier. That classifier is easy to miss and is the difference between a JAR that runs on its own and one that needs its dependencies on the classpath.
Where SchemaSpy Stops Being the Right Tool
The report is a snapshot, not a live view. The README frames CI/CD generation as the answer to staleness, but that is a workflow you build, not a feature the tool provides. If your schema changes several times a day and people read the report as reference material, the report is wrong between runs, and nothing in the README describes a freshness indicator or a diff against the previous run.
The credential handling is also plain. The quick start passes -u and -p on the command line, which means the password can land in shell history and in process listings. The repository does include a schemaspy.properties_template at the top level, which suggests a properties file path exists, but the README excerpt does not document its keys, so the exact mapping from command-line flags to properties entries is something you would have to confirm from the documentation site rather than from the README.
Scale is another boundary the README does not address. Nothing in the README describes how the tool behaves on a schema with thousands of tables, or how long diagram rendering takes at that size. Graphviz is doing the layout work, and layout is the part that tends to get slow. Treat large-schema performance as unverified rather than assumed to be fine.
Finally, the output is HTML files on disk. There is no query interface, no API and no search endpoint beyond whatever the generated pages offer. Teams that want to ask questions of the metadata programmatically are looking at the wrong layer.
SchemaSpy Alternatives and the Difference in Approach
The most direct alternative is a database client that draws diagrams on demand, such as DBeaver or pgAdmin. The difference is where the work happens. Those tools render a diagram inside a session that a person is driving, and the diagram lives in that session or in a saved project file. SchemaSpy writes a directory of static HTML that anyone can open in a browser, with no client installed and no database connection. That is the whole trade: you give up interactivity and get portability.
A second alternative is a documentation platform that ingests schema metadata into a searchable catalog. Those systems keep a persistent index that can be queried and linked to, which SchemaSpy does not attempt. The cost is that you now operate a service and a data store, and the metadata has to be pushed into it. SchemaSpy's model is the inverse: no service, no index, one process per run.
A third option is writing your own queries against the information schema and templating the output. That gives you exactly the fields you want and nothing else. It also means you own the diagram layout, the anomaly checks and the HTML, which is precisely the work SchemaSpy already does. The README's claim that the tool knows about design best practices and flags missing indexes and orphan tables is the part you would be reimplementing.
Licence, Maintenance and the Cost of Upgrading
SchemaSpy is distributed under LGPL-3.0. The repository carries both COPYING and COPYING.LESSER at the top level, which is the usual pairing for that licence. For most users this changes nothing: you run the JAR or the container and read the HTML it produces. The question that matters is whether you modify SchemaSpy itself, because the LGPL's obligations attach to the library and to modified versions of it, not to the report files it generates. I am not giving legal advice here; if you plan to embed SchemaSpy in a product or ship a patched build, have someone read the actual licence text in COPYING.LESSER rather than a summary.
On maintenance, the repository is not archived, and the last push was on 2026-03-05. The release history is uneven: v7.0.2 arrived on 2025-09-20, while the previous releases, v6.2.4 and v6.2.3, date to July and June 2023. That gap is worth knowing about before you plan around a steady upgrade cadence.
Upgrade cost is mostly the driver and the flags. The Dockerfile pins specific driver versions (MYSQL_VERSION=8.4.0, MARIADB_VERSION=1.1.10, POSTGRESQL_VERSION=42.7.2, JTDS_VERSION=1.3.1), so a container upgrade can move the driver under you. If you supply your own driver with -dp, you own that compatibility. The README does not document a rollback procedure for a report directory, which is fine because the output is disposable: keep the previous directory if you want a comparison point, and regenerate.
Editorial conclusion
Adopt SchemaSpy if you need a static, shareable snapshot of a schema and can run a JAR or a Docker image against a read-only replica. Do not adopt it if you need continuous, queryable documentation or a hosted UI, because the report is regenerated per run and the project publishes no server mode. Before committing, verify which database type string matches your engine via -dbhelp, confirm the JDBC driver version you download matches your server, and check whether your report will be regenerated by hand or from a scheduled job, since the README documents neither a rollback path nor a hosted deployment.
Frequently asked questions
What is SchemaSpy used for?
SchemaSpy is a database metadata analyzer that reads structural information from a database and writes an HTML report for visualizing, navigating and understanding a data model. The README also describes using it to collect statistics about the database structure and to detect design anomalies such as missing indexes, implied relationships and orphan tables.
How do I install SchemaSpy?
Download the JAR file from the releases page or use the Docker image; the README describes it as a standalone application without a GUI, so there is no installer. To run it against PostgreSQL you also download a JDBC driver and pass it with the -dp flag.
How do I use SchemaSpy?
Run the JAR with a database type flag, a JDBC driver path, connection details and an output directory, then open index.html inside that directory. The README's example uses -t pgsql11, -dp, -db, -host, -port, -u, -p and -o. Run with -dbhelp to list the supported database types.
Is SchemaSpy free and open source?
The project is published under LGPL-3.0, and the repository contains both COPYING and COPYING.LESSER at the top level. The README presents the JAR and the Docker image as the distribution channels, with no paid tier mentioned.
Is SchemaSpy open source?
Yes. The licence is LGPL-3.0, and the repository carries COPYING and COPYING.LESSER at the top level alongside the source tree. The README points readers to the documentation site and the releases page rather than to any commercial offering.
How does SchemaSpy compare with other tools?
The README does not name competing products, so no direct comparison is documented. What it does describe is the shape of the output: a static HTML report with ER diagrams generated from database metadata, which differs from a diagram drawn interactively inside a database client.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/schemaspy-schemaspy)