NASA-CMR-STAC: A Daily-Updated Dataset List for CMR's STAC Endpoint
A list of geospatial datasets on NASA's Common Metadata Repository (CMR)
At a glance
- What is it?
- The repository is a mirror of NASA CMR's STAC catalog as TSV and JSON, refreshed daily, not a STAC client or API. It answers one question: which datasets exist, and how do I get their IDs into a script?
- Who is it for?
- Adopt it if you need a flat, machine-readable inventory of CMR STAC collections for scripting, filtering or building your own index, and you are comfortable reading a TSV over HTTP. Do not adopt it if you need search, download or authentication against CMR itself; the repository is a list, and the README points to CMR for everything else.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 1 day ago.
- What is it written in?
- Mainly Jupyter Notebook, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What NASA-CMR-STAC Actually Is, and Who It Is For
NASA's Common Metadata Repository (CMR) is a metadata catalog of NASA Earth Science data. CMR-STAC is the SpatioTemporal Asset Catalog view of that metadata. The problem this repository addresses is discovery at the list level: if you want to know which collections exist before you query anything, you need to enumerate them, and CMR's own interfaces are built for search rather than for bulk listing. The README states the repo compiles the list of all geospatial datasets on CMR-STAC as a CSV and as a JSON file, and that the list is updated daily.
The audience is narrow and specific. It is for people who write scripts that need a starting inventory: a data engineer building a picker, a notebook author who wants a DataFrame of collection ids, or someone who needs to diff the catalog over time. It is not for someone who wants to search by bounding box, filter by cloud cover or download granules. Those are CMR operations, and this repository does not wrap them.
How the List Is Produced and What the Files Contain
The repository is a Jupyter Notebook project at its core. The top-level entries include nasa_cmr_catalog.ipynb, nasa_cmr_catalog.py and nasa_cmr_catalog.tsv, which tells you the shape of the pipeline: a Python script generates the catalog, the notebook is the interactive view of the same work, and the TSV is the checked-in output. The README describes the result as two formats, a tab separated values file and a JSON file, and says the list is updated daily.
The data flow is therefore one-directional. CMR-STAC is read, the collection metadata is flattened into rows, and the result is committed to the repository. Consumers read the committed file rather than calling CMR themselves. That is the whole mechanism, and it is worth being explicit about the consequence: the repository is a snapshot, and its freshness depends entirely on a scheduled job that is not documented in the README. The README does not describe the schedule, the credentials or the failure behaviour of that job.
Installing It and Reading the Catalog in a Notebook
There is nothing to install to consume the list. The README's usage section gives a single example that reads the TSV directly from the raw GitHub URL into a Pandas DataFrame. The URL is the one the README uses, and the separator is a literal tab.
import pandas as pd
url = 'https://github.com/giswqs/NASA-CMR-STAC/raw/master/nasa_cmr_catalog.tsv'
df = pd.read_csv(url, sep='\t')
df.head()What you should see is a DataFrame whose columns come from the flattened CMR-STAC collection records, with one row per dataset. The README does not list the column names, so print df.columns before you write code against them.
If you want to regenerate the list rather than read it, the repository ships requirements.txt, and its entire contents are one line.
pip install -r requirements.txt
python nasa_cmr_catalog.pyThe only declared dependency is leafmap, which is a mapping and geospatial analysis package rather than a CMR client. That is a signal about how the script works: it is a notebook-oriented workflow, and the README does not document command line arguments for nasa_cmr_catalog.py. Expect to read the script before running it against anything you care about.
The Refresh Cadence Is the Main Risk
The README says the list is updated daily. It does not say what triggers that update, whether it runs on GitHub Actions, what happens when CMR-STAC changes its schema, or how a failed run is surfaced. The .github directory exists at the top level, which is consistent with a scheduled workflow, but the README does not document it and this review cannot confirm it.
That matters because the value of the repository is entirely in its freshness. If the daily job stops, the TSV and JSON keep serving stale collection lists, and nothing in the file itself tells you when it was generated. There is no timestamp column documented in the README. A consumer who needs to know whether a collection is current has to check CMR directly, which defeats part of the purpose of using a mirror.
The second failure mode is schema drift. CMR-STAC is an evolving specification, and a flattened TSV is a lossy projection of it. When the upstream collection schema gains or renames fields, the generated columns change, and any downstream code that hardcodes column names breaks silently. The README gives no compatibility statement and no versioning of the output format.
Why You Would Still Use It Instead of Querying CMR
The alternative is to skip the mirror and call CMR's search API or the CMR-STAC endpoint directly. That approach gives you live data, filtering, pagination and the full record rather than a flattened row. It is the right choice when you need to search by time, place or attribute, or when you need the STAC item geometry rather than a collection summary.
The difference in approach is between a cached index and a query engine. This repository trades completeness and freshness guarantees for a file you can fetch with one HTTP request, load into Pandas, and join against your own tables without writing pagination logic. For a one-off inventory, a documentation table, or a comparison of collection ids across catalogs, that trade is reasonable. For anything where a stale row would produce a wrong answer, it is not. Related projects listed in the README follow the same pattern for other catalogs, including aws-open-data-stac and Planetary-Computer-Catalog, which suggests the author treats this as a family of index files rather than as an application.
Licence, Maintenance and Upgrade Cost
The repository is MIT licensed, which permits reuse, modification and redistribution with the licence and copyright notice preserved. That is permissive enough for embedding the TSV in a commercial pipeline. The licence covers the repository's own code and generated files; it does not grant rights to the NASA data the catalog describes, and the README does not discuss data licensing. Treat the two separately and read NASA's own terms before redistributing the underlying datasets.
Upgrade cost is low in the ordinary case, because there is no library to pin and no API surface to migrate. You read a file over HTTP. The cost appears when the output schema changes: any code that indexes columns by name has to be revisited. The repository provides no changelog and no releases, so there is no signal to watch. The practical check is to compare df.columns from a fresh read against what your code expects. The last push date for this repository is not available in the information reviewed here, so the current maintenance state cannot be stated; the README's claim of daily updates is the only cadence information available.
Editorial conclusion
Adopt it if you need a flat, machine-readable inventory of CMR STAC collections for scripting, filtering or building your own index, and you are comfortable reading a TSV over HTTP. Do not adopt it if you need search, download or authentication against CMR itself; the repository is a list, and the README points to CMR for everything else. Before relying on it, open nasa_cmr_catalog.tsv and confirm it contains the collection ids you expect, then check nasa_cmr_catalog.py to see how the list is generated, because that script is the only place the refresh logic lives.
Frequently asked questions
What is NASA's Common Metadata Repository (CMR)?
The README describes CMR as a metadata catalog of NASA Earth Science data. CMR-STAC is the SpatioTemporal Asset Catalog view of that metadata, and this repository compiles the list of geospatial datasets available through it.
What is the STAC data model?
The README does not explain the STAC data model itself; it points to the CMR SpatioTemporal Asset Catalog documentation and treats CMR-STAC as the source of the collection list. What this repository exposes is a flattened TSV and JSON projection of those collections.
What is STAC used for?
The README does not describe STAC's general use cases. It only states that CMR-STAC is where the geospatial datasets are listed, and that this repository compiles that list into a TSV file and a JSON file for programmatic use.
What are the benefits of using STAC?
The README does not list benefits of STAC. The benefit it claims for this repository is narrower: compiling the CMR-STAC dataset list into TSV and JSON makes it easier to find and use them programmatically, with the list updated daily.