EdgarTools caps its HTTP client on purpose, and its release cadence outruns its last push
Read and analyze SEC EDGAR filings in Python. 10-K, 8-K, XBRL financials, Form 3/4/5, 13F, ADV — clean API, well-typed, MIT-licensed.
At a glance
- What is it?
- EdgarTools turns SEC EDGAR filings into typed Python objects, with an MCP server and XBRL financial statements on top of a library that identifies you to the SEC by email. Its dependency policy, its two entry points and its documentation hosting are the parts worth reading closely.
- Who is it for?
- EdgarTools earns its place if you want SEC filings as typed Python objects inside your own process rather than JSON from a hosted endpoint, and if the forms you care about are in the covered set, which spans financial statements, insider forms, fund holdings and proxy material. Verify three things before you build a pipeline on it.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
You identify yourself to the SEC before any data arrives
The install step is one command, and the second step is the one that has consequences:
pip install edgartoolsfrom edgar import *
set_identity("[email protected]")EDGAR asks every requester to identify themselves with an email address, and this library makes that a one time call rather than a per request header. The framing is deliberate: no key, no signup, no rate limit tier, set it once. That is also the whole authentication story, which is worth being clear about, because there is nothing else to configure and nothing to rotate.
What that means in practice is that a real address leaves your machine on every fetch, attached to requests that identify you as the person pulling the data. For a personal research script that is a fair trade. For anything running inside a company, the address is the company's problem to answer for rather than yours alone, and the documentation does not discuss that distinction.
It is also the reason the library can claim no quotas. There is no vendor tier to be placed in, so rate limiting is something you configure for yourself rather than something you buy.
Two entry points, and the global one returns whoever filed last
The use cases show two different ways in, and the difference is easy to skim past.
The company scoped path takes a ticker and goes through the company object:
financials = Company("MSFT").get_financials()
financials.balance_sheet() # all line items
financials.income_statement() # revenue, net income, EPSThe other path is a module level call with no company in it at all:
thirteenf = get_filings(form="13F-HR").latest().obj()
thirteenf.holdings # every portfolio position as a DataFrameRead that second snippet carefully. There is no fund named, so the object returned belongs to whichever 13F filer submitted most recently across the whole database, and the holdings you get are that filer's portfolio rather than a named institution's. The same shape applies to the 8-K example, where get_filings(form="8-K").latest().obj() returns one filer's reported event items with no company attached. For exploration that is a fast way in; for a report, the company scoped path is the one that identifies whose data you are holding.
The third path is neither of those. Company("AAPL").get_facts() hands back a query object that can be filtered by concept, which is how a revenue history across years is pulled without naming the form at all.
The dependency cap is a policy decision written in a comment
The most revealing lines in the package manifest are comments rather than entries. Sitting above the dependency list is a note explaining that an upper cap is deliberate, that the HTTP client upstream is dormant with its last release at 0.28.1 in December 2024 and its issues and discussions closed in February 2026, and that any future release would be unexpected and must not be auto adopted.
So the library's own maintainers decided that a security or feature release from the HTTP client, if one ever appeared, should not silently change behaviour underneath their users. That is an unusual amount of care to take over a transitive dependency, and the comment says why: the migration path is a package named httpx2, published from a project called edgartools-q2iz, and the note ends by saying to lift or retarget the cap. The sentence stops there.
Two things follow for anyone reading the manifest. The transport layer of a financial data library is a fork with a different name, published under an unrelated project, which is a supply chain fact worth knowing about even though the code is MIT licensed and inspectable. And the cap means a future upstream fix, including a security fix, will not arrive through a normal upgrade. The visible part of the dependency list is cut off after that comment, so what else is pinned cannot be read here.
MkDocs configuration on a readthedocs.io address
The documentation is split across several places in the tree and the pieces do not obviously belong to one tool. There is a mkdocs.yml at the root, which is the configuration file for the MkDocs generator. There is also a .readthedocs.yaml, which is the configuration for hosting on readthedocs.io, and the README's links all point at that domain rather than at a project page. The link form used in the quick start is the hosted one, with a language segment and a version segment in the path.
So the docs are built by one generator and served from a host whose default association is a different documentation tool. The result works for a reader, since the links in the README resolve to hosted pages, but it means the local preview command and the published site are governed by two configuration files that have to agree with each other.
The rest of the tree is broader than a library package. There is an edgar/ directory for the code itself, plus data/, notebooks/, tools/, engineering/, scripts/, examples/, tests/ and docs/. Two of those deserve a note. The examples directory contains its own README and a script whose whole subject is table width, alongside a second README about the same topic, so rendering behaviour has been fiddly enough to document separately. And engineering/ suggests some of the effort went somewhere other than the published package.
The comparison table prices the library against a hosted API
The documentation includes a head to head table against one named competitor, and every row is an argument about where the work happens rather than about the data.
On cost the row reads free and MIT against 49 dollars a month and up. On format, typed Python objects that turn into DataFrames against JSON you parse yourself. On where it runs, inside your process with no key, no quotas and no vendor lock-in, against a hosted API with a key and rate tiers. On coverage, 20 plus typed forms against 15 plus structured endpoints. On open source, inspect, fork and self-host against proprietary.
The AI and MCP row is the interesting one, and it is also the sloppiest. The EdgarTools cell says built in, and the competitor cell is empty. Nothing in the row says what the built in part consists of, even though the feature list names it: a built in MCP server and LLM ready text, plus HTML converted to clean text and markdown for retrieval, and full text search over filings.
The section closes with a recommendation that begins by saying that in Python this library gives you typed objects, AI native output and the full SEC corpus, and then stops partway through the phrase free, open, and i. So the strongest claim in the comparison is the one that is not finished being written.
Twenty plus forms, and section extraction for the ones that have no XBRL
The coverage claim is a count, and the count is twenty plus forms. The named ones run from the obvious annual and quarterly reports through current reports, institutional holdings on 13F, insider transactions on Form 4, proxy statements, S-1 registrations, fund reports on N-CSR, money market data on N-MFP, fund portfolios on N-PORT, beneficial ownership on Schedule 13D and G, and the smaller offering and transaction forms including Form D, Form C and Form 144.
Underneath the financial statements there is a second capability that matters for documents which have no structured data at all: section extraction, with Risk Factors and Management's Discussion and Analysis named as the examples, plus EX-21 subsidiary exhibits and auditor information. That is the part that makes a proxy statement or an annual report usable by something other than a human reader, and it feeds the same text conversion the MCP server exposes.
The filing text pipeline is described as HTML to clean text and markdown, with full text search on top, and there is ticker and CIK lookup with industry and exchange filtering for narrowing the universe before you pull anything.
History is the other half of the pitch. The claim is that EDGAR has every filing back to 1994 and that this library exposes the complete history rather than a recent window, which is the difference between a financial database with a start date and one without.
Three releases in eight days, and a tag newer than the last push
The release record is dense at the end and sparse before it. The three most recent versions are 5.60.0 published on 2026-10-02, 5.59.1 on 2026-09-26, and 5.59.0 also on 2026-09-26. Two of those three shipped on the same day, which is the signature of a patch released to fix something and a correction released to fix the fix.
There is a dating oddity worth checking rather than assuming. The last push recorded on the default branch main is 2026-09-22, which is ten days before the 5.60.0 release was published. A release that is newer than the last commit on the branch it ships from is either built from a tag that is not on main, or built from another branch, and the repository does not say which.
The changelog arrangement matches the cadence. Instead of one growing file there is a changelog.d directory for fragments, a CHANGELOG.md for what has been assembled, and a separate CHANGELOG-archive.md for what has been retired, so a project cutting several releases a week needs a way to record changes without one file becoming the bottleneck.
For a library that parses regulatory filings, that cadence has a second meaning. Filing formats and XBRL taxonomies change, and a library pinned to an old release is a library whose parser may quietly disagree with the filings it is reading.
A type checker config, a claimed thousand tests, and a committed agent folder
The tooling choices are visible in the root and they line up with the claims. There is a pyrightconfig.json, which is the configuration for a strict type checker, backing the promise of type hints throughout. There is a .pre-commit-config.yaml, so formatting and checks run before a commit lands rather than in review. There is a .codefactor.yml next to a CodeFactor badge in the README, which is a hosted code quality service that measures the repository from outside.
The test claim is 1000 plus tests. It is stated in the feature list rather than demonstrated anywhere in the visible documentation, and the tests/ directory is present at the root, so the number is at least consistent with the layout.
Two root entries are unusual for a published library. A .claude/ directory and a CLAUDE.md file sit beside pyproject.toml, which means configuration aimed at one commercial coding agent is committed to a package that anyone installs. That is harmless and increasingly normal, but it does mean part of the repository is written for a tool rather than for users of the library.
The packaging floor is Python 3.10, with classifiers running through 3.14 and both CPython and PyPy named. The license is MIT, recorded in the manifest and in a LICENSE.txt file at the root, and the intended audiences listed in the metadata are financial and insurance, developers, and science and research, which is a more precise statement of who this is for than the README's opening paragraph makes.
Editorial conclusion
EdgarTools earns its place if you want SEC filings as typed Python objects inside your own process rather than JSON from a hosted endpoint, and if the forms you care about are in the covered set, which spans financial statements, insider forms, fund holdings and proxy material. Verify three things before you build a pipeline on it. First, the dependency policy: the HTTP client is capped because upstream went dormant, so an upstream revival would not reach you automatically. Second, the documented entry points, since a company scoped call and a global latest call return very different things and the difference is easy to miss. Third, your own compliance posture, because the library requires a real email address on every request to EDGAR and sends it as your identity. Nothing here needs a key, and nothing here needs a vendor.
Frequently asked questions
What is EdgarTools?
A Python library for reading SEC EDGAR filings as structured data, turning a filing into a typed object whose data is available as pandas DataFrames and clean text. It covers 20 plus form types including 10-K, 10-Q, 8-K, 13F, Form 4 and proxy statements, with financial statements, insider trades and fund holdings reachable in a few lines.
how to install edgartools
One command: pip install edgartools. The package requires Python 3.10 or newer, is MIT licensed, and is built with hatchling. After installing you call set_identity with an email address, because EDGAR asks every requester to identify themselves and this library sends it for you on each request.
how to use edgartools
Everything starts from a Company or a Filing, and calling .obj() returns a typed object for that form. Financial statements come from Company("MSFT").get_financials(), insider trades from Company("TSLA").get_filings(form="4").latest().obj(), and there is also a module level get_filings(form=...) that returns the most recent filing of that type across the database.
Is EdgarTools free to use?
Yes. The metadata records the MIT license and the documentation's comparison table lists the cost as free against 49 dollars a month and up for the hosted alternative. There is also no key, no signup and no rate limit tier, since the library talks to EDGAR directly from your own process rather than through a hosted endpoint.
is edgar tools safe
The documentation makes no security claim. What it states is that the library runs inside your own process with no key and no vendor, and that it is MIT licensed, described as inspectable, forkable and self hostable. What it does require is a real email address identifying you to the SEC on every request, which is a disclosure decision rather than a technical risk.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dgunning-edgartools)