osquery: turning an operating system into a SQL database
SQL powered operating system instrumentation, monitoring, and analytics.
At a glance
- What is it?
- osquery exposes processes, sockets, users and kernel state as queryable tables. It is a strong fit for engineers who already think in SQL and want host telemetry without writing a platform-specific agent, but the daemon, the fleet layer and the schema are three separate decisions.
- Who is it for?
- Adopt osquery if your team writes SQL comfortably and you want host state as rows rather than as a bespoke agent format; the daemon and shell split means you can explore interactively before scheduling anything. Do not adopt it expecting a finished fleet console, because the project states it does not endorse, recommend or test the managers listed in its README.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository last received commits 4 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem osquery solves, and the people it is aimed at
Most host instrumentation ends up as a custom agent per platform. You write one collector for Linux that parses /proc, another for macOS that shells out to system tools, and a third for Windows. Each one returns a bespoke format, and answering a new question means shipping new collector code.
osquery inverts that. The README describes it as a framework that "exposes an operating system as a high-performance relational database", where tables represent "running processes, loaded kernel modules, open network connections, browser plugins, hardware events or file hashes". The person this is built for is an engineer who already knows SQL and would rather write a join than a parser. Security and compliance work is the obvious home for it, which is why the repository topics include intrusion-detection, monitoring and security, but the same property helps anyone doing inventory or incident triage.
The trade-off is real. You gain a query language and a stable schema across three platforms. You lose the ability to model anything that is not a table, and you inherit the schema's update cadence as your own.
How the daemon, the shell and the scheduler fit together
There are two binaries to understand, and the README is explicit that the same queries can be run three ways: ad-hoc in the osqueryi shell, on a schedule via the daemon, or from a custom application through the osquery Thrift APIs.
osqueryi is the interactive path. You open it and type SQL. This is where the schema becomes legible, because a wrong column name fails immediately rather than silently at 3am in a scheduled pack.
osqueryd is the scheduled path. It runs queries on an interval across a set of hosts. The README frames this as monitoring operating system state across a fleet, which is the point at which you need somewhere for the results to go. osquery itself provides the collection and the query engine; shipping and storing results is the fleet manager's job.
The third path is the plugin and extensions API. The README states that SQL tables are implemented through it, so a table you need but do not have is a plugin, not a fork.
The example queries in the README show how much of the value comes from joins rather than single-table scans. One joins listening_ports to processes on pid to get process name, port and PID for listeners on all interfaces. Another groups arp_cache by mac and filters with HAVING count(mac) > 1 to surface ARP anomalies. That second pattern is the one worth internalising: osquery is often less about reading a table and more about finding rows that break an assumption.
Installing osquery and running a first query
The README does not inline install commands. It points at osquery.io/downloads for "the latest stable builds and for repository information and installation instructions", and says building from source is encouraged with a separate build guide. So the exact package command depends on your platform and on which repository you add; check that page rather than copying a command from a blog post.
What is stable across platforms is the interface. Once installed, the shell is where you start. This query lists local users, and it is the first thing the README shows:
SELECT * FROM users;You should see one row per account, with the columns the users table defines in the schema.
The next query is the more interesting one. It finds processes whose executable no longer exists on disk, which is a common signal for a binary that was deleted after launch:
SELECT * FROM processes WHERE on_disk = 0;An empty result is normal on a healthy machine. A populated one is worth investigating.
The third example joins two tables to answer a question neither can answer alone:
SELECT DISTINCT processes.name, listening_ports.port, processes.pid
FROM listening_ports JOIN processes USING (pid)
WHERE listening_ports.address = '0.0.0.0';This returns process name, port and PID for everything listening on all interfaces. Note the address literal: '0.0.0.0' is the all-interfaces wildcard, so this deliberately excludes listeners bound only to loopback. If you want the loopback case too, that is a different filter, and getting it wrong is the most common way this query misleads you.
Before writing anything into a schedule, run it in the shell. The schema is large and column names are not always what you would guess.
Where osquery is the wrong tool
The schema is a snapshot of host state, not an event stream. If your question is "what happened between 14:02 and 14:07", osquery answers it only to the extent that some table retained the evidence, and many do not. A short-lived process that has exited is not in processes. A connection that closed is not in listening_ports. The README's own framing, exploring operating system state, is honest about this: it is a database you query, not a log you tail.
That has a direct cost. Detection that depends on catching a transient event requires polling at an interval short enough to catch it, which raises the load the daemon places on every host. There is no interval that is both cheap and complete.
The second limit is the fleet layer. The README lists Fleet, Kolide, OSCTRL and Zentral as fleet managers and then states plainly that "the osquery project does not endorse, recommend, or test these". If you need a console, alert routing and a query history, you are choosing a second product, and that product's maintenance, licence and data handling are now part of your risk surface. Teams that pick osquery expecting a finished management platform are the ones most likely to be disappointed.
The third limit is platform asymmetry. osquery is available for Linux, macOS and Windows, but the tables are not identical across them. The README's own macOS example queries the launchd table, which has no meaning on Linux. A cross-platform pack is a portability exercise, not a copy-paste.
osquery compared with an agent that ships its own schema
The closest alternative in practice is a host agent with a fixed event model and a bundled backend, which is how most endpoint and SIEM agents work. The difference is where the abstraction sits.
With a fixed-schema agent, the vendor decides what is collected. You get a stable set of fields and a pipeline that already knows how to ship and store them. Adding a new question usually means filing a request or writing a custom check in whatever DSL the agent supports.
With osquery, the abstraction is the SQL table. Adding a new question means writing a query against tables that already exist, and the answer is available the moment you type it in osqueryi. The cost moves to the parts the project deliberately leaves open: where results go, how they are retained, and who owns the console.
That is the real decision. Choose osquery when the questions change faster than a vendor's release cycle and your team is fluent in SQL. Choose a fixed-schema agent when you want the collection, transport and storage decided for you and you are willing to accept its field list as the boundary of what you can ask.
Release cadence, licence and the cost of staying current
The README describes a numbered X.Y.Z scheme with minor releases planned roughly every two months, tracked on the Milestones page, and patch releases for unforeseen bugs. It also describes a testing window: a release is considered "in testing" while downloads are hosted and repositories updated, and is marked stable on GitHub when enough testing has occurred, which the README says usually takes two weeks. Recent releases follow that pattern, with 5.23.1 on 2026-06-24, 5.23.0 on 2026-04-25 and 5.22.1 on 2026-02-25. The last push to the repository was on 2026-09-17.
For an operator, the practical consequence is a roughly bimonthly decision point. Upgrading is not just a binary swap if you maintain custom packs, because a column rename or a table change breaks queries silently unless you test them. Budget time for that, not just for the install.
On licensing, the repository's top level contains LICENSE, LICENSE-Apache-2.0 and LICENSE-GPL-2.0, and the metadata reports the licence as NOASSERTION. The README says only that contributions are licensed as defined in the LICENSE file. That is a mixed-licence layout, and the split matters if you redistribute osquery inside a product. Read the LICENSE file and the two named licence files directly; this is not something to infer from a badge or a package manager's summary field.
Security announcements are tracked in tagged release notes and aggregated into SECURITY.md, which is the file to watch rather than a general changelog.
What to check before you commit to osquery
Three things decide whether this works for you, and all three are answerable before you deploy anything.
First, open the schema at osquery.io/schema and confirm the tables you need exist for each platform you run. The README links the schema directly, and it is the authoritative list. A detection idea that depends on a table present only on macOS is a macOS-only detection.
Second, prototype the queries in osqueryi on one representative host per platform. Run them against a machine you understand, so you can tell a false negative from a quiet host. The launchd example in the README filters on run_at_load = 1 AND keep_alive = 1, and you will only trust that filter after seeing what it returns on a real Mac.
Third, decide the fleet layer before you schedule anything. The README's list of managers is a starting point and nothing more, and the project says so. If you have no answer for where query results land, osqueryd is collecting into a void.
Editorial conclusion
Adopt osquery if your team writes SQL comfortably and you want host state as rows rather than as a bespoke agent format; the daemon and shell split means you can explore interactively before scheduling anything. Do not adopt it expecting a finished fleet console, because the project states it does not endorse, recommend or test the managers listed in its README. Before rolling it out, verify two things on your own hardware: that the tables you need exist for your platform in the current schema, and that your chosen fleet manager's licence and data path fit your environment.
Frequently asked questions
What is osquery daemon and shell?
They are the two ways to run the same queries. osqueryi is an interactive shell for ad-hoc exploration of operating system state, while osqueryd is a scheduler that executes queries to monitor state across a set of hosts. The README also notes that custom applications can launch queries through the osquery Thrift APIs.
What query language does osquery use?
SQL. osquery exposes the operating system as a relational database, and SQL tables represent concepts such as running processes, loaded kernel modules, open network connections and file hashes. The README's examples include joins between listening_ports and processes, and sub-queries against arp_cache.
How does osquery work?
Tables are implemented through a plugin and extensions API, and a variety of tables already exist with more being written. You query those tables in the shell, on a schedule, or from your own application via the Thrift APIs. The README's launchd example shows a table that is specific to macOS.
How do I install osquery?
The README does not give inline install commands. It directs readers to osquery.io/downloads for the latest stable builds, repository information and installation instructions, and separately encourages building from source using the project's build guide.
Is osquery a good tool?
That depends on whether SQL against host state matches the questions you need to answer. It is well suited to exploring and monitoring operating system state across Linux, macOS and Windows, but its tables are a snapshot rather than an event stream, and the project states it does not endorse, recommend or test the fleet managers listed in the README.
What is osquery used for?
The repository topics name intrusion-detection, monitoring, security and SQL, and the README describes ad-hoc exploration of operating system state, scheduled monitoring across a set of hosts, and launching queries from custom applications. Its example queries cover users, deleted executables, listening ports, macOS LaunchDaemons and ARP anomalies.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/osquery-osquery)