Open-source project
rizsotto/Bear avatar
rizsotto/Bear

Bear: recording compile_commands.json from builds you cannot modify

Generate compile_commands.json for any C or C++ build

6,507 stars369 forksRustGPL-3.0

At a glance

What is it?
Bear wraps an existing C or C++ build and writes the JSON compilation database that clangd needs. It is the right tool when the build system cannot export one, and the wrong tool when CMake or Meson already can.
Who is it for?
Adopt Bear when your build system cannot export a compilation database: Make, autotools, custom scripts, embedded or HPC toolchains, or a build you are only exploring. Skip it when CMake, Meson or Bazel can produce compile_commands.json directly, and skip it when you need a guarantee about which interception method ran on your platform, because the README defers that to the documentation site.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly Rust, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Bear solves: clangd with no build-system cooperation

Clang tooling needs to know how each translation unit is compiled. clangd, clang-tidy and the rest read a JSON compilation database, a file that records the working directory, the source file and the full argument vector for every compile step. Build systems that generate it themselves are the easy case. Everything else is not.

A Makefile does not emit one. Neither does an autotools tree, a hand-written shell script, or a vendor SDK shipped with a cross-compiler. You can read the build log and reconstruct the commands by hand, and people did that for years, badly and incompletely. Bear's answer is to sit in front of the build and record what actually happened. The README frames the audience directly: Make, autotools and custom script builds, third-party code you are exploring, CI jobs, and unusual toolchains such as embedded targets, HPC compilers and launchers like ccache, distcc and icecc.

The last category is the interesting one. When a wrapper rewrites the compiler invocation, a database derived from the build files would be wrong, because the build files never mention the wrapper. Bear records the invocation as it was executed, which is the only way to get the injected flags right.

Two modes: intercepting a live build and parsing a build that never ran

Bear has two distinct data paths, and the README treats them as separate features rather than one feature with a fallback.

The first intercepts the build as it runs. Bear captures the compiler invocations during execution and writes the database afterwards. This is the mode that sees wrapper-injected flags and entries for generated sources, because those only exist at run time. The second interprets commands a build system plans to run, formatting them into a database as if they had executed. The README's example is the dry-run case: parse `make -n` output. A saved CI log is the other obvious input.

The distinction matters more than it looks. A dry run tells you what the build system intends. An interception tells you what the process tree actually did, including anything a wrapper inserted. If your build uses ccache or distcc, the dry-run path is structurally blind to whatever those tools add, and the README does not claim otherwise. Pick the mode that matches the question you are asking.

The repository layout backs this up. The workspace has a `crates/` directory, a `man/` directory for the man pages the README points at, a `site/` directory behind the documentation site, and `tests/` including a `tests/dogfooding` POSIX-sh release-time harness that the Cargo workspace explicitly excludes from its glob. That last detail is a small signal about how the project tests itself against real builds.

Installing Bear and producing your first compile_commands.json

The README does not give a single install command. It says Bear is packaged for many distributions, links to a Repology page for the package name `bear-clang`, and tells you to check your distribution's package manager first. If that fails, `INSTALL.md` covers building from source. The workspace is Rust, edition 2024, so a source build goes through Cargo.

On a distribution that ships it, installation is a package-manager command. The exact name varies by distribution; the Repology page is the README's chosen way to look it up.

After installation, the README's usage line is a single prefix on your existing build command:

bash
bear -- <your-build-command>

Bear writes `compile_commands.json` to the current working directory. Bear's own options go before `--`; everything after it belongs to the build command. Run the build from the directory where you want the database to land, because the README ties the output location to the working directory, not to a flag.

If you cannot run the build at all, only a log or a dry run, the second path takes text on standard input:

bash
make -n | bear parse-sh

For anything beyond these two invocations, the README points at `bear --help` and the man pages. The documentation site is where it says limitations, known issues and platform-specific usage live, so treat the README as a starting point rather than the full reference.

Where Bear is the wrong tool

The README says it plainly: if your project uses CMake, Meson or Bazel, those tools can export a compilation database directly, and you should use that when it is available. That is not a hedge. It is the correct advice, and it is worth taking seriously.

A build system that generates the database has information Bear has to recover by observation. It knows about targets that were not built, it can emit entries without executing anything, and it does not depend on the platform's interception mechanism. Bear's interception is a runtime property, and runtime properties vary. The README's limitations section is short but pointed: each operating system and build system imposes constraints on how Bear can run and which interception method is available, and those constraints affect the final output. The specifics are delegated to the documentation site, which means anyone evaluating Bear for a new platform has to go read them rather than assume parity across Linux, macOS, the BSDs and Windows.

There is a second, quieter failure mode. Bear records what ran. If your build skipped a translation unit, or if a target failed before its compile step, there is nothing to record for it. A database produced by interception is a database of the build you actually executed, not the build the project describes. That is fine for a full build and misleading for a partial one.

Bear against compiledb, and when the difference shows

The search data around this project keeps returning one comparison: compiledb versus Bear. Both aim at the same output, and the difference is in where the information comes from.

compiledb is a Python tool in the same family, and it works from the build system's own output rather than from a live interception. Bear does both, which is the actual distinction: it can intercept a running build, and it can parse dry-run or log text through `bear parse-sh`. A tool that only reads build output cannot see flags a compiler wrapper injects at run time, because those never appear in the printed command. A tool that only intercepts cannot produce anything from a CI log you saved last week.

That is the whole comparison, and it is enough to decide. If your build runs in an environment where you can execute it and wrappers are in play, interception is the more faithful source. If you have text and nothing else, the parsing path is the only option, and both tools are in the same territory. Bear's other practical difference is packaging: the README points at Repology for distribution packages, so on many systems it arrives through the system package manager rather than a language-specific one.

Licence and the cost of keeping it current

Bear is GPL-3.0-or-later, declared in both the README badge set and the Cargo workspace metadata, with the SPDX identifier `GPL-3.0-or-later` in the workspace `Cargo.toml`. That is a copyleft licence. If you are only running the binary to generate a database for your own project, the licence governs distribution of Bear itself, not your source tree. If you intend to link Bear's crates into another program or ship a modified binary, the obligations are different, and the COPYING file in the repository root is the text that governs. This is not legal advice; read COPYING and, for anything commercial, ask someone qualified.

The upgrade cost is low by the standards of build tooling, because Bear is not part of your build. It wraps it. You can pin a version, upgrade when convenient, and a regression shows up as a worse database rather than a broken build. The release history is active: 4.2.0 on 2026-08-01, 4.2.1 on 2026-08-16, 4.2.2 on 2026-09-05, with the last push to the repository on 2026-09-18. The workspace pins its dependency versions centrally in `[workspace.dependencies]`, and the release profile uses `lto = true`, `strip = true` and `codegen-units = 1`, which is a build tuned for a small shipped binary rather than fast iteration.

One cost is not in the code. The README's limitations section points to the documentation site for the per-platform and per-build-system constraints. Keeping Bear working across a fleet of mixed toolchains means tracking that page, not just the releases.

Editorial conclusion

Adopt Bear when your build system cannot export a compilation database: Make, autotools, custom scripts, embedded or HPC toolchains, or a build you are only exploring. Skip it when CMake, Meson or Bazel can produce compile_commands.json directly, and skip it when you need a guarantee about which interception method ran on your platform, because the README defers that to the documentation site. Verify two things before you rely on it: that your distribution packages it, and that the generated database contains entries for generated sources, not just hand-written translation units.

Frequently asked questions

How do I install Bear on my system?

Check your distribution's package manager first; the README links to a Repology page under the package name bear-clang to look up what your distribution ships. If there is no package, build it from source following INSTALL.md.

Can Bear generate a compilation database without running the build?

Yes. The README gives the example `make -n | bear parse-sh`, which reconstructs the database from the build system's dry-run output, and says the same path works for a saved build log when you cannot run the build at all.

Should I use Bear if my project already uses CMake or Meson?

The README says no: CMake, Meson and Bazel can export a compilation database directly, and it recommends using that when it is available. Bear is for builds that cannot produce one themselves.

Which platforms does Bear run on?

The README lists Linux, macOS, FreeBSD, OpenBSD, NetBSD, DragonFly BSD and Windows. It also states that each operating system and build system constrains which interception method is available, and points to the documentation site for those per-platform details.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. README
  4. Releases
  5. rizsotto/Bear on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/rizsotto-bear.svg)](https://hysenlabs.com/projects/rizsotto-bear)