# CloudQuery: a SQL-queryable cloud asset inventory for platform teams

> CloudQuery syncs metadata from AWS, Azure, GCP and 70+ other sources into a destination of your choice using Apache Arrow. It is a strong fit for teams who want cloud config data in their own warehouse, and a poor fit for anyone expecting a hosted SaaS with a fixed schema.

**cloudquery/cloudquery** — Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.

- Repository: https://github.com/cloudquery/cloudquery
- Website: https://cloudquery.io
- Stars: 6,532 · Forks: 562
- Language: Go
- License: MPL-2.0
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/cloudquery-cloudquery

## The problem CloudQuery solves: cloud metadata stuck behind per-provider APIs

Every cloud provider exposes its own inventory API. AWS has one shape, Azure another, GCP a third, and each SaaS security tool you already pay for adds a fourth. Writing scripts against those APIs means handling pagination, rate limits, schema differences and authentication for each one separately. The README frames the alternative bluntly: "No more writing scripts to hit poorly documented APIs." CloudQuery's pitch is that it normalizes all of that into tables you query with SQL.

The intended audience is specific. The README says CloudQuery is "a cloud asset inventory built for platform teams." The three use cases it lists are cloud asset inventory, cloud security posture management (CSPM) and cloud FinOps. Those are jobs typically owned by a platform, cloud security or FinOps function, not by an application developer. If your goal is to answer "which S3 buckets are public" or "what are we spending on unused load balancers" across several providers at once, this is the category of tool you are looking at.

## How a sync works: source plugins, destination plugins and Apache Arrow

CloudQuery is built as a plugin system. A source plugin reads from a provider (AWS, Azure, GCP, GitHub, Wiz, Finout and others), and a destination plugin writes to a store. The CLI orchestrates the two. The repository layout reflects this: there is a cli/ directory, a plugins/ directory holding source and destination plugins, and a scaffold/ directory for generating new ones. A plugin SDK lives in a separate repository, github.com/cloudquery/plugin-sdk.

Data moves between plugins over Apache Arrow, which the README cites as the basis for its sync performance claim and which appears in go.mod as github.com/apache/arrow-go/v18. Destinations in the recent release list include MongoDB, and the repository topics name BigQuery alongside SQL generally. The README states that syncs run on your own infrastructure and that "Your cloud data never touches CloudQuery's servers." That is an architectural claim about where the process executes, not a statement about the plugin distribution model.

One detail worth reading carefully: the README's open source section says the framework, SDK, CLI and "some integrations" are open source, and links to zip and CSV files listing code that moved from open to closed source under MPL 2.0. So the plugin ecosystem is not uniformly open. Which plugins fall on which side is something you have to check per integration, and the README does not enumerate them inline.

## Installing the CloudQuery CLI and running a first sync

The README gives a single Homebrew command for macOS and Linux. The documentation site covers Linux and Windows install guides separately.

```bash
brew install cloudquery/tap/cloudquery
```

After that, the README points to the quickstart guide at cloudquery.io/docs/cli/getting-started for "step-by-step instructions on completing your first sync." The repository itself contains a config.yaml at the top level, which is the shape of file you edit to declare which source plugins to run and where the data should land. The README does not reproduce the full contents of that file, so treat the quickstart as the authoritative source for the exact keys.

The conceptual flow is: install the CLI, add a source plugin and a destination plugin, write a config that names both, then run a sync. Because the README does not print the config schema, do not guess at key names from this article. The documentation at cloudquery.io/docs/ is where the plugin configuration reference lives, and each source plugin has its own docs page under cloudquery.io/hub/plugins/source. A first sync against a large AWS account is not a small operation, so start with a narrow set of tables rather than everything the plugin exposes.

## Where CloudQuery is the wrong tool

The most important limitation is structural: CloudQuery is not a service. It runs where you run it, writes where you point it, and you own the scheduling, the storage and the failure handling. If your team wants a dashboard that is live the moment you sign in, with no warehouse to provision and no cron job to babysit, this is the wrong category of product. The README's privacy claim ("Your cloud data never touches CloudQuery's servers") is a benefit only if you are willing to operate the pipeline.

SQL is the interface. Every use case the README lists assumes you can write queries against normalized tables, and the CSPM material links to building dashboards with Grafana. A team without SQL skills, or without an existing BI or query layer, gets a pile of synced tables and no way to read them.

Plugin coverage is uneven by design. The README claims 70+ sources, but the open source section makes clear that some integrations are closed source, and the hub is where you check coverage per provider. If a provider you depend on has no plugin, or only a closed one you cannot inspect, CloudQuery does not help you. Finally, the README does not document rollback behavior for a failed or partial sync, and it does not describe what happens to destination tables when a source schema changes between plugin versions. Those are questions to raise with the documentation before you put this in a production path.

## CloudQuery vs Airbyte and AWS Config

The two comparisons people search for most are against Airbyte and against AWS Config, and the difference in approach is real.

Airbyte is a general-purpose ELT connector framework. Its connectors target application and database data (SaaS APIs, Postgres, warehouses) and the emphasis is breadth of generic connectors. CloudQuery's plugins are domain-specific: the README describes "Specialized plugin coverage" with "normalization, rate limit handling, and more" for cloud infrastructure, security and FinOps sources. If your problem is moving Salesforce records into Snowflake, Airbyte is the more natural fit. If your problem is enumerating every EC2 instance, IAM role and S3 bucket across accounts into queryable tables, CloudQuery's plugins are built for that shape of data.

AWS Config is a managed AWS service, and it only covers AWS. CloudQuery's README positions itself across "AWS, Azure, GCP, and 70+ cloud and SaaS sources." That multi-cloud and multi-source unification is the core selling point, and it is also the reason the tool is more work: you assemble the pipeline instead of consuming a console. AWS Config also gives you continuous evaluation against rules inside AWS; CloudQuery gives you tables and expects you to write the checks, for example as SQL against your warehouse.

## Licence, maintenance and the cost of keeping plugins current

The repository is licensed MPL-2.0 and is not archived. The last push was on 2026-09-21, and the most recent releases listed are plugins-destination-mongodb v3.2.0 on 2026-09-17 and cli v6.42.3 on 2026-09-15. Those dates indicate current activity, and the release cadence matters here because each source plugin versions independently. Upgrading the CLI does not automatically upgrade your plugins, and plugin versions are pinned in your config.

MPL-2.0 is a file-level copyleft licence: modifications to MPL-covered files must be made available under the same licence, while larger works that combine MPL files with other code can be licensed differently. That is a general description of the licence family, not legal advice for your situation. The practical wrinkle is the one the README itself raises: the framework, SDK, CLI and "some integrations" are open source, and code that moved to closed source is enumerated in the linked zip and CSV files rather than in the README text. Before you build a dependency on a specific plugin, confirm its licence and whether its source is in this repository.

The ongoing cost is upgrade work. Cloud provider APIs change, new resource types appear, and plugin releases follow. Every plugin version bump is a chance that a table's schema changed under a query you wrote. Budget for pinning plugin versions deliberately and testing syncs against a non-production destination before rolling changes out.

## Conclusion

Adopt CloudQuery if you already run a warehouse or database and want cloud configuration data inside it under your own control. Do not adopt it if you need a managed service with zero infrastructure, or if your team has no SQL skills. Before committing, verify that a source plugin exists for every provider you care about, check which plugins are open source versus closed, and confirm your destination plugin supports the tables you plan to sync.

## FAQ

### What is CloudQuery and what does it do?

CloudQuery is a cloud asset inventory built for platform teams. It syncs cloud infrastructure metadata from sources such as AWS, Azure and GCP into a destination you control, where you query it with SQL.

### Is CloudQuery open source?

The README states that the CloudQuery framework, SDK, CLI and some integrations are open source under MPL-2.0, and links to zip and CSV files listing code that moved from open to closed source. Not every integration is open source, so check the specific plugin you need.

### Is CloudQuery free?

The repository is licensed MPL-2.0 and the README does not describe a paid tier or pricing model. It also does not say that every plugin is free, and some integrations are closed source, so the cost of a given setup depends on which plugins you use.

### How does CloudQuery compare with Airbyte?

Airbyte is a general-purpose ELT connector framework, while CloudQuery's plugins are specialized for cloud infrastructure, security and FinOps sources with normalization and rate limit handling. CloudQuery is aimed at turning cloud asset metadata into SQL-queryable tables rather than at generic application data replication.

### How does CloudQuery compare with AWS Config?

AWS Config is a managed AWS-only service, while the CloudQuery README positions the tool across AWS, Azure, GCP and 70+ cloud and SaaS sources. CloudQuery runs on your infrastructure and gives you tables to query, rather than a console with built-in rule evaluation.

## Sources

- [cloudquery/cloudquery on GitHub](https://github.com/cloudquery/cloudquery)
- [License: MPL-2.0](https://github.com/cloudquery/cloudquery/blob/main/LICENSE)
- [Project website](https://cloudquery.io)
- [README](https://github.com/cloudquery/cloudquery/blob/main/README.md)
- [Releases](https://github.com/cloudquery/cloudquery/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/cloudquery-cloudquery
