Open-source project
canonical/cloud-init avatar
canonical/cloud-init

cloud-init: The Standard Tool for Cloud Instance Initialization

Official upstream for the cloud-init: cloud instance initialization

3,826 stars1,157 forksPythonNOASSERTION

At a glance

What is it?
cloud-init is the canonical method for bootstrapping a new cloud instance at first boot, reading cloud metadata and user-supplied configuration to set up networking, storage, SSH keys, and installed packages. It is the default initialization system on most major Linux distributions shipped by public cloud providers.
Who is it for?
cloud-init is the right tool when you need repeatable, provider-agnostic instance initialization at boot time. It is the wrong choice for ongoing configuration management of running systems.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What cloud-init Does During Instance Boot

cloud-init solves a problem that arises at the intersection of immutable images and dynamic cloud environments: a disk image that boots in one environment must be configured differently for each instance it creates. The image cannot know in advance what hostname it will get, what SSH public key should grant access, or what network address it will use.

cloud-init runs early in the boot sequence, identifies the cloud or platform it is running on, reads three categories of data, and applies configuration based on them. The first category is cloud metadata, provided by the cloud platform, which contains instance-specific details such as the instance ID and network configuration. The second is user data, which is optional and supplied by the person launching the instance, typically containing packages to install, files to write, or commands to run. The third is vendor data, also optional, supplied by the cloud provider to customize behavior for their platform.

After reading this data, cloud-init applies the configuration. The README states this may involve setting up network and storage devices, configuring SSH access keys, and many other aspects of a system. The last push was on 2026-09-25, and the most recent release is 26.2 from 2026-07-29.

Datasources: Matching the Platform to the Configuration Source

cloud-init uses the concept of a datasource to identify where to read cloud metadata from. Different cloud providers and hypervisors expose their metadata through different mechanisms: some provide an HTTP endpoint at a link-local address, others use a special volume attached at boot, and some use kernel command-line parameters.

The datasource detection happens automatically at boot. cloud-init runs a script called `ds-identify` (visible in the repository's `tools/ds-identify` file referenced in the Makefile) to determine which datasource is active before any cloud-specific configuration runs.

The documentation at docs.cloud-init.io lists the supported datasources, which cover all major public cloud providers as well as private cloud systems and bare-metal installations. If a cloud or distribution is not supported, the README advises contacting that distribution or provider and directing them to the project.

For developers working on the cloud-init codebase itself, the project uses a Python test suite:

bash
python3 -m pytest -v tests/unittests cloudinit

The project depends on jinja2 for template rendering, pyyaml for configuration parsing, jsonschema for cloud-config validation, and requests for HTTP communication with metadata services.

User Data and the cloud-config Format

The most common way to customize an instance is through user data passed to the instance at launch time. cloud-init supports several user data formats. The cloud-config format, which starts with `#cloud-config`, is a YAML document that declares the desired state: which packages to install, which users to create, which files to write, and which commands to run.

The distinction between cloud-init and cloud-config matters for day-to-day use. cloud-init is the initialization system: the process that runs at boot, reads data, and applies configuration. cloud-config is one specific format that user data can take, expressed as YAML. The `#cloud-config` header is how cloud-init recognizes that a block of user data should be parsed as cloud-config YAML rather than treated as a script.

The `runcmd` module, which appears in the Google search data for this project, is a cloud-config directive that specifies a list of commands to run during the initialization phase. Packages listed under the `packages` key are installed through the distribution's package manager. Network configuration can be supplied through user data using the `network` key.

cloud-init uses jinja2 for template processing, which means user data can reference instance metadata variables before configuration is applied.

cloud-init Compared to Ansible

cloud-init and Ansible solve different parts of the configuration management problem. cloud-init runs once, at first boot, and configures the instance from a blank state. It has no persistent agent, no inventory of hosts, and no concept of idempotent state enforcement over time. Once an instance is up and its initialization phase is complete, cloud-init has done its job.

Ansible is a configuration management and automation tool designed to run against existing, running systems. It uses an inventory of hosts, connects over SSH, and applies playbooks that can be re-run repeatedly. Ansible is appropriate for enforcing a desired state across many instances over their lifetime, not just at creation.

The two tools are often used together: cloud-init handles the first-boot configuration that makes an instance reachable and minimally configured, and Ansible takes over for ongoing configuration management. Using Ansible alone at first boot is possible but requires the instance to already have SSH access, which cloud-init is often responsible for setting up through SSH key injection.

Running cloud-init clean resets the instance's cloud-init state, which allows re-running the initialization process. This is mainly used during development and testing of cloud-config configurations, not in production workflows.

Known Limitations: One-Shot Execution and Debugging

cloud-init is designed for one-shot execution at first boot. By default, it records which modules have run and skips them on subsequent boots. This behavior is intentional: re-running cloud-init on a running system would overwrite configurations that may have changed since the initial setup. Developers who need to test their user data must either use a fresh instance or run `cloud-init clean` to reset the run state.

Debugging cloud-init problems requires examining log output. The cloud-init documentation references `/var/log/cloud-init.log` and `/var/log/cloud-init-output.log` as the primary log files. When cloud-init fails early in boot, the instance may come up with missing network configuration or no SSH keys, making it inaccessible. The failure has to be diagnosed from the cloud provider's console output or serial console access.

User data is limited in size by the cloud provider. AWS, for example, imposes a 16 KB limit on user data. Large configurations must be staged through a URL or split across modules. The cloud-config YAML format, while expressive, has no mechanism for conditionals or loops in its basic form; complex logic requires either shell scripts in `runcmd` or a configuration management layer applied after cloud-init completes.

Disabling cloud-init entirely on a cloud instance is possible but has side effects. On many providers, cloud-init is responsible for injecting SSH authorized keys and configuring the network at boot. Disabling it without an alternative key injection mechanism will lock out SSH access to the instance.

Release History, Licensing, and Maintenance

cloud-init follows a version numbering scheme tied to the year. Release 26.2 was tagged on 2026-07-29, and release 26.1 was tagged on 2026-02-28. The previous release was 25.3 in September 2025, indicating a cadence of roughly two releases per year.

The repository is published under a dual license. The `LICENSE-Apache2.0` and `LICENSE-GPLv3` files are both present, and the repository's license field is listed as NOASSERTION in the metadata, reflecting the dual-license situation rather than a single SPDX identifier. Downstream distributions that package cloud-init should verify which license terms apply to their use case.

The canonical repository is at github.com/canonical/cloud-init, maintained by Canonical. The project accepts contributions through GitHub pull requests, with contribution guidelines described in the documentation at docs.cloud-init.io. Questions can be directed to the `#cloud-init` channel on Matrix or through GitHub Discussions.

The package requires Python 3, jinja2, oauthlib (for the MAAS datasource and webhook reporting), configobj, pyyaml, requests, jsonpatch, and jsonschema. These are listed in `requirements.txt` and are standard dependencies for a Python-based configuration tool.

Editorial conclusion

cloud-init is the right tool when you need repeatable, provider-agnostic instance initialization at boot time. It is the wrong choice for ongoing configuration management of running systems. Before deploying, verify that your cloud or hypervisor appears in the supported datasources list at docs.cloud-init.io, and confirm whether disabling it will break your provider's key injection or network configuration workflow.

Frequently asked questions

Should I disable cloud-init?

Disabling cloud-init on a cloud instance will prevent it from injecting SSH authorized keys and configuring networking at boot on many providers. Before disabling it, confirm which features your cloud platform relies on cloud-init to provide. The cloud-init documentation at docs.cloud-init.io describes how to disable specific modules or the entire service safely.

What is the difference between cloud-init and cloud-config?

cloud-init is the initialization system that runs at boot and applies configuration to a new instance. cloud-config is one specific format for user data, written as a YAML document starting with #cloud-config, which cloud-init recognizes and parses. Other user data formats, such as shell scripts, are also supported by cloud-init.

What are the key differences between cloudinit and ansible?

cloud-init runs once at first boot to configure a new instance from scratch, with no persistent agent and no inventory management. Ansible is a configuration management tool designed to run against existing running systems, applying playbooks repeatedly over time. The two are often used together, with cloud-init handling first-boot setup and Ansible managing ongoing configuration.

How do I configure cloud-init?

Configuration is supplied through user data passed to the instance at launch time. The cloud-config format, a YAML document starting with #cloud-config, lets you declare packages to install, files to write, users to create, and commands to run. Detailed configuration reference is at docs.cloud-init.io.

Official sources

  1. canonical/cloud-init on GitHub
  2. Issues
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/canonical-cloud-init.svg)](https://hysenlabs.com/projects/canonical-cloud-init)