Library / SDK
protocolbuffers/protobuf avatar
protocolbuffers/protobuf

Protocol Buffers: a field guide to protoc, runtimes and release pinning

Google's language-neutral, platform-neutral mechanism for serializing structured data, combining the protoc compiler with runtimes for many programming languages.

72,078 stars16,297 forksC++License varies

At a glance

What is it?
Protocol Buffers is Google's language-neutral serialization format, split into a compiler and per-language runtimes. This guide covers how the pieces fit together, how to install protoc and a runtime, and where the release process expects you to pin.
Who is it for?
Adopt Protocol Buffers when several services or languages must exchange the same structured records and you can commit to regenerating code from .proto files as part of the build. Do not adopt it for one-off configuration files or for wire data you must inspect by hand, since the binary representation is not human-readable without the schema.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

The problem Protocol Buffers solves, and who actually needs it

Protocol Buffers is a serialization mechanism for structured data. You describe a message in a .proto file, run the protocol compiler over it, and get generated classes in your target language. Those classes write and read a compact binary encoding of the same record. The README calls the format language-neutral and platform-neutral, and the practical consequence is that a Go service, a Java service and a Python script can all read the bytes one of them wrote without agreeing on a text format or a hand-written parser.

The people who benefit most are teams running more than one language across a service boundary, or teams whose stored records change shape over time. The format is extensible: you add fields to the schema and old readers skip what they do not recognize. That property is the reason the project exists, and it is the reason schema files, not code, become the source of truth.

It is a poor fit for a single application that only ever talks to itself in one language. The generated code, the compiler step and the schema files are overhead you pay for interoperability you are not using. A configuration file read by one process does not need a compiler.

How the compiler and the runtimes divide the work

The repository is organized around one split: protoc is written in C++ and lives under src, while each language gets its own runtime directory. The README's table lists C++, Java, Python, Objective-C, C#, Ruby, PHP and Dart as in-tree directories, and points Go, JavaScript and Dart at separate repositories (protocolbuffers/protobuf-go, protocolbuffers/protobuf-javascript, dart-lang/protobuf).

The data flow is one direction. You write a .proto file. protoc parses it and emits source code for the language you asked for. Your program links the runtime library for that language and calls the generated classes. The runtime owns the binary encoding and decoding; the generated code owns the field names and types. Nothing at runtime reads the .proto file, so a schema change requires a regeneration step, not a configuration change.

Two consequences follow from that layout. First, protoc and the runtime are versioned together, which is why the README tells non-C++ users to download a pre-built protoc from the release page and pair it with the runtime for their language. Second, the compiler is not language-specific: one protoc binary can emit Java, Python and C++ from the same schema, which is what makes cross-language agreement practical.

Installing protoc and running a first schema

For non-C++ users the README gives the simplest path: download a pre-built binary from the GitHub release page. Each release carries zip packages named protoc-$VERSION-$PLATFORM.zip, and the archive contains the protoc binary plus the standard .proto files distributed with protobuf. The README notes these pre-built binaries exist only for released versions; if you want the main branch at HEAD, or you are modifying protobuf, or you are a C++ user, it points you at building from source via src/README.md.

The repository ships a worked example under examples/, including addressbook.proto and per-language programs such as add_person.py and list_people.py. The README directs new users to the tutorials in the developer guide at protobuf.dev/getting-started and to that examples directory for code. If you build through Bazel instead of calling protoc yourself, the README documents two paths. With Bazel 8 or newer and Bzlmod, you declare the dependency in MODULE.bazel:

python
bazel_dep(name = "protobuf", version = <VERSION>)

The README notes you can optionally override the repository name, for example repo_name = "com_google_protobuf", for compatibility with WORKSPACE. Legacy WORKSPACE users instead add an http_archive for com_google_protobuf and then load protobuf_deps, rules_java_dependencies, rules_java_toolchains and py_repositories, in that order. The README flags that from 30.x onward the rules_java and rules_python load statements are needed to set those up properly. If you are not on Bazel, the README does not give a build-system path here and defers C++ builds to src/README.md.

Working from main is explicitly not the easy path

The README is unusually direct about this. It states that if you work from the head revision of main, your build will occasionally be broken by source-incompatible changes and by insufficiently-tested behavior. It then says that if you are using C++ or otherwise building protobuf from source as part of your project, you should pin to a release commit on a release branch, because even release branches can experience instability between release commits.

That is a real constraint, not a disclaimer. It means the supported surface is a release commit, and the recommendation to pin applies to the release branch itself, not just to main. A team that tracks a branch head and rebuilds continuously is choosing a configuration the project documents as unstable.

The second limitation is scope. The README covers installation and points elsewhere for everything about using the format: the developer guide hosts the tutorials, the doc site hosts the complete documentation, and the version support policy lives at protobuf.dev/version-support. If you need to know how long a given language library stays supported, this repository does not answer it. You have to follow that link. The README also does not document rollback behavior for schema changes, so the safety of reverting a field is something you will not find stated here.

Where JSON and Avro differ in approach

The comparison people reach for most often is JSON, and the difference is structural rather than a matter of degree. JSON carries field names in every payload and is readable without any external artifact. Protocol Buffers carries field numbers and types, and the names live only in the .proto file. That is why the binary output is compact and why you cannot interpret it without the schema. If your data is inspected in logs, edited by hand, or consumed by a browser that expects text, JSON is the better choice and protobuf adds a decoding step for no gain.

Apache Avro takes a third position. Avro schemas are commonly stored alongside the data, typically in the file header or a schema registry, so a reader can resolve the writer's schema at read time. Protobuf expects both sides to compile from a shared .proto, which makes the schema a build-time dependency rather than a data-time one. That distinction matters for archival: a protobuf payload without the matching .proto is opaque, whereas an Avro container file carries what a reader needs. The trade-off is that protobuf avoids shipping schema metadata on every record.

Apache Thrift is closer to protobuf in shape, with an IDL and a compiler emitting multiple languages, and it bundles an RPC layer where protobuf leaves that to gRPC. The README does not compare protobuf to any of these, so treat the choice as a question about where your schema lives and who can regenerate code, not about which format is faster.

Release cadence, upgrade cost and licensing

The recent release history shows a predictable pattern: v36.0 on 2026-08-20, preceded by v36.0-rc2 on 2026-08-03 and v36.0-rc1 on 2026-07-09. Release candidates arrive roughly a month before the final tag, which gives you a window to compile your schemas against the next version before it lands. The last push to the repository was on 2026-08-20, the same day as the v36.0 release. That is recent, and the repository is not archived, but the README's own warning about instability between release commits is the more useful signal for planning than the push date.

The upgrade cost is concentrated in two places. Regenerating code is mechanical: rerun protoc or bump the Bazel dependency and rebuild. The harder part is that protoc and the runtime must move together, and the version support policy governs how long each language library stays supported, so a language you depend on may fall out of support on a different schedule than the compiler. Pinning to a release commit and testing the regenerated output against existing stored data is the only way to catch an encoding change before production does.

On licensing, the repository contains a LICENSE file at the top level and the README carries a 2008 Google LLC copyright notice, but the licence identifier is not stated in the README or the repository description. Read LICENSE directly rather than assuming a specific licence, and check the terms for any third-party code under patches/ or the vendored build files before redistributing.

Editorial conclusion

Adopt Protocol Buffers when several services or languages must exchange the same structured records and you can commit to regenerating code from .proto files as part of the build. Do not adopt it for one-off configuration files or for wire data you must inspect by hand, since the binary representation is not human-readable without the schema. Before you start, verify three things: which runtime package your language's source directory tells you to install, which protoc version matches that runtime, and whether your build system is Bazel with Bzlmod, legacy WORKSPACE, or neither, because the README documents only the first two Bazel paths and leaves CMake to src/README.md.

Frequently asked questions

What is Protobuf used for?

It serializes structured data in a language-neutral and platform-neutral binary format. You define messages in .proto files, compile them with protoc, and use the generated classes plus the runtime for your language to read and write records that other languages can also decode.

Is Protobuf better than JSON?

The README does not compare the two. The structural difference is that protobuf payloads carry field numbers and rely on the .proto file for names and types, so they are compact but not readable without the schema, while JSON is readable as-is. Which is better depends on whether your data is inspected by hand.

What is Protobuf and gRPC?

The README describes protobuf only as a serialization mechanism for structured data and does not describe gRPC. The repository's documentation covers the format, the protoc compiler and the per-language runtimes, so any relationship to gRPC is outside what this source states.

How to install protobuf?

For non-C++ users the README says the simplest way is to download a pre-built protoc binary from the GitHub release page, where each release has zip packages named protoc-$VERSION-$PLATFORM.zip containing the compiler and the standard .proto files. You then install the runtime for your chosen language following the instructions in that language's source directory.

How to use protobuf in Python?

Install the Python runtime following python/, then compile your schema with protoc and import the generated module, as the examples directory does with addressbook.proto, add_person.py and list_people.py. The README points to the tutorials at protobuf.dev/getting-started for the full walkthrough.

How to use protobuf in C++?

The README says C++ users should follow src/README.md to install protoc along with the C++ runtime, and that C++ users or anyone building protobuf from source as part of their project should pin to a release commit on a release branch. The examples directory includes add_person.cc and list_people.cc.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
  4. Release notes
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/protocolbuffers-protobuf.svg)](https://hysenlabs.com/projects/protocolbuffers-protobuf)
Community notes

Community notes