PROTOBUF
EXPLAINED
Define your data and APIs once in a schema, generate consistent code for every language, and evolve them for years without breaking existing clients. This guide takes it apart, byte by byte.
Protocol Buffers (Protobuf) is a schema-driven format for serializing structured data.
A Shared Contract
The schema is a single source of truth for your APIs and data. Every team and every language generates code from the same definitions.
Type Safety
Shared schemas let generated code catch many shape and type mismatches before data crosses a service boundary.
Compatibility
Schemas can evolve without breaking what's already deployed. Clients built against older versions keep working.
How it works
"Protobuf" refers to two things: an Interface Definition Language (IDL) for writing schemas, and the encodings those schemas produce. Every message has two possible encodings. One is a compact binary form; the other is a standardized JSON mapping. Both encodings are part of the specification, and either can go over the wire. Code generated from a schema reads and writes both.
JSON writes out a field name for every value, so you can read a JSON payload without the schema. So can any tool that speaks JSON. The binary encoding writes field numbers instead. A decoder with the same schema knows field 2 is name, so the bytes never carry the word.
The Binary section explains how those bytes are laid out.
From schema to code
You don't write any of the encoding above by hand. A compiler reads the schema and generates native code for each language you target: structs in Go, classes in TypeScript and Java, and so on. The generated code gives you typed constructors, the binary serialization, and the JSON mapping, so a message feels like any other object in your language.
Next
Basics
Messages, fields, field numbers, enums, packages, and collections, ending with generating code from a schema and putting it to work.