# chrislusf/gleam: a Go map/reduce DAG engine that runs standalone or across agents

> Gleam is a distributed execution system written in pure Go, where you define a DAG of datasets and computation steps and then run it locally or through a Master, Agent and Executor setup. The code is small enough to read, but the operational surface is real.

**chrislusf/gleam** — Fast, efficient, and scalable distributed map/reduce system, DAG execution, in memory or on disk, written in pure Go, runs standalone or distributedly.

- Repository: https://github.com/chrislusf/gleam
- Stars: 3,566 · Forks: 292
- Language: Go
- License: Apache-2.0
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/chrislusf-gleam

## What problem chrislusf/gleam solves, and for whom

Gleam targets a specific kind of team: one that writes Go, has outgrown a single-process loop over its data, and does not want to move the whole pipeline into Scala or Java to get distributed execution. The README describes it as a distributed map/reduce system with DAG execution, in memory or on disk, running standalone or distributedly. The unit of work is a flow: you name it, attach datasets and computation steps, and it becomes a directed acyclic graph. The README states that the default way to execute that graph is locally, and that this works in most cases. Distributed mode is the deliberate second option, not the default.

The pitch is not raw throughput. It is that the same Go program can be run on one machine during development and spread across agents in production, and that user-defined computation can be written in Go, in Unix pipe tools, or in any streaming program. If your job is to join two CSV files, tokenize text, or run a sort benchmark, the Go API is the shortest path. If your job is a long-lived SQL warehouse, this is not that.

## The DAG model: datasets, vertices and edges

The README is explicit about the abstraction. Gleam code defines the flow, specifying each dataset as a vertex and each computation step as an edge, building up a directed acyclic graph. A dataset is produced by an executor, managed by an agent, and by default travels only through memory and network rather than touching disk. The README notes that keeping data in memory lets the flow have back pressure and supports stream computation naturally. Disk persistence is optional.

Four roles make up distributed mode. The Driver is the program you write; it defines the flow and talks to Master, Agents and Executors. The Master is a single server that collects resource information from Agents and stores it transiently, and the README says it can be restarted. Agents run on any machine that can run computations, periodically send resource usage updates to the Master, start Executors when the Driver has been assigned them, and manage the datasets each Executor generates. Executors read inputs from external or previous datasets, process them, and write a new dataset.

Two design choices stand out. First, each executor runs in a separate OS process, and the README argues this avoids the garbage collection problems it attributes to other languages, because memory is managed by the OS and one machine can host many more executors. Second, the README claims the Master and Agent servers consume about 10 MB of memory and that Gleam tries to adjust required memory size automatically from data size hints. Treat the memory figure as a documented claim, not a measured result.

## Installing chrislusf/gleam and running a first flow

The README points to the Gleam Wiki Installation page as the place to get the project, and the repository carries a go.mod declaring module github.com/chrislusf/gleam with a go 1.25.0 directive. There is no retrieved release list, so the practical path is to fetch the module through the Go toolchain rather than to look for a versioned binary. The README's own word count example reads /etc/passwd, splits it into two shards, tokenizes, appends a count of one, reduces by key, sorts and prints the top five. The important structural detail is the registration step: mapper and reducer functions are registered with gio.RegisterMapper and gio.RegisterReducer before the flow is built, and gio.Init() must run first so that when the command line invokes the mapper or reducer, it executes and exits instead of building the flow again.

```go
var (
	isDistributed = flag.Bool("distributed", false, "run in distributed or not")
	Tokenize      = gio.RegisterMapper(tokenize)
	AppendOne     = gio.RegisterMapper(appendOne)
	Sum           = gio.RegisterReducer(sum)
)

func main() {
	gio.Init()
	flag.Parse()
	f := flow.New("top5 words in passwd").
		Read(file.Txt("/etc/passwd", 2)).
		Map("tokenize", Tokenize).
		Map("appendOne", AppendOne).
		ReduceByKey("sum", Sum).
		Sort("sortBySum", flow.OrderBy(2, true)).
		Top("top5", 5, flow.OrderBy(2, false)).
		Printlnf("%s\t%d")

	if *isDistributed {
		f.Run(distributed.Option())
	} else {
		f.Run()
	}
}
```

Running the binary with no flags executes the flow locally. The README states that the -distributed option needs a simple setup described later in the wiki, so expect to configure the Master and Agents before that flag does anything useful. You can also mix in external processes rather than Go functions; the README shows a variant that pipes through tr, sort and uniq:

```go
flow.New("word count by unix pipes").
	Read(file.Txt("/etc/passwd", 2)).
	Map("tokenize", mapper.Tokenize).
	Pipe("lowercase", "tr 'A-Z' 'a-z'").
	Pipe("sort", "sort").
	Pipe("uniq", "uniq -c").
	OutputRow(func(row *util.Row) error {
		fmt.Printf("%s\n", gio.ToString(row.K[0]))
		return nil
	}).Run()
```

## Where the operational model bites: single Master, transient state

The Master is described as one single server. That is a single point of failure for the control plane, and the README does not describe a standby, a quorum, or a failover path. It does say the Master stores transient resource information and can be restarted, which tells you the design tolerates a restart but not a partition. If the Master is unavailable when a Driver starts, the Driver cannot ask it for available Executors on Agents, and the flow does not begin.

A second limitation is the boundary between local and distributed execution. The default local path is the one the README calls sufficient for most cases, and the distributed path requires a setup the README defers to the wiki. That means the two modes are not interchangeable without work: a flow that runs fine with f.Run() may need network reachability between Driver, Master, Agents and Executors before distributed.Option() is useful. The README does not document rollback, partial failure handling, or what happens to a dataset when an Executor dies mid-flow. The README is also silent on security: there is no mention of authentication or transport encryption between the four roles. If you are deploying Agents across machines you do not fully control, that silence is a reason to keep the cluster inside a trusted network and read the wiki before assuming otherwise.

## Gleam versus Spark for a Go shop

The obvious alternative for DAG-shaped batch work is Apache Spark. The difference in approach is not subtle. Spark is a JVM system with a Scala core, a Python and SQL surface, and its own execution and memory model; you write transformations in one of those languages and the runtime schedules them. Gleam inverts that: the Driver is your Go program, the flow is built by method chaining in Go, and the runtime is a set of Go processes you start yourself. The README makes the readability argument directly, saying the Go code is much simpler to read than Scala, Java or C++.

That trade cuts both ways. Spark brings a managed ecosystem, a SQL layer and years of operational tooling; Gleam brings a smaller surface you can read end to end, plus the ability to drop into Unix pipes for steps that are easier to express as shell commands. If your team is Go-first and your pipeline is a handful of map, reduce, sort and join steps, Gleam keeps everything in one language and one binary. If your analysts need SQL, or you need a cluster manager that other teams already operate, Spark is the lower-friction choice. The repository also lists a sql/ directory and a plugins/ directory with cassandra_reader, kafka_reader, orc and parquet examples, so there is more surface than the README's word count suggests, but the README does not document those plugins in detail.

## Maintenance, licensing and what you are signing up for

The repository is not archived, and the last push was on 2026-09-18, days before this article. That is a live repository. It is also a repository with no retrieved releases, so there is no changelog to read and no version to pin. Upgrades therefore mean tracking the master branch and re-reading the wiki when something changes. The go.mod pins a go 1.25.0 directive and a long dependency list that includes cloud.google.com/go/storage, github.com/IBM/sarama, github.com/aws/aws-sdk-go, github.com/colinmarc/hdfs, github.com/gocql/gocql, github.com/xitongsys/parquet-go and google.golang.org/grpc, among others. Every one of those is a transitive upgrade surface you inherit. Budget for dependency review, not just for reading Gleam's own code.

The licence is Apache-2.0, which is permissive and includes an explicit patent grant. That is a standard choice for infrastructure code and places no copyleft obligation on your own Driver program. This is a description of the licence identifier, not legal advice; if your organisation has rules about patent clauses or attribution notices, run the LICENSE file past whoever handles that.

## Conclusion

Adopt chrislusf/gleam if your team already writes Go, your pipeline is a DAG of map, reduce, sort and join steps, and you want the option of running the same flow locally or through a Master and Agent deployment. Do not adopt it if you need a managed service, a Python or SQL-first workflow, or a project with published release notes to pin against; there are no retrieved releases here. Before committing, verify that your Go toolchain satisfies the go 1.25.0 directive in go.mod, read the Installation wiki page for the distributed setup, and run the examples/word_count_in_go example in both modes to confirm the Master and Agent topology works on your network.

## FAQ

### What is chrislusf/gleam?

It is a distributed map/reduce and DAG execution system written in pure Go. You define a flow of datasets and computation steps, then run it locally or through a Master, Agent and Executor setup.

### How do you install chrislusf/gleam?

The README points to the Gleam Wiki Installation page, and the repository declares the module github.com/chrislusf/gleam in go.mod, so the module can be fetched with the Go toolchain. There are no retrieved releases to pin against.

### Can chrislusf/gleam run without a distributed cluster?

Yes. The README states that the default way to execute the DAG is to run locally and that this works in most cases. Distributed mode via distributed.Option() is the second option and needs a Master and Agents.

### What licence does chrislusf/gleam use?

The repository is licensed under Apache-2.0, a permissive licence that includes a patent grant. The LICENSE file is the authoritative text.

### Does chrislusf/gleam require Java or Scala?

No. The README describes it as built in Go, with user-defined computation written in Go, Unix pipe tools, or any streaming program. The repository's go.mod declares a go 1.25.0 directive.

## Sources

- [chrislusf/gleam on GitHub](https://github.com/chrislusf/gleam)
- [Issues](https://github.com/chrislusf/gleam/issues)
- [License: Apache-2.0](https://github.com/chrislusf/gleam/blob/master/LICENSE)
- [README](https://github.com/chrislusf/gleam/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/chrislusf-gleam
