# Apache Cassandra: a partitioned row store for writes that outgrow one machine

> Apache Cassandra is a Java, Apache-2.0 partitioned row store whose README walks you from a tarball to a single-node cluster in four commands. This review covers the data model, the install path, the operational cost, and where a relational database remains the better answer.

**apache/cassandra** — Open source transactional distributed database. Linear scalability and proven fault-tolerance on commodity hardware or cloud infrastructure without compromising performance.

- Repository: https://github.com/apache/cassandra
- Website: https://cassandra.apache.org/
- Stars: 10,107 · Forks: 4,122
- Language: Java
- License: Apache-2.0
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/apache-cassandra

## What Apache Cassandra solves, and who it is for

The README describes Apache Cassandra as a highly-scalable partitioned row store in which rows are organized into tables with a required primary key. That single sentence carries the whole pitch. Partitioning means Cassandra can distribute your data across multiple machines in an application-transparent way, and it will automatically repartition as machines are added and removed from the cluster. The row-store half means the shape of the data stays familiar: unlike a pure key-value system, Cassandra organizes data by rows and columns, and the Cassandra Query Language is a close relative of SQL.

The audience is therefore narrow and specific. If you have a write-heavy workload that has outgrown a single server, and you are willing to design your tables around the queries you will run rather than normalizing first, Cassandra is aimed at you. If you want to keep writing joins and ad hoc aggregations, it is not. The README's own summary of CQL is the most honest framing available: think of it as SQL minus joins and subqueries, plus collections. Everything downstream of that subtraction, from schema design to how you paginate, follows from it.

## Partitioning, replication and the shape of a CQL table

The mechanism visible in the README is partition-based distribution. Every table has a required primary key, and the partitioner uses it to decide which machines hold which rows. Because the application does not address machines directly, adding or removing a node causes Cassandra to repartition automatically, and the client code does not change. This is the property that separates it from a manually sharded relational setup, where the shard map usually lives in application code and becomes the hardest thing to change later.

The cost of that design shows up in the schema. A partition key determines locality, so a poorly chosen one produces a hot node while the rest of the cluster idles. The README's example keyspace uses SimpleStrategy with a replication factor of 1, which is appropriate for the single-node walkthrough it accompanies but not a description of a production cluster. Replication class and factor are declared per keyspace at creation time, which means the decision is made once, in a statement, and is not something the README presents as revisitable. Anyone evaluating Cassandra should read that as a design commitment rather than a tuning knob.

## Installing Cassandra from the tarball and running a first query

The README's getting-started guide assumes a binary distribution. You unpack the archive and change into the resulting directory:

```bash
tar -zxvf apache-cassandra-$VERSION.tar.gz
cd apache-cassandra-$VERSION
```

Then start the server. Running the startup script with the -f argument keeps Cassandra in the foreground and logs to standard out, and the README notes it can be stopped with ctrl-C:

```bash
bin/cassandra -f
```

With the node running, the interactive client is the next step:

```bash
bin/cqlsh
```

The README shows the prompt you should see, including a connection banner naming the cluster, the cqlsh version, the CQL spec version and the native protocol version. From there, create a keyspace with a replication strategy and a replication factor, switch to it, create a table, insert a row and read it back:

```bash
cqlsh> CREATE KEYSPACE schema1
       WITH replication = { 'class' : 'SimpleStrategy', 'replication_factor' : 1 };
cqlsh> USE schema1;
cqlsh:Schema1> CREATE TABLE users (
                 user_id varchar PRIMARY KEY,
                 first varchar,
                 last varchar,
                 age int
               );
cqlsh:Schema1> INSERT INTO users (user_id, first, last, age)
               VALUES ('jsmith', 'John', 'Smith', 42);
cqlsh:Schema1> SELECT * FROM users;
```

The README states that a session resembling this output means your single node cluster is operational. Two prerequisites are named rather than pinned: Java, with supported versions listed in build.xml under the java.supported property, and Python for cqlsh, checked in bin/cqlsh by a function named is_supported_version. Because the README points at those files instead of naming versions, verify them against your checkout before you start, or the first failure you hit will be an environment problem rather than a database problem.

## Where Cassandra is the wrong tool

The absence of joins and subqueries is not a missing feature waiting on a roadmap. It is the reason the read path can stay predictable across machines, and it means any question that requires combining two tables has to be answered by the application, by denormalizing the data into a table shaped for that query, or by a separate system. Teams that treat Cassandra as a drop-in replacement for a relational database and only discover this after the schema is written tend to accumulate query logic in application code that the database used to do for them.

Transactions are the second boundary. The README's description is of a partitioned row store, and nothing in the getting-started guide suggests cross-partition transactional guarantees. If your workload depends on multi-table atomicity, that requirement has to be designed around rather than assumed away.

There is also a size threshold below which the whole exercise is unnecessary. Automatic repartitioning across machines only pays for itself when you have multiple machines and a reason to add or remove them. A dataset that fits comfortably on one server gains partition-key design work, a distributed failure surface, and an operational learning curve in exchange for scalability it does not need yet. The README does not document rollback or downgrade procedures for a cluster, so the decision to move in should be treated as one-way in practice.

## Cassandra compared with a relational database

The most useful comparison is with the relational system you would otherwise reach for, because the difference is architectural rather than a matter of degree. In a relational database you normalize first and let the query planner join at read time; the schema is stable and the queries are flexible. Cassandra inverts that. You model tables around the queries you intend to run, the partition key encodes where the data lives, and the read path does not assemble results from multiple tables. Flexibility moves from the database to the schema design phase, and it does not come back later.

The second difference is distribution. A relational database scales up until the machine runs out, then usually scales reads with replicas while writes remain on one primary. Cassandra's partitioning is built for spreading both reads and writes across commodity hardware or cloud infrastructure, and rebalancing when the set of machines changes. That is a real advantage for write-heavy workloads and a real liability for anyone whose access patterns are still being discovered, because changing a partition key after data exists is not the kind of migration the README's example workflow prepares you for.

## Maintenance, licensing and the cost of staying current

The repository is not archived, and the last push was on 2026-09-21, so the project is under current development. That matters for upgrade planning: a moving trunk means the version you deploy is a snapshot of an active codebase, and the README's own banner in the example session shows a snapshot version string rather than a fixed release. There is no releases section in the repository material, so this review cannot tell you what the current stable version is or how frequently releases land. Check the official downloads page linked from the README before you pick a version.

Licensing is Apache-2.0, per the LICENSE.txt file at the repository root and the licence badge at the top of the README. That is a permissive licence, and the practical implication for most teams is that redistribution and modification carry few obligations beyond the notice requirements the licence itself sets out. This is a description of the licence identifier, not legal advice; if you are embedding Cassandra in a product, read LICENSE.txt and NOTICE.txt in the repository rather than a badge.

The maintenance cost that the README does not address is the one that dominates in practice: partition key design, replication strategy selection per keyspace, and capacity planning as nodes are added. The getting-started guide deliberately uses SimpleStrategy with a replication factor of 1, which is the right choice for a tutorial and the wrong choice for anything with an availability requirement. Budget for schema review as an ongoing activity, not a one-time step.

## Conclusion

Adopt Apache Cassandra when your access pattern is known in advance, your writes need to spread across machines, and you can accept an eventually consistent, no-join data model. Do not adopt it for ad hoc analytical queries, multi-table transactions, or a small dataset that fits on one server, because CQL has no joins or subqueries and the README offers no rollback story. Before committing, verify the Java version against the java.supported property in build.xml, check that your Python satisfies the is_supported_version check in bin/cqlsh, and confirm that your schema's partition keys match the queries you actually run, since that choice is what decides whether the cluster distributes load or concentrates it.

## FAQ

### How do you install Apache Cassandra on Windows?

The README does not give a Windows-specific procedure. It documents unpacking the binary tarball, running bin/cassandra -f to start the server, and connecting with bin/cqlsh, and it names Java and Python as prerequisites whose supported versions are listed in build.xml and bin/cqlsh respectively. Windows users have to map those Unix-style commands onto their environment themselves.

### How do you use Apache Cassandra with Docker?

The README does not document a Docker workflow. It links to a Docker Hub badge for the cassandra image, but the getting-started guide only covers the binary tarball: unpack, run bin/cassandra -f, then connect with bin/cqlsh.

### How do you use a keyspace in Apache Cassandra?

The README's example creates one with CREATE KEYSPACE schema1 WITH replication = { 'class' : 'SimpleStrategy', 'replication_factor' : 1 }, then switches to it with USE schema1 before creating a table. Replication class and factor are set at keyspace creation.

### How do you use the Cassandra Query Language to read and write data?

The README shows the sequence in cqlsh: CREATE TABLE with a required primary key, INSERT INTO with the column list and values, then SELECT * to read the rows back. It describes CQL as SQL minus joins and subqueries, plus collections.

## Sources

- [apache/cassandra on GitHub](https://github.com/apache/cassandra)
- [Issues](https://github.com/apache/cassandra/issues)
- [License: Apache-2.0](https://github.com/apache/cassandra/blob/trunk/LICENSE)
- [Project website](https://cassandra.apache.org/)
- [README](https://github.com/apache/cassandra/blob/trunk/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/apache-cassandra
