ArcticDB: a serverless DataFrame store for Python time series
ArcticDB is a high performance, serverless DataFrame database built for the Python Data Science ecosystem.
At a glance
- What is it?
- ArcticDB writes Pandas DataFrames straight to S3 or LMDB through a Python API backed by a C++ engine, with versioning and no server process. The catch is the licence: production and Database Service use require a paid agreement.
- Who is it for?
- Adopt ArcticDB when your workload is time-indexed Pandas frames that you want to keep in object storage without running a database server, and when the BSL terms are acceptable for your deployment. Skip it if you need open-source licensing, if your queries are ad hoc SQL joins across many tables, or if you cannot sign a commercial agreement for production.
- Can I use it commercially?
- Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
- Is it still maintained?
- Yes. The repository received new commits within the last day.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The problem ArcticDB targets: Pandas frames that outgrow a single process
A quant research desk accumulates time series. Prices, option chains, alternative data feeds. The natural container is a Pandas DataFrame, and the natural storage is Parquet on object storage. That combination breaks down in specific ways. You end up with a directory of files whose partitioning scheme you invented yourself, a metadata layer you also invented yourself, and no way to ask what a table looked like last Tuesday. Reading a date range means listing objects and filtering filenames.
ArcticDB is aimed at that gap. The README describes it as a "high performance, serverless DataFrame database built for the Python Data Science ecosystem", launched in March 2023 as the successor to Arctic. The unit of storage is a symbol, and the README states that each symbol is maintained as a separate entity with no shared data, which is what lets the system scale horizontally across symbols. A worked example in the README claims a 20-year history of more than 400,000 unique securities can live in a single symbol.
The audience is narrow and identifiable: Python users who already think in DataFrames, work with time-indexed data, and do not want to operate a database server. If your data is relational and your queries are joins, this is the wrong shape of tool. If your data is a wide frame indexed by timestamp and your queries are slices, it fits.
How ArcticDB works: symbols, libraries, and compressed reads straight from storage
The object model has three levels. An Arctic instance points at a storage location. Inside it you create libraries. Inside a library you write symbols, where a symbol is a named DataFrame. The README's quickstart shows exactly this sequence: construct adb.Arctic, call create_library('travel_data'), then index the instance with ac['travel_data'] to get the library object and call lib.write("my_data", df).
There is no server in that path. The client reads compressed data directly from storage, and the README gives the reason plainly: there is no server to overload, so your data is always available. That design decision has consequences. Concurrency control has to live in the storage layer rather than in a coordinator process, and the README asserts that persistent data structures mean once a version of a symbol has been written it can never be corrupted by subsequent updates. That is a strong claim about the storage format, and it is the mechanism behind the time travel feature, where you can read previous versions of a symbol or create snapshots.
Storage support splits into S3, LMDB and Azure Blob Storage across Linux, Windows and macOS. The README lists tested S3 backends: AWS S3, Ceph, MinIO on Linux, Pure Storage S3, Scality S3 and VAST Data S3. LMDB is the local-disk option, addressed with an lmdb:/// URI.
The schemaless behaviour is worth understanding before you rely on it. Append, update and modify operations are not constrained by the existing schema, and there is built-in support for sparse data storage. That flexibility is real, but it also means the database will not stop you from writing a frame that is inconsistent with what came before. Schema discipline becomes your responsibility, not the engine's.
Installing ArcticDB and writing your first symbol
The README gives two install paths. Prebuilt wheels exist for the current version on PyPI for Python 3.9 through 3.14, and on conda-forge for Python 3.10 through 3.14. Platform coverage differs between the two, which matters if you are on Linux arm64 or macOS x86_64: the README's table shows PyPI wheels are not available for those two combinations, while conda-forge covers them.
$ pip install arcticdbOr, if you need one of the combinations PyPI does not cover:
$ conda install -c conda-forge arcticdbImport it and point an instance at a local LMDB directory. This is the fastest way to confirm the install works, because it needs no credentials:
import arcticdb as adb
ac = adb.Arctic("lmdb:///<path>")
ac.create_library('travel_data')
print(ac.list_libraries())The list_libraries call should return a collection containing travel_data. Now build a frame with a DatetimeIndex, which is the index type the engine is built around, and write it:
import numpy as np
import pandas as pd
NUM_COLUMNS = 10
NUM_ROWS = 100_000
df = pd.DataFrame(
np.random.randint(0, 100, size=(NUM_ROWS, NUM_COLUMNS)),
columns=[f"COL_{i}" for i in range(NUM_COLUMNS)],
index=pd.date_range('2000', periods=NUM_ROWS, freq='h'),
)
lib = ac['travel_data']
lib.write("my_data", df)
data = lib.read("my_data")That snippet is the README's own example. The read returns the frame back. For S3 the constructor changes but the rest does not. The README shows two forms, one that lets AWS derive credentials and one with explicit values:
ac = adb.Arctic('s3://MY_ENDPOINT:MY_BUCKET?aws_auth=true')
ac = adb.Arctic('s3://MY_ENDPOINT:MY_BUCKET?region=YOUR_REGION&access=ABCD&secret=DCBA')The query parameters are aws_auth, region, access and secret. Note that the README does not document what happens when credentials are wrong, and it does not describe a rollback procedure for a partially written symbol.
Where ArcticDB is the wrong tool
The licence is the first constraint, and it is not a footnote. ArcticDB is released under the Business Source License 1.1. The README states that use in production, including business or commercial environments, or for a Database Service, requires a paid licence from ArcticDB Limited. The README itself notes that the BSL is not certified as an open-source licence, though it says most OSI criteria are met. Each version carries a conversion date to Apache 2.0, listed in a table in the README; the entries shown run from version 1.0 converting on March 16, 2025 through version 3.0 converting on September 13, 2025. The table in the README is truncated at version 4.0, so you cannot read the conversion date for the current 6.x line from the README alone. If your organisation has a blanket policy against non-OSI licences, this project is excluded regardless of its technical merits.
The second constraint is query shape. ArcticDB is not a SQL engine. It offers filter, aggregate and column creation with a Pandas-like syntax, and the README describes a Python-centric API. If your question is a join across five tables, or a window function over a schema you did not design, a columnar SQL engine will serve you better. ArcticDB's scaling axis is symbols, not tables you join.
The third is operational. Serverless sounds like less work, and for the database tier it is. But you still own the storage bucket, the credentials, the library layout and the decision about how many symbols to create. The README does not document a migration tool for moving a library between backends, and it does not describe how to detect or repair a symbol written by a client that died mid-write. Those gaps are where the operational cost hides.
ArcticDB compared with DuckDB and ClickHouse
The two names that come up most often next to ArcticDB are DuckDB and ClickHouse, and the three differ in ways that are not about speed.
DuckDB is an in-process analytical SQL engine. You give it Parquet, CSV or its own file format, and you query with SQL. Its storage is local files, and the mental model is a database you embed in a script or a notebook. ArcticDB's storage is remote-first: S3 or Azure Blob, with LMDB as the local case. The API is Pandas rather than SQL. If your team writes SQL and thinks in relations, DuckDB matches that. If your team writes Python and thinks in DataFrames, ArcticDB matches that. DuckDB is MIT licensed; ArcticDB is BSL with a paid production tier. That difference alone decides the question for many organisations.
ClickHouse is a server. You deploy it, it owns the storage, and clients connect over the network. It is built for SQL analytics over large tables and it has its own operational surface: replication, shards, upgrades. ArcticDB removes that surface entirely by pushing the state into object storage and the compute into the client. The trade is that you give up a query planner that can join and optimise across tables, and you take on the responsibility of organising your data into symbols.
A reasonable rule: pick ClickHouse when you need a shared SQL endpoint that many services query. Pick DuckDB when the data fits on one machine and you want SQL. Pick ArcticDB when the data is time-indexed frames that belong in object storage and you want to stay in Python.
Maintenance, upgrades and what the licence costs you
The repository is not archived, and the last push was on 2026-09-28. Recent releases include v6.27.0+man0 and v6.26.0, both dated 2026-09-14, and v6.26.0+man0 from 2026-09-09. The presence of two release streams is worth noting: the plain version numbers and the +man0 suffixed ones, which suggests a Man Group internal track alongside the public one. Nothing in the README explains the difference, so treat the +man0 builds as unverified for your purposes and pin to the plain version.
The upgrade cost is mostly Python-side. ArcticDB ships as a wheel with a compiled C++ extension, so a version bump means a new binary for your platform and Python version. The README's compatibility table is the thing to check before upgrading: if you are on Linux arm64 or macOS x86_64 and you install from PyPI, you are relying on a combination the table does not mark as available. Building from source is possible; the repository carries a Makefile with linux-release and linux-debug presets, a cpp/ directory, vcpkg configuration and a pyproject.toml configured for cibuildwheel. That is a real build system, and it is not a five-minute job.
On the licence, the practical point is that the BSL converts each version to Apache 2.0 on a stated date. That means an old version eventually becomes permissively licensed, but you would be running an old version to get there. The README directs commercial enquiries to [email protected] and references an ArcticDB Software License Agreement. This is not legal advice; the terms are short enough to read, and the production and Database Service wording is the part to read carefully.
Editorial conclusion
Adopt ArcticDB when your workload is time-indexed Pandas frames that you want to keep in object storage without running a database server, and when the BSL terms are acceptable for your deployment. Skip it if you need open-source licensing, if your queries are ad hoc SQL joins across many tables, or if you cannot sign a commercial agreement for production. Before committing, verify three things against your own data: that your Python version has a prebuilt wheel for your platform, that your S3-compatible backend behaves like the ones the README lists, and that your intended use falls outside the paid-licence trigger. A local LMDB library created with adb.Arctic("lmdb:///<path>") is the cheapest way to check the write and read path before touching object storage.
Frequently asked questions
What is ArcticDB?
It is a serverless DataFrame database for the Python data science ecosystem, launched in March 2023 as the successor to Arctic. It reads and writes Pandas DataFrames and NumPy arrays to S3, Azure Blob Storage or LMDB through a Python API backed by a C++ engine.
Is ArcticDB free to use?
The source is available under the Business Source License 1.1, but the README states that production use, including business or commercial environments, and use for a Database Service require a paid licence from ArcticDB Limited. Each version converts to Apache 2.0 on a date listed in the README's table.
How does ArcticDB compare with DuckDB?
DuckDB is an in-process SQL engine over local files, while ArcticDB is a Pandas-oriented store whose primary targets are S3 and Azure Blob Storage, with LMDB for local disk. The licensing also differs: DuckDB is permissively licensed, whereas ArcticDB requires a paid licence for production use under the BSL.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/man-group-arcticdb)