# elastic/eland: pandas-style DataFrames on top of Elasticsearch

> Eland gives Python users a pandas-like DataFrame whose data stays in Elasticsearch, plus a CLI for uploading trained models. It is a good fit for analysts who already run Elasticsearch and a poor one for anyone without a cluster.

**elastic/eland** — Python Client and Toolkit for DataFrames, Big Data, Machine Learning and ETL in Elasticsearch

- Repository: https://github.com/elastic/eland
- Website: https://eland.readthedocs.io
- Stars: 693 · Forks: 113
- Language: Python
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/elastic-eland

## The gap Eland fills between pandas and Elasticsearch

Pandas assumes the data fits in memory on the machine running the code. Elasticsearch assumes the data is large and distributed. Getting from one to the other usually means writing scroll queries, paging through hits, and rebuilding a DataFrame by hand, or exporting to CSV and losing the query layer entirely. Eland closes that gap with a DataFrame object that looks like pandas but keeps the data in the cluster. The README describes it as "a Python Elasticsearch client for exploring and analyzing data in Elasticsearch with a familiar Pandas-compatible API", and the key sentence is that "the data resides in Elasticsearch and not in memory". The audience is narrow but real: Python users who already operate an Elasticsearch cluster and want notebook-driven analysis over an index too large to load locally. The second audience is machine learning engineers who need to push a trained model into Elasticsearch for inference, which Eland also handles.

## How the DataFrame defers work to the cluster

An Eland DataFrame is a wrapper around an index pattern, not a materialized table. You construct it with a host string or an Elasticsearch client plus an index pattern, and every subsequent operation is translated into Elasticsearch queries rather than executed in Python. That is why the README can point at df.info() and show a memory usage of 80.0 bytes next to an Elasticsearch storage usage of 5.043 MB: the local object holds metadata, and the rows stay on the server. Filtering follows the same path. A boolean expression such as df.Carrier == "Kibana Airlines" combined with a numeric comparison and df.Cancelled == True becomes a query, and only the result set is fetched. The practical consequence is that the cluster does the work, so latency and cost move to Elasticsearch, and the local machine only sees what you ask for. It also means the DataFrame is not a snapshot. If the index changes underneath, a later operation sees the new data.

## Installing Eland and running a first query

The package is on PyPI, so the install is a single pip command. The README gives this exact form, using the module invocation rather than the bare pip binary.

```bash
python -m pip install eland
```

If you plan to upload NLP models rather than tabular ones, the README says to install the PyTorch extras instead. That pulls in the PyTorch version Eland expects, which matters because the README states you need to install the appropriate version of PyTorch to import an NLP model.

```bash
python -m pip install 'eland[pytorch]'
```

On Debian-based systems the README warns that transitive dependencies may need system packages first, and lists them explicitly. Other distributions will need different package names and a different package manager.

```bash
sudo apt-get install -y \
  build-essential pkg-config cmake \
  python3-dev libzip-dev libjpeg-dev
```

Once installed, connecting to a local node takes one line. The README's example connects to an Elasticsearch instance on http://localhost:9200 and binds the flights index pattern.

```python
import eland as ed

df = ed.DataFrame("http://localhost:9200", es_index_pattern="flights")
print(df.head())
```

What you should see is a table shaped like a pandas DataFrame, with the column headers taken from the index mapping. The README's sample output shows 27 columns and a timestamp column typed as datetime64[ns], which tells you Eland maps Elasticsearch field types onto pandas dtypes rather than returning raw JSON. For Elastic Cloud, the README passes an Elasticsearch client instead of a URL, constructed with cloud_id and basic_auth. If you would rather not install anything, the project publishes a Docker image at docker.elastic.co/eland/eland, which the README shows running interactively with --network host.

## Uploading trained models, and the version rule that bites

The second half of Eland is model upload. The README says Eland provides tools to upload trained machine learning models from scikit-learn, XGBoost and LightGBM into Elasticsearch, and the Docker example invokes eland_import_hub_model with a --url, a --hub-model-id and a --task-type to bring a Hugging Face model into the cluster for named entity recognition. That is the piece most likely to trip you up, because the compatibility rules differ by feature. For NLP with PyTorch, the README requires the Eland minor version to match the minor version of the Elasticsearch cluster. For everything else, matching the major version is sufficient, and Eland 8.x is the line for Elasticsearch 8.x clusters. A team that upgrades its cluster without upgrading Eland can therefore break model import while ordinary DataFrame queries keep working, which is an unpleasant failure mode to debug because the symptom appears in one code path only.

## Where Eland is the wrong tool

Eland requires an Elasticsearch 9+ cluster for the current line. If your data is in a warehouse, a Parquet file, or Postgres, Eland has nothing to offer, and reaching for it means standing up a cluster you did not need. The pandas compatibility is also partial by construction. The README says the DataFrame "has the same API as pandas.DataFrame except all data is in Elasticsearch", and points readers at a separate DataFrame API reference for the supported surface. Anything outside that surface has no server-side translation, so it either raises or forces you to pull data locally, which defeats the purpose. There is a version trap on the pandas side too: the README states support for pandas 1.5 and 2, so a project already on pandas 3 is outside the documented range. Finally, the deferral model means every operation is a network round trip. Iterating row by row, or looping over many small filters, turns into many queries, and the cluster absorbs that load. Eland is for aggregate work over large indexes, not for tight interactive loops over small ones.

## How Eland differs from querying Elasticsearch directly

The obvious alternative is the Elasticsearch Python client that Eland itself builds on. Eland uses the low level client for connections, and the README links to that client's own connection and authentication options, so the two share a transport layer. The difference is the programming model. With the plain client you write a query body, send it with search, and parse hits yourself, which gives you full control over the query DSL, aggregations, and pagination but leaves you assembling DataFrames by hand. Eland trades that control for a DataFrame interface: you write df[df.AvgTicketPrice > 900.0] and it becomes a query. The trade is real in both directions. If you need an aggregation shape Eland does not express, the plain client is the better tool. If you need a columnar workflow over an index and want to stay in a notebook, Eland saves you from writing the translation layer. Note that Eland does not replace the client, it sits on top of it, so you can pass a configured Elasticsearch instance into an Eland DataFrame and keep your existing auth setup.

## Licence, releases and what maintenance costs

Eland is Apache-2.0, and the setup.py classifier confirms it as an OSI-approved Apache Software License. The practical implication for most teams is that you can ship it inside a commercial product without a copyleft obligation, but the NOTICE.txt file in the repository root is part of the distribution and should travel with any redistribution. That is a packaging detail, not legal advice, and anyone with an actual compliance question should read LICENSE.txt and NOTICE.txt rather than a summary. On maintenance, the repository is not archived and the last push was on 2026-06-16, so it is being worked on. The release cadence visible in the recent tags is uneven: v9.2.0 landed on 2025-10-30, while v9.0.1 and v8.18.2 both landed on 2025-04-30. Two parallel lines, 9.x for Elasticsearch 9 and 8.x for Elasticsearch 8, means a team on the older cluster is tracking a branch that sees fewer releases. The upgrade cost is dominated by the version-matching rules rather than by API churn: you check your cluster major version, check your pandas version against the documented 1.5 and 2 range, and for NLP work check that the Eland minor matches the cluster minor.

## Conclusion

Adopt Eland if your data already lives in Elasticsearch 9+ and you want to keep pandas-style analysis without pulling the index into local memory; the same package installs with python -m pip install eland. Do not adopt it if you have no cluster, need pandas 3, or rely on pandas APIs the DataFrame reference does not list, since unsupported methods will not silently fall back to local execution. Before committing, check that your cluster and Eland major versions match, that your pandas version is 1.5 or 2, and that the operations you depend on appear in the Eland DataFrame API documentation.

## FAQ

### What is elastic/eland?

It is a Python client and toolkit that exposes an Elasticsearch index through a pandas-compatible DataFrame API, keeping the data in Elasticsearch rather than in local memory. It also provides tools to upload trained models from scikit-learn, XGBoost and LightGBM into Elasticsearch.

### How do I install elastic/eland?

The README gives python -m pip install eland for the base package and python -m pip install 'eland[pytorch]' when you need to upload NLP models. A conda-forge package is also available via conda install -c conda-forge eland.

### Which Elasticsearch and Python versions does elastic/eland support?

The README states support for Python 3.10 through 3.13, pandas 1.5 and 2, and Elasticsearch 9+ clusters, with Eland 8.x used for Elasticsearch 8.x. For the NLP with PyTorch feature the Eland minor version must match the cluster minor version; other features only need the major version to match.

### Can I use elastic/eland without installing it?

The README documents a Docker image at docker.elastic.co/eland/eland that can be run interactively, and shows running eland_import_hub_model through it without an interactive shell. The Dockerfile builds on python:3.13-slim.

## Sources

- [elastic/eland on GitHub](https://github.com/elastic/eland)
- [License: Apache-2.0](https://github.com/elastic/eland/blob/main/LICENSE)
- [Project website](https://eland.readthedocs.io)
- [README](https://github.com/elastic/eland/blob/main/README.md)
- [Releases](https://github.com/elastic/eland/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/elastic-eland
