# graphql/dataloader: batching and caching in the Node.js data fetching layer

> DataLoader is a small MIT-licensed utility that coalesces per-key loads into one batch call per tick of the event loop. It is a reference implementation of a Facebook idea, and it is deliberately narrow.

**graphql/dataloader** — DataLoader is a generic utility to be used as part of your application's data fetching layer to provide a consistent API over various backends and reduce requests to those backends via batching and caching.

- Repository: https://github.com/graphql/dataloader
- Stars: 13,392 · Forks: 518
- Language: JavaScript
- License: MIT
- Published: 2026-09-21 · Updated: 2026-09-21 · Language: en
- Canonical page: https://hysenlabs.com/projects/graphql-dataloader

## What graphql/dataloader actually solves

The problem is the N+1 request pattern. A resolver or service function knows one key at a time, so it asks the backend for one key at a time, and a request that touches 200 entities makes 200 round trips. DataLoader sits between that code and the backend. The README describes it as "a generic utility to be used as part of your application's data fetching layer to provide a simplified and consistent API over various remote data sources such as databases or web services via batching and caching." The audience is Node.js service authors, and the README names graphql-js as the common case while stating the utility is broadly useful elsewhere.

The project is honest about its lineage. It is a port of the Loader API written at Facebook in 2010, which became an implementation detail of the Ent framework and later underpinned Facebook's GraphQL server. The README also points at Haxl, Facebook's Haskell data loading library, as the same idea in another language, and invites ports to open an issue so the repository can link them. That framing matters when you evaluate it: this is a reference implementation of a concept, not a framework with an ecosystem around it.

## Batching, caching and the per-request loader rule

A DataLoader instance is constructed with a batch loading function and represents a unique cache. The README is explicit that instances are typically created per request when used inside a web server such as express, "if different users can see different things." That sentence is the whole caching policy. There is no TTL, no invalidation API, no size limit in the documented surface. If a loader is shared across requests, a value loaded for one user is served to the next one from memory. The per-request rule is not a style preference; it is the correctness boundary.

Batching is the primary feature, and the README says so directly. Loads issued within a single frame of execution, a single tick of the event loop, are coalesced and handed to the batch function as one array. The README's worked example shows four sequential-looking loads collapsing to at most two backend round trips. The scheduling mechanism is named in the source layout: enqueuePostPromiseJob. The README notes this matches the original 2010 PHP implementation.

The batch function carries two hard constraints. The returned array must be the same length as the key array, and each index must correspond to the same index in the key array. The README's example is worth reading twice: keys [ 2, 9, 6, 1 ] come back from a backend as 9, 1, 2 with 6 omitted, and the batch function must reorder and pad with null or an Error to restore alignment. Error instances are a legal value in that array, which is how per-key failures are expressed without failing the whole batch.

## Installing graphql/dataloader from npm and a first batch function

The README's Getting Started section gives one install command. It is a plain npm package with no build step for consumers; the published files list index.js, index.js.flow, index.d.ts, README.md, LICENSE and PATENTS, and the package declares main as index.js with typings at index.d.ts.

```bash
npm install --save dataloader
```

The README notes DataLoader assumes a JavaScript environment with global ES6 Promise and Map, available in all supported versions of Node.js. Next, create a loader by passing a batch function. This is the README's own example, with myBatchGetUsers standing in for your backend call.

```js
const DataLoader = require('dataloader');

const userLoader = new DataLoader(keys => myBatchGetUsers(keys));
```

Then load individual values. The loader's API is per key, but concurrent calls in the same tick are coalesced.

```js
const user = await userLoader.load(1);
const invitedBy = await userLoader.load(user.invitedByID);
console.log(`User 1 was invited by ${invitedBy}`);
```

The README's batch function example shows the alignment rule in practice, returning an Error for any key the backend did not resolve.

```js
async function batchFunction(keys) {
  const results = await db.fetchAllKeys(keys);
  return keys.map(key => results[key] || new Error(`No result for ${key}`));
}

const loader = new DataLoader(batchFunction);
```

What you should see: one call to your backend per tick containing every key requested in that tick, and a per-key Error for missing rows rather than a rejected promise for the whole batch.

## Custom batch schedules and the latency they add

The default scheduler fires on the next tick, which the README frames as adding no latency while still capturing related requests. When that is wrong, batchScheduleFn in the options object takes over. It receives a callback and is expected to call it in the immediate future.

The README gives two examples. The first collects requests over a 100ms window and, in the README's own words, adds 100ms of latency.

```js
const myLoader = new DataLoader(myBatchFn, {
  batchScheduleFn: callback => setTimeout(callback, 100),
});
```

The second is a manually dispatched scheduler that pushes callbacks into an array and exposes a schedule method, leaving dispatch to application code. The README's stated reason for both is that requests may be spread across subsequent ticks by an existing setTimeout, or that you want manual control regardless of the run loop. The trade-off is real and the README does not soften it: a wider window means bigger batches and more waiting. If your backend has a query size limit, the manual scheduler is also the only documented way to cap batch size, and you would have to implement that capping yourself.

## Where graphql/dataloader is the wrong tool

The cache is the sharpest limitation. It is per instance, unbounded within that instance, and has no documented eviction, TTL or invalidation mechanism. A long-lived loader in a process that handles many requests will hold every key it has ever loaded. The README's answer is to create instances per request, which means the cache is a request-scoped deduplication layer, not a shared cache. If you wanted a shared cache, DataLoader is not it, and the README does not pretend otherwise.

Ordering discipline is the second failure mode. If your backend returns rows in its own order, or omits missing keys, or returns a different length array, the loader will map values to the wrong keys. That is a silent data-correctness bug, not a crash, and it is easy to ship because the happy-path examples in the README use small, well-behaved responses.

Third, this is a JavaScript and Node.js library. The README positions it as a reference implementation intended to be ported to other languages, and points at Haxl for Haskell. If your service is not Node.js, you are looking at a design document, not a dependency. There is also no retry, timeout, circuit breaker or observability hook in the documented options; batching and caching are the whole surface.

## How DataLoader differs from a query builder such as Knex

The repository ships an examples directory with CouchDB.md, GoogleDatastore.md, Knex.md, Redis.md, RethinkDB.md and SQL.md. That list is the clearest statement of what DataLoader is not. It does not talk to a database. It does not build queries, manage connections or model schemas. It wraps whatever function you give it.

Knex is the useful contrast because the repository treats it as a companion, not a competitor. Knex builds and executes SQL and gives you a query builder with a connection pool. To use it with DataLoader you would write a batch function that takes the key array, issues one whereIn-style query, and reorders the rows to match the key order. DataLoader contributes the coalescing and the request-scoped cache; Knex contributes the SQL and the transport. The SQL example file in the repository exists precisely to show that pairing. If you already have a data access layer you like, DataLoader is the thin layer above it, and that is a smaller commitment than adopting a new ORM.

## Maintenance, releases and the MIT licence

The repository is not archived, and the last push was on 2026-09-16. The most recent release is v2.2.3 from 2024-12-03, preceded by v2.2.2 on 2023-02-13 and v2.2.1 on 2023-02-02. The gap between the 2024 release and the 2026 push suggests ongoing repository activity that is not shipping tagged versions, which for a library this small is a reasonable state to be in. The version in package.json is 2.2.3, so the published package matches the latest release.

Upgrade cost is low by construction. The documented API is a constructor, a load method, and a batchScheduleFn option, and the package ships index.d.ts for TypeScript consumers and an index.js.flow file for Flow. Releases are managed with changesets, and the publish script runs a prepublish step from resources/prepublish.sh. The test script chains lint, a Flow check with max-warnings 0, and jest, so a change to the loader that breaks the batch contract should fail CI before it reaches you.

The licence is MIT. The published files list also includes PATENTS. MIT is permissive and imposes no source disclosure on your application; the PATENTS file is a separate grant, and if patent terms matter to your organisation, read that file rather than assuming MIT covers it. Nothing here is legal advice.

## Conclusion

Adopt DataLoader if you are writing a Node.js service that resolves many keys from one backend per request and you can keep a loader instance per request. Do not adopt it if you need a cross-request cache, a retry layer, or a scheduler that survives process boundaries; the README documents none of those. Before wiring it in, check the two batch function constraints in the README (array length equals key length, index alignment) against your backend's response shape, because a backend that reorders or omits results silently breaks the mapping.

## FAQ

### What is graphql/dataloader?

It is a JavaScript utility for a Node.js service's data fetching layer that reduces requests to a backend through batching and caching. The README describes it as a port of the Loader API developed at Facebook in 2010, and notes it is often used when implementing a graphql-js service.

### How do I install graphql/dataloader?

The README's Getting Started section gives one command, npm install --save dataloader. It assumes a JavaScript environment with global ES6 Promise and Map, which the README says is available in all supported versions of Node.js.

### How do I use graphql/dataloader?

Create a DataLoader by passing a batch function that accepts an array of keys and returns a promise for an array of values in the same order, then call load per key. The README says instances are typically created per request when used inside a web server such as express if different users can see different things.

### How do I access graphql/dataloader?

It is an npm package, so it is accessed by installing it and requiring it in your Node.js code. The README's first example is const DataLoader = require('dataloader'); followed by new DataLoader with a batch function.

## Sources

- [graphql/dataloader on GitHub](https://github.com/graphql/dataloader)
- [Issues](https://github.com/graphql/dataloader/issues)
- [License: MIT](https://github.com/graphql/dataloader/blob/main/LICENSE)
- [README](https://github.com/graphql/dataloader/blob/main/README.md)
- [Releases](https://github.com/graphql/dataloader/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/graphql-dataloader
