# ollama-js: the JavaScript client for Ollama, local and cloud

> A thin typed wrapper around the Ollama REST API with three different ways to authenticate. The manifest version is the first thing that will surprise you.

**ollama/ollama-js** — Ollama JavaScript library

- Repository: https://github.com/ollama/ollama-js
- Website: https://ollama.com
- Stars: 4,382 · Forks: 478
- Language: TypeScript
- License: MIT
- Published: 2026-10-08 · Updated: 2026-10-08 · Language: en
- Canonical page: https://hysenlabs.com/projects/ollama-ollama-js

## What the package is and what it installs as

The repository is ollama/ollama-js and the npm package it publishes is named ollama. That is not a trick, it is a name squat on a generic term that the maintainers evidently judged worth owning, and it means the install command does not match the repository name.

```
npm i ollama
```

The library is MIT licensed, written in TypeScript, and published by the Ollama organization itself with a homepage pointing at ollama.com. It has 4378 stars and 478 forks, which puts it well inside the range where the surrounding project is genuinely established rather than a weekend experiment.

The core usage is four lines.

```javascript
import ollama from 'ollama'

const response = await ollama.chat({
  model: 'llama3.1',
  messages: [{ role: 'user', content: 'Why is the sky blue?' }],
})
console.log(response.message.content)
```

The default export is an object with methods, not a class, so this is the shape for the overwhelmingly common case. The import is bare specifier style with no braces, which means the package relies on the exports map and a bundler or a modern resolver rather than on CommonJS interop.

## Streaming and the generator contract

Streaming is a flag rather than a separate method, which is the design decision that matters most here. Setting stream to true on any call changes the return type from a resolved response into an AsyncGenerator, so the same function serves both patterns.

```javascript
import ollama from 'ollama'

const message = { role: 'user', content: 'Why is the sky blue?' }
const response = await ollama.chat({
  model: 'llama3.1',
  messages: [message],
  stream: true,
})
for await (const part of response) {
  process.stdout.write(part.message.content)
}
```

Note that the await is still there. The call resolves, and what you get back is the generator. This is worth internalizing because it is easy to write await on a streaming call, forget to iterate, and end up with a program that exits cleanly having produced nothing. There is no error for that failure mode, which makes it a quiet bug rather than a loud one.

The parameter documentation is consistent about this. Every documented method repeats the same line, that stream when true returns an AsyncGenerator, and the response types for the model management calls are named ProgressResponse rather than ChatResponse, reflecting that pull and push report progress rather than content.

## Two cloud paths and one CLI step

The Cloud Models section describes offloading to Ollama's hosted models while keeping the local workflow, and it splits into two approaches that look similar and behave differently.

The first is a local Ollama that knows about the cloud. You sign in once from the command line, which is a CLI command rather than anything the JavaScript library performs, then pull a model with the cloud suffix in its name.

```
ollama signin
```

```
ollama pull gpt-oss:120b-cloud
```

After that, the client code is ordinary. You construct an Ollama instance with no arguments and call chat with the cloud model name, and the offload happens somewhere behind the local daemon rather than in your code.

```javascript
import { Ollama } from 'ollama'

const ollama = new Ollama()
const response = await ollama.chat({
  model: 'gpt-oss:120b-cloud',
  messages: [{ role: 'user', content: 'Explain quantum computing' }],
  stream: true,
})
for await (const part of response) {
+  process.stdout.write(part.message.content)
+}
```

The second path bypasses the local daemon entirely by pointing the client at the cloud host and supplying a bearer token yourself.

```javascript
import { Ollama } from 'ollama'
+
+const ollama = new Ollama({
+  host: 'https://ollama.com',
+  headers: { Authorization: 'Bearer ' + process.env.OLLAMA_API_KEY },
+})
+
+const response = await ollama.chat({
+  model: 'gpt-oss:120b',
+  messages: [{ role: 'user', content: 'Explain quantum computing' }],
+  stream: true,
+})
+
+for await (const part of response) {
+  process.stdout.write(part.message.content)
+}
```

The model names differ between the two, with the cloud-suffixed form in the local path and the bare form in the direct API path. That asymmetry is undocumented beyond the examples. If you switch between them, the failure will look like a model-not-found error rather than an authentication one, which sends people debugging the wrong thing.
+
+```
+export OLLAMA_API_KEY=your_api_key
+```
+
+The key is created in the web settings rather than issued by this repository, and the header is constructed by string concatenation in the README's example. If you are copying that into anything real, a dedicated token service is the safer route than an environment variable spliced into an Authorization header.

## The manifest says 0.0.0 and the releases say 0.6.4

The package manifest on the default branch carries version 0.0.0. The releases page carries v0.6.2 from 2025-10-30, v0.6.3 from 2025-11-13, and v0.6.4 from 2026-09-29. Both facts are true and the reason is the release workflow: the version is stamped into the published tarball during release rather than maintained in the committed file, so the branch copy stays at the placeholder.

The practical consequence is narrow but real. Anything that reads the version from the repository file, which includes some dependency auditing tools and any script you write that checks what you have installed against the source tree, will read 0.0.0. Anything that reads the installed package metadata will read the real number. Trust the installed package and the tag, and ignore the manifest when the two disagree.

The manifest also carries a description with different capitalization from the repository description, Javascript against JavaScript. That is a trivial difference, but it is the kind of thing that makes you wonder whether you are looking at the same package you read about last month.

The runtime dependency list is a single entry, whatwg-fetch at ^3.6.20. Everything else is a dev dependency, including unbuild as the bundler, vitest for tests, typescript at ^5.3.2, and eslint at ^8.29.0 with the typescript-eslint plugins at ^5. The eslint major version and the corresponding plugin major are both behind the current releases of those tools, which is consistent with the .eslintrc.cjs legacy configuration file sitting at the repository root rather than a flat config. Nothing here affects consumers, but it does tell you the toolchain is maintained on a deliberate cadence rather than continuously updated.

## Exports, the browser build, and a spec ordering issue

The exports map declares two entry points plus a wildcard. The root resolves to dist/index.cjs for require, dist/index.mjs for import, and dist/index.d.ts for types. A subpath for the browser does the same with dist/browser.cjs, dist/browser.mjs and dist/browser.d.ts. Then a single ./* entry forwards everything else to the same path.

The README documents the browser entry as its own import.

```javascript
import ollama from 'ollama/browser'
```

So the package ships a separate build for environments without node, which matters because the root build and the browser build may resolve fetch differently.

Here is the detail that would trip up a maintainer. In both the root and the browser entries, the types condition is listed last, after require and import. The node module resolution specification expects types first, because resolvers evaluate conditions in the order they appear in the object and match the first one that applies. With types last, a resolver that understands types will still eventually reach it, but a resolver that stops at the first match or that applies a shortcut for TypeScript declaration lookup can resolve the JavaScript file first and hand the type checker the wrong thing. This is the kind of ordering detail that produces a type error which looks like a bug in your own code, and the fix is to move types to the front of each entry.

The wildcard entry also means the package does not restrict deep imports, so a path that exists in dist can be reached by consumers even though it was never intended as public API. The tree contains a src directory, a test directory and an examples directory, along with both .npmignore and package-lock.json, so what actually ships is decided by that ignore file rather than by the exports map alone.

## Parameters that reveal the API's shape

The reference section for chat and generate is the most complete part of the README and it is where the API's current ambitions show through.

Several parameters exist in both methods and are worth naming because they are not obvious. think accepts a boolean or one of the strings high, medium or low, a three-way union for one toggle, and the documentation notes it requires model support. logprobs returns log probabilities for tokens and top_logprobs sets how many of them, again gated on model support. keep_alive takes either a number of seconds or a duration string with a unit suffix, and the examples given are 300ms, 1.5h and 2h45m. That last one matters more than it looks, because it is the mechanism for controlling how long a model stays resident in memory after a request, which is the main lever on a machine running a local model.

The generate method adds three parameters marked experimental and scoped to image generation models: width and height in pixels, and steps as a count of diffusion steps. Their presence in the shared request type means a multimodal model is reachable through the same endpoint as text generation.

One field deserves a flag. Inside the messages array, the documentation lists role, content, images and then tool_name, described as adding the name of the tool that was executed to inform the model of the result. Alongside it, at the top level of the request, there is a separate tools parameter described as a list of tool calls the model may make. Those two do not obviously describe the same mechanism. A conventional tool-calling flow has the model request a tool and the caller return the result, so a single tool_name field on an already-sent message looks like a partial or legacy shape rather than the standard result-passing field. If you are building agent loops on this, check the type definitions and the Ollama REST API documentation for the current field names before committing to tool_name, because the README documents it as optional and does not explain when it applies.

## Signals about maintenance

The repository has 91 open issues against 4378 stars. For a library at this adoption level that is a moderate number rather than an alarming one, and it is roughly a fifth of the fork count, which suggests the usual pattern where a large installed base files proportionally fewer issues than the star count suggests.

The default branch is main and the most recent push is dated 2026-09-29, which is the same timestamp as the v0.6.4 release, meaning the last commit and the last release are the same event. Before that there is a gap of more than ten months between v0.6.2 in late October 2025 and v0.6.3 in mid November 2025, then another ten months of nothing until v0.6.4. This is a project that ships on its own schedule rather than continuously, and it lines up with the fact that the client tracks a REST API whose changes are driven by the server side.

The topic list is javascript, js and ollama, which is thin but honest. There is no typescript topic even though the source is TypeScript and type definitions are a primary reason to use the package.

What the repository does not have is a changelog file. The releases exist on the platform with version numbers and nothing else, so there is no place in the tree to look for what changed between 0.6.3 and 0.6.4. For a library whose purpose is to track a moving server API, that is the most useful document it is missing.

## Conclusion

ollama-js is deliberately thin. There is no client-side logic to speak of, no retry policy, no connection pool, and exactly one runtime dependency, which is a fetch polyfill. What the library gives you is type definitions and an object shape that matches the REST API, which is genuinely useful, because the alternative is hand-writing that mapping in every project. Three things are worth internalizing before you build on it. The npm package is named ollama, not ollama-js, so that is what you install. The manifest on the default branch says version 0.0.0 while the tagged releases are at v0.6.4, so never read the version from the file to decide what you have. And the cloud story has two separate authentication paths, one implicit after a CLI signin and one explicit with an API key against a different host, and picking the wrong one produces authentication errors that do not explain themselves.

## FAQ

### What is Ollama software used for?

It runs large language models locally on your own machine, downloading models on demand and serving them over a local REST API. The ollama JavaScript library is the client for that API, so it is the layer your Node or browser code uses to talk to a model running on localhost or to Ollama's hosted cloud.

### Is using Ollama safe?

The models run on your own hardware and no prompt leaves your machine unless you opt into Ollama's cloud models, which the README describes as offloading to ollama.com. If you use the cloud path, your prompts go to Ollama's servers and their terms apply rather than your local isolation guarantees. The library itself is MIT licensed and adds only a fetch polyfill as a runtime dependency.

### Why does importing Ollama in JavaScript need a second entry point for the browser?

The package ships separate builds. The root export resolves to dist/index.mjs or dist/index.cjs and the browser subpath resolves to dist/browser.mjs or dist/browser.cjs, and the README tells you to import from ollama/browser to use it without node. The library has one runtime dependency, whatwg-fetch, which is a polyfill for environments whose fetch is missing or non-standard.

## Sources

- [License: MIT](https://github.com/ollama/ollama-js/blob/main/LICENSE)
- [ollama/ollama-js on GitHub](https://github.com/ollama/ollama-js)
- [Project website](https://ollama.com)
- [README](https://github.com/ollama/ollama-js/blob/main/README.md)
- [Releases](https://github.com/ollama/ollama-js/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ollama-ollama-js
