# Chatty by Addy Osmani: Running WebLLM Models Privately in the Browser

> Chatty is a Next.js chat interface that runs open-weight language models on your own GPU through WebGPU, with no server-side inference. It suits engineers who want a local, offline-capable chat UI, and it is a poor fit for anyone without a capable GPU or a Chromium-based browser.

**addyosmani/chatty** — ChattyUI - your private AI chat for running LLMs in the browser

- Repository: https://github.com/addyosmani/chatty
- Stars: 836 · Forks: 105
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/addyosmani-chatty

## What Chatty Is For, and Who Should Care

Chatty is a browser application for talking to open-weight language models without sending prompts to a server. The README describes it as "your private AI that leverages WebGPU to run large language models (LLMs) natively & privately in your browser." The target user is someone who wants the interface conventions of ChatGPT or Gemini (the README says the UI is "inspired by popular AI interfaces") but does not want inference to leave the machine.

The practical audience is narrow and specific. You need a GPU with enough memory: the README states that 7B models require about 6GB of GPU memory while 3B models require around 3GB. You also need WebGPU, which the README says is enabled by default in Chrome and Edge and can be turned on in Firefox and Firefox Nightly. If either condition fails, the project is the wrong tool regardless of how the interface looks.

Where it earns its place is document Q&A. Chatty can load PDFs and other non-binary files and answer questions about them, using XenovaTransformerEmbeddings and LangChain's MemoryVectorStore, with the README stating the documents are never processed outside the local environment. For a contract, a datasheet, or a private code file, that property is the whole point.

## How WebLLM and In-Browser Embeddings Fit Together

The architecture is a client-side stack with no inference backend. The repository's package.json lists @mlc-ai/web-llm as a dependency alongside next, react, langchain, @xenova/transformers, and idb-keyval. The README credits WebLLM, HuggingFace, open source LLMs, and LangChain as the foundations.

Model weights are fetched from HuggingFace on first use and then executed through WebGPU on the local GPU. The README states that once the initial model download completes, the app works without an active internet connection. That download is the expensive step, and it is per model rather than per session.

File chat runs a separate pipeline. Documents are embedded locally with the all-MiniLM-L6-v2 model through XenovaTransformerEmbeddings, and the vectors are held in LangChain's MemoryVectorStore. The README also notes that smaller models may not process file embeddings as efficiently as larger ones, which is a useful warning: the retrieval step and the generation step have different quality ceilings.

Chat history and exports are handled in the browser. The presence of idb-keyval in the dependency list is consistent with the README's claim that you can access and manage conversation history, and the export feature writes messages to json or markdown. The README does not document where history is stored beyond the browser, nor does it describe a sync or backup path, so treat local browser storage as the only copy.

## Installing Chatty Locally with npm

The README gives a five-step source install. It requires Node.js 18 or newer and npm. Clone the repository, enter the directory, install dependencies, and start the development server:

```bash
git clone https://github.com/addyosmani/chatty
cd chatty
npm install
npm run dev
```

After npm run dev completes, the README says to open localhost:3000 and start chatting. The first thing you will see in the browser is a model download rather than an immediate reply, because weights are pulled from HuggingFace before inference can run. Expect that step to take time on a slow connection.

If you only want to evaluate the interface, the README points to the hosted instance at chattyui.com, which avoids the local toolchain entirely. That is the fastest way to check whether the UI suits you, though it does not tell you how the app behaves on your own GPU.

## Running Chatty with Docker and docker compose

The repository ships a Dockerfile and a docker-compose.yaml. The compose file defines a single service named chatty, builds from the local Dockerfile, tags the image as chatty, and maps port 3000 to 3000. The README gives both a direct Docker path and a compose path:

```bash
docker build -t chattyui .
docker run -d -p 3000:3000 chattyui
```

The compose route is shorter:

```bash
docker compose up
```

The README adds that after making changes you can run docker-compose up --build to rebuild.

One caveat is stated by the project itself. A note in the README says the Dockerfile has not yet been optimized for a production environment, and points readers to the Next.js Docker example if they want to do that work. The Dockerfile is a straightforward node:18 image that installs dependencies, copies the source, runs npm run build, exposes 3000, and starts with npm run start. That is enough to get the app running, but the maintainers are explicit that it is not a hardened deployment artifact. Note also that containerizing the app does not change where inference happens: the model still runs in the user's browser, so the container serves the UI rather than the GPU work.

## GPU Memory, Browser Support, and Other Hard Limits

The most consequential limitation is hardware. The README's own numbers, roughly 6GB of GPU memory for 7B models and 3GB for 3B models, exclude a large share of laptops, and integrated GPUs that lack WebGPU support are excluded outright. There is no CPU fallback described in the README, so a machine without a suitable GPU simply cannot run the models.

Browser support is the second constraint. WebGPU is on by default in Chrome and Edge, and the README says it can be enabled in Firefox and Firefox Nightly, linking to MDN's compatibility table for current status. If your organization standardizes on a browser without WebGPU, Chatty is not an option until that changes.

There are functional gaps too. The roadmap lists multiple file embeddings and a prompt management system as unchecked items, so embedding several files at once and selecting reusable system prompts are not available yet. The README does not document a rollback mechanism for the in-browser chat history, and it does not describe encryption of that history, so anyone with access to the browser profile can likely read it. For a tool whose selling point is privacy, that is a gap worth naming rather than glossing over.

## How Chatty Differs from Server-Hosted Chat UIs

The obvious alternative is a server-hosted interface such as Open WebUI or a locally run Ollama instance paired with a chat front end. Those tools also run models locally, but the inference process lives on the host machine and the browser is a thin client. That difference matters in two directions.

On the plus side for server-hosted tools, a single machine with a strong GPU can serve several users, models persist between browser sessions, and there is no per-browser download. Chatty gives that up. Every user downloads weights into their own browser, and the GPU work is tied to the tab.

On the plus side for Chatty, there is no server to expose. The README's central claim is that all models run client side and no server-side processing occurs. A server-hosted stack has to be secured, updated, and kept reachable; Chatty's attack surface for prompt data is the browser itself. If your reason for running models locally is to avoid standing up infrastructure, Chatty's approach is the cleaner one. If your reason is to share one expensive GPU across a team, it is the wrong approach.

## Licence, Maintenance, and Upgrade Cost

Chatty is MIT licensed, which permits commercial use, modification, and redistribution provided the copyright notice and permission notice are preserved. That is permissive and imposes no copyleft obligations on your own code. As always, the licence text in the LICENSE file governs, and this is not legal advice.

The repository is not archived, and the last push was on 2026-08-09. The most recent release listed is 0.3.0 from 2025-01-22, which added DeepSeek R1 reasoning models; 0.2.0 from 2024-07-25 added Llama3.1 and Qwen2 support; 0.1.0 was the initial release in June 2024. So model support has been added in discrete jumps rather than continuously, and the gap between the last release and the last push suggests work in the repository that has not been tagged.

Upgrade cost is mostly dependency churn. The app pins Next.js 14.2.3, React 18, LangChain 0.1.x, and @mlc-ai/web-llm 0.2.x. LangChain's JavaScript packages have moved quickly, and a major version bump there would likely require touching the file-chat code. The WebLLM dependency is the other moving part, since it determines which models can be loaded at all. Budget for re-testing model loading after any upgrade to either package.

## Conclusion

Adopt Chatty if you have a discrete or recent integrated GPU, a Chromium browser, and a reason to keep prompts and documents off a server: local file Q&A and offline chat are the two features that justify the setup. Skip it if your machine has no WebGPU-capable GPU, you need a shared team deployment, or you expect a hardened production container, because the Dockerfile is a development image by the maintainers' own note. Before committing, run Chatty on the exact browser and GPU you plan to use, load the largest model you intend to run, and confirm the initial download and memory footprint are acceptable on that hardware.

## FAQ

### Is Chatty an AI?

Chatty is a chat interface that runs open-weight language models in your browser through WebLLM and WebGPU. The AI models themselves come from HuggingFace; Chatty is the client that loads and runs them locally.

### Is Chatty a legit app?

The project is MIT licensed, published on GitHub under addyosmani/chatty, and also available as a hosted instance at chattyui.com. The README states that all models run client side and that no server-side processing occurs.

### How do I set up Chatty locally?

The README requires Node.js 18 or newer and npm, then has you clone the repository, run npm install, and start the development server with npm run dev. The app is then available at localhost:3000.

### What is the Chatty AI app?

It is a Next.js chat application that runs large language models natively in the browser using WebGPU, with support for chat history, file Q&A, voice input, and exporting conversations to json or markdown.

### How do I use Chatty?

Open the app in a WebGPU-capable browser such as Chrome or Edge, let the model download complete, and start typing. The README notes that the first model download is required before offline use is possible.

## Sources

- [addyosmani/chatty on GitHub](https://github.com/addyosmani/chatty)
- [Issues](https://github.com/addyosmani/chatty/issues)
- [License: MIT](https://github.com/addyosmani/chatty/blob/main/LICENSE)
- [README](https://github.com/addyosmani/chatty/blob/main/README.md)
- [Releases](https://github.com/addyosmani/chatty/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/addyosmani-chatty
