# Windows-Copilot-API wraps your real Copilot session, and publishes it on port 8000

> This project reverse-engineers the consumer Copilot web chat and republishes it two ways: a Python client and a local server that speaks the OpenAI format. The mechanism is a browser sign-in that saves a real profile and a token, refreshed from a warm-up message that also clears the human-verification check. The API needs no key, which also means it authenticates nobody.

**sums001/Windows-Copilot-API** — Reverse engineered Windows Copilot into an OpenAI-compatible API. Access GPT-4 and GPT-5 models through a simple REST interface without API keys or billing.

- Repository: https://github.com/sums001/Windows-Copilot-API
- Stars: 1,247 · Forks: 404
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/sums001-windows-copilot-api

## A consumer chat session behind an OpenAI-shaped API

The claim is that there is no API key, no credits and no paid plan, because the underlying resource is the free Copilot chat you are already signed in to at copilot.microsoft.com. The project automates that web experience and republishes it in two forms.

The first is a library. You construct a client, it loads your saved session, and `client.chat("Hi")` returns a reply object carrying the text and a conversation id you can pass back to keep a thread going. A streaming method yields the answer as it is typed.

The second is a server on `http://localhost:8000/v1` that speaks the OpenAI wire format, so the official `openai` SDK and any OpenAI-compatible application work by pointing `base_url` at it. Two endpoints are exposed: a POST for chat completions supporting streaming and an optional `conversation_id`, and a GET for models.

The project states plainly that it is unofficial, not affiliated with or endorsed by Microsoft, that it automates the consumer Copilot web experience for personal use, and that it should be used responsibly and within Microsoft's terms.

## Sign-in stores a real browser profile and mints a token with a throwaway chat

Authentication is a Playwright browser that opens, waits for you to log in with a Microsoft or a Google account, and closes itself when it detects success. The setup block that gets you there is three commands:

```bash
# Install dependencies
pip install -r requirements.txt

# Install the browser Playwright needs (one-time)
playwright install chromium

# Sign in once: a browser opens, log into your Microsoft or Google account
python -m copilot login
```

Then the interesting part happens: after sign-in it sends one short warm-up message, and that single request mints the chat token and passes the human-verification check in the same step. A brief finishing-setup message appears, and a tiny throwaway conversation lands in your chat history.

That last detail is worth pausing on. The setup step is not a silent handshake; it writes something visible into the account you use every day.

If a checkbox appears during login, you are asked to click it in that window, so the interactive step stays human. The sequence is logged to `session/login.log` when something goes wrong, and a bundled diagnostic is described as both fixing common issues around captcha and clearance and producing a shareable report.

The result is a session directory, described as git-ignored and never shared, that is reused on every run, which is why the first request works immediately. What it holds is a signed-in browser profile plus a cached token for your own account. That is the most sensitive artefact this project creates, and its safety rests entirely on the directory not being copied or shared.

## A container cannot renew its own clearance, so it fails every half hour

The Docker path exists, and it has an inherent limitation that the page states rather than hides.

You must sign in on the host first, because the login step opens a visible browser and a headless container cannot do that. Once `session/` is populated, `docker compose up --build` bind-mounts it and reuses the clearance that was earned on the host.

The container can refresh the chat token headlessly. It cannot earn fresh clearance without a visible browser. So when the clearance expires, quoted as roughly thirty minutes, the server returns a `503` and the fix is to re-run the login command on the host to refresh `session/`.

That turns into an operational loop: a periodic interactive login on the host every half hour, for a server that is otherwise running unattended. It is a reasonable trade for a local tool and an awkward one for anything scheduled.

The compose file also sets `restart: unless-stopped`, so the container comes back after a reboot and then fails on the first request until someone signs in again.

## The compose file publishes port 8000 and the API key is ignored

The security surface of this project is short and sharp, and it lives in the deployment rather than the code.

The compose file maps `8000:8000` with no host address restriction, and sets `HOST` to `0.0.0.0` inside the container. Combined with the documented behaviour that the API key is required by the SDK but ignored by the server, the result is an unauthenticated HTTP endpoint that will spend your signed-in Copilot session, reachable by anything that can route to your host.

The fix is one character, and it is worth being explicit about it: bind the published port to loopback, the way the same file already restricts nothing else. The instructions do say to prefer the local address, and the Python server path in the page prints `127.0.0.1:8000`, so the default-by-hand is fine. The Compose path is the one to check.

There is a second, quieter exposure. The session directory holds a signed-in browser profile for your Microsoft or Google account. It is git-ignored, which prevents the obvious accident, and nothing in the documentation suggests it is encrypted at rest.

## Rate limits are two numbers in a compose file, described as self-imposed

The compose environment sets `RATE_LIMIT_RPM` to `12` and `RATE_LIMIT_BURST` to `4`, with a comment pointing at `server/config.py` and describing the limits as self-imposed and tunable to your account.

Self-imposed is the honest word. There is no quota negotiated with a provider here, because there is no provider contract; the numbers are the project's guess at what a consumer session tolerates. Raising them is a decision the operator makes about their own account, and the defaults are a starting point rather than a limit enforced from the other side.

That framing matters for how you use the thing. If you point an agent loop or a batch job at this server, the ceiling you will actually hit is your account's tolerance, and the ceiling you will hit first is whatever Microsoft decides it is.

The interface is small enough that a self-imposed limiter is proportionate: one chat endpoint, one models endpoint, one model id.

## One model id is exposed, whatever the chat window is running

The project's one-line summary advertises access to GPT-4 and GPT-5 models through the API. The endpoint table tells a different, smaller story: `GET /v1/models` lists the single `copilot` model.

So a client cannot select a model, and nothing in the request shape suggests a model parameter is honoured. What you get is whatever the consumer chat is serving, addressed under one id. That is the honest description of an interface that sits in front of a consumer UI rather than a model catalogue.

It also means the compatibility is at the wire level, not the catalogue level. Code that lists models, picks one by capability or checks pricing will not find what it expects, while code that sends chat completions and reads text back will work unchanged.

Multi-turn state is handled explicitly rather than implicitly: a conversation id comes back on the first reply and is accepted on the next request, so threads survive across calls without server-side session state.

## Four dependencies, one of which impersonates a browser's TLS fingerprint

The dependency list is four lines: `curl_cffi`, `playwright`, `fastapi` and `uvicorn`, all with lower bounds only.

`curl_cffi` is the interesting one. It is an HTTP client built around curl's TLS stack with browser impersonation, which is what makes a scripted request look like a browser's handshake rather than a Python client's. For a project whose whole job is talking to a consumer web endpoint through a headless browser, that is the library doing the load-bearing work, and it is worth knowing it is there rather than discovering it in a network trace.

The rest follows from the same design. Playwright at 1.60 or later drives the browser for sign-in and the diagnostic, FastAPI and Uvicorn serve the local API.

The container keeps the two in step: the base image is the official Playwright Python image tagged to the same minor version, so the preinstalled Chromium matches the pin. The Dockerfile also installs dependencies before copying the source, which is a layer-caching decision rather than a correctness one, and sets `HOST` and `PORT` before handing control to `app.py`.

## A region-evasion claim sits in the feature list

One of the reasons the project gives is that the signed-in path works in regions where anonymous Copilot is blocked, with India named as the example. It is listed under why use this, as a capability.

Stated plainly, what that means is that signing in with an account extends the availability of a service that otherwise gates access by region. Whether that is within the terms that apply to your account is a separate question, and it is the same question the disclaimer raises.

The repository itself is small and conventional: a `copilot/` package, a `server/` package, an `app.py` entry point, tests, assets, and seven numbered examples that walk from direct chat through direct conversation and streaming to the server via raw HTTP, server-side streaming and the OpenAI SDK. The default branch is master, the licence is MIT, there are no GitHub releases, and the last push is dated 2026-06-27.

The age of that last commit is the practical caveat: a reverse-engineered interface against a consumer web service is the kind of thing that stops working when the service changes, and this one has had no commits in three months.

## Conclusion

First, read Microsoft's terms for the consumer Copilot service and decide for yourself whether automated access to that chat endpoint is something your account may do; the project's own disclaimer asks users to stay within them, and that judgement is yours to make. Second, bind the server to loopback. The compose file publishes port 8000 on every interface while ignoring the API key, which means anything that can reach your host can spend your signed-in session. Third, understand that session/ holds a real browser profile and token for your own account, git-ignored but still on disk. Fourth, expect re-authentication roughly every half hour in a container. It is a useful tool for one person on one machine, and not something to expose on a network.

## FAQ

### What does Windows-Copilot-API provide?

It reverse-engineers the consumer Copilot web chat and republishes it two ways: a Python client with chat and streaming methods, and a local OpenAI-compatible server on port 8000 exposing POST /v1/chat/completions, which supports streaming and an optional conversation_id, plus GET /v1/models. The licence is MIT and the project is unofficial.

### Do I need an API key for Windows-Copilot-API?

No key, no credits and no paid plan, because it uses your already signed-in Copilot session. If you point the OpenAI SDK at the local server, the SDK still requires an api_key value, and the server ignores it.

### How do I sign in to Windows-Copilot-API?

Run `python -m copilot login`, which opens a visible browser and closes itself once sign-in is detected. It then sends one short warm-up message that mints the chat token and clears the human-verification check in the same step, which leaves a small throwaway chat in your history. The session is saved under session/, which is git-ignored.

### Can I run Windows-Copilot-API in Docker?

Yes, but sign in on the host first. `docker compose up --build` maps port 8000 and bind-mounts session/, and the container refreshes the chat token headlessly. It cannot earn fresh human-verification clearance without a visible browser, so when clearance expires, quoted as roughly thirty minutes, requests return 503 until you log in again on the host.

### What models does Windows-Copilot-API expose?

The models endpoint lists a single model named copilot, so a client cannot select a model. The compatibility is at the wire level rather than the catalogue level, and multi-turn state is carried by passing the conversation_id returned with a reply back on the next request.

### Is Windows-Copilot-API official Microsoft software?

No. The page states that it is an unofficial project, not affiliated with or endorsed by Microsoft, that it automates the consumer Copilot web experience for personal use, and that it should be used responsibly and within Microsoft's terms. Those terms are what you should read before deciding whether to use it.

## Sources

- [Issues](https://github.com/sums001/Windows-Copilot-API/issues)
- [License: MIT](https://github.com/sums001/Windows-Copilot-API/blob/master/LICENSE)
- [README](https://github.com/sums001/Windows-Copilot-API/blob/master/README.md)
- [sums001/Windows-Copilot-API on GitHub](https://github.com/sums001/Windows-Copilot-API)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sums001-windows-copilot-api
