# Ecoute: real-time dual-channel transcription for Windows

> Ecoute transcribes your microphone and your speakers at the same time in a Tkinter textbox, running either a local Whisper tiny.en model or the OpenAI Whisper API. It is a Windows-only Python project, last pushed on 2026-04-08.

**SevaSk/ecoute** — Ecoute is a live transcription tool that provides real-time transcripts for both the user's microphone input (You) and the user's speakers output (Speaker) in a textbox.

- Repository: https://github.com/SevaSk/ecoute
- Website: https://github.com/SevaSk/ecoute
- Stars: 6,049 · Forks: 828
- Language: Python
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/sevask-ecoute

## What Ecoute transcribes, and for whom

Ecoute captures two audio channels at once and writes both into a single textbox. One channel is your microphone, labelled You. The other is your speaker output, labelled Speaker. That second channel is the part most transcription tools skip: it is the audio your machine is playing, so it catches the other side of a call, a video, or anything else routed through the sound card.

The README frames the purpose narrowly: "Ecoute is designed to help users in their conversations by providing live transcriptions." There is no transcript export, no speaker diarization, no meeting integration. The repository topics list gpt-35-turbo, whisper-ai and windows, which is a fair summary of the scope. If you want a running on-screen record of a conversation you are having at your desk, this is the target case. If you want a searchable archive of recorded meetings with named participants, it is not.

## How the audio pipeline is split across four modules

The repository layout shows four Python modules at the top level, and their names describe the data flow: AudioRecorder.py captures sound, AudioTranscriber.py turns it into text, TranscriberModels.py selects the model backend, and main.py wires them together and drives the customtkinter interface. custom_speech_recognition/ is a vendored copy of a speech recognition library rather than an external dependency, and tiny.en.pt is the local Whisper model file shipped with the repository.

The two backends in TranscriberModels.py behave differently. Without the --api flag, Ecoute runs the local tiny.en Whisper model, which the README says was chosen "due to its low resource consumption and fast response times." With the flag, transcription is sent to the OpenAI Whisper API instead. requirements.txt reflects both paths: faster-whisper and ctranslate2 for local inference, openai for the API, and torch pinned to a CUDA 12.1 wheel index. The README notes the system "might take a few seconds for the system to warm up before the transcription becomes real-time," which is consistent with loading a local model before the first inference.

The capture side is the constraint worth understanding. The README states Ecoute "is currently configured to listen only to the default microphone and speaker set in your system." There is no device picker in the interface. Changing input means changing your Windows default device before launching.

## Installing Ecoute on Windows with FFmpeg and PyAudioWPatch

The prerequisites are Python 3.8.0 or newer, FFmpeg, and a Windows machine. The README says the project has not been tested on other operating systems, so treat the Windows requirement as hard rather than incidental.

FFmpeg is the dependency most likely to trip you up. The README's route is Chocolatey, run from an administrator PowerShell:

```bash
Set-ExecutionPolicy Bypass -Scope Process -Force; [System.Net.ServicePointManager]::SecurityProtocol = [System.Net.ServicePointManager]::SecurityProtocol -bor 3072; iex ((New-Object System.Net.WebClient).DownloadString('https://community.chocolatey.org/install.ps1'))
```

Once Chocolatey is present, FFmpeg installs with a single command:

```bash
choco install ffmpeg
```

The README explicitly says to run these in a PowerShell window with administrator privileges. Then clone the repository and install the Python dependencies:

```bash
git clone https://github.com/SevaSk/ecoute
cd ecoute
pip install -r requirements.txt
```

The requirements file pulls PyAudioWPatch, which is the Windows-specific PortAudio binding that makes speaker-loopback capture possible. On Linux this package is not the usual choice, which is one more reason the Windows constraint is structural. If you want the API backend, write the key file. The README offers a one-liner that creates keys.py in the ecoute directory:

```bash
python -c "with open('keys.py', 'w', encoding='utf-8') as f: f.write('OPENAI_API_KEY=\"API KEY\"')"
```

Replace API KEY with a real OpenAI key that can access the Whisper API. The alternative is creating keys.py by hand containing OPENAI_API_KEY="API KEY". Then start it:

```bash
python main.py
```

Expect a few seconds of warm-up before transcription keeps up with live speech. The README recommends the API variant for speed and language coverage:

```bash
python main.py --api
```

With the flag you should see the same two-panel textbox, but transcripts appear faster and non-English speech is handled. Without it, you are on the local tiny.en model.

## tiny.en, English-only, and the cost of the API flag

The default local model is the weakest link. The README states that without --api, Ecoute uses the 'tiny' version of Whisper, and that it "may not be as accurate as the larger models in transcribing certain types of speech, including accents or uncommon words." The shipped file is tiny.en.pt, and the README confirms the model "is set to English," so non-English speech and dialects will transcribe poorly or not at all. The README says multi-language support is being worked on for future versions; there is no release in the repository that delivers it.

The API flag fixes accuracy and language coverage but changes the economics. The README is direct: "using the Whisper API will consume more OpenAI credits than using the local model." Every second of captured microphone and speaker audio is uploaded, and both channels are transcribed, so a one-hour conversation is two hours of audio against your account. The README calls the trade-off potentially worthwhile, but the decision is yours and it is per-use, not one-time.

The device limitation is the third failure mode. Because Ecoute binds to the system default microphone and speaker, a headset that is not set as default, a virtual audio cable, or a second monitor with its own speakers will simply not be captured. The README's instruction is to set the device you want as default in system settings before running.

## How Ecoute differs from OBS plus a transcription plugin

The nearest alternative for capturing both sides of a conversation on Windows is OBS Studio with a speech-to-text plugin. OBS gives you explicit source selection: you add a microphone source and a desktop audio source and route each one where you want it, so the default-device restriction that Ecoute documents does not apply. The trade-off runs the other way. OBS is a streaming and recording application first, and getting a live two-speaker transcript out of it means configuring a plugin, a scene, and often a separate transcription service. Ecoute ships that wiring already assembled in four Python files, at the cost of the flexibility OBS gives you.

If your goal is a permanent, diarized record of a meeting platform call rather than an on-screen live feed, neither tool is the right shape. The README's own sponsorship note points at Recall.ai, which it describes as pulling speaker data and separate audio streams from Zoom, Google Meet and Microsoft Teams to produce speaker-attributed transcripts. That is a hosted API with a different architecture: it joins the meeting rather than listening to your sound card. Ecoute's advantage is that it never talks to the meeting platform at all, so it works on any audio your machine plays.

## Maintenance status, licence and upgrade cost

The repository is not archived, and the last push was on 2026-04-08. There are no releases, so the only way to get the current code is to clone main. That means upgrades are a git pull plus a re-run of pip install -r requirements.txt, and you should expect the dependency set to move: torch is pulled from a CUDA 12.1 wheel index, and ctranslate2 is pinned to 3.24.0, so a Python or CUDA environment change can break the local inference path even when the Ecoute source is unchanged.

The licence is MIT, which permits commercial use, modification and redistribution provided the copyright notice and permission notice are retained. That is a permissive licence, and the practical implication is that you can vendor the code into an internal tool without a licensing conversation. It does not affect the OpenAI terms that apply once you set OPENAI_API_KEY and send audio to the Whisper API. Those are separate terms between you and OpenAI, and the README does not discuss them. This is not legal advice; read the OpenAI API terms yourself if you plan to transcribe other people's speech.

## Conclusion

Adopt Ecoute if you work on Windows, want both sides of a conversation transcribed in one window, and are willing to accept the tiny.en model's accuracy or pay for the Whisper API. Do not adopt it if you need multi-language transcription without the API, if your audio does not run through the default microphone and speaker, or if you are not on Windows, since the README states it has not been tested elsewhere. Before relying on it, verify that FFmpeg is on PATH, that your target devices are set as system defaults, and whether you are willing to spend OpenAI credits on every --api run.

## FAQ

### Does Ecoute transcribe both my microphone and my speakers?

Yes. The README describes it as providing real-time transcripts for both the user's microphone input, labelled You, and the user's speakers output, labelled Speaker, in a textbox. It listens only to the default microphone and speaker configured in your system.

### What does the --api flag change in Ecoute?

It switches transcription from the local tiny Whisper model to the OpenAI Whisper API, which the README says significantly enhances speed and accuracy and works in most languages rather than just English. It also consumes more OpenAI credits than the local model.

### Can Ecoute transcribe languages other than English?

Only with the --api flag. Without it, the README states the Whisper model is set to English and may not accurately transcribe non-English languages or dialects.

### Does Ecoute work on macOS or Linux?

The README lists Windows as a prerequisite and notes it has not been tested on other operating systems. The dependency list includes PyAudioWPatch, a Windows-specific audio binding.

## Sources

- [Issues](https://github.com/SevaSk/ecoute/issues)
- [License: MIT](https://github.com/SevaSk/ecoute/blob/main/LICENSE)
- [Project website](https://github.com/SevaSk/ecoute)
- [README](https://github.com/SevaSk/ecoute/blob/main/README.md)
- [SevaSk/ecoute on GitHub](https://github.com/SevaSk/ecoute)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/sevask-ecoute
