# jarvis-ai-assistant: ten Python voice modules and a Windows-shaped dependency list

> A voice-controlled Python desktop assistant organised as ten swappable subsystems rather than one script. The install is three commands, but the dependency list is Windows-only in places, the top level carries committed logs and a driver binary, and the licence line contradicts the licence field.

**AnubhavChaturvedi-GitHub/jarvis-ai-assistant** — Voice-controlled AI desktop assistant in Python. Speech recognition, text to speech, real-time web search, image generation, computer vision and WhatsApp automation, inspired by Iron Man's JARVIS.

- Repository: https://github.com/AnubhavChaturvedi-GitHub/jarvis-ai-assistant
- Website: https://www.youtube.com/@NetHyTech
- Stars: 351 · Forks: 70
- Language: Python
- License: GPL-3.0
- Published: 2026-09-18 · Updated: 2026-09-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/anubhavchaturvedi-github-jarvis-ai-assistant

## Two entry points, one microphone, and Chrome in the list

Installation is three commands and there is no packaging step:

```bash
git clone https://github.com/AnubhavChaturvedi-GitHub/jarvis-ai-assistant.git
cd jarvis-ai-assistant
pip install -r requirements.txt
```

Running it is one of two lines, and the choice is about presentation rather than capability:

```bash
python jarvis.py
```

```bash
python ui.py
```

jarvis.py is the entry point with intent routing, ui.py is the desktop interface, and the second is the one the page recommends over a terminal. The prerequisites are Python 3.10 or newer, a working microphone and speakers, and Google Chrome, which is needed by the browser automation modules rather than by the voice path. Chrome being a prerequisite for the whole project, and not only for the module that drives a browser, is the first thing to know before installing.

Interaction is a wake word followed by ordinary speech. The page never says what the wake word is, so you find out by running it.

## Ten modules that are meant to be used one at a time

The design decision stated up front is that this is built as separate, swappable subsystems rather than one giant script, so you use only the parts you need. The module table is where that shows. NetHyTechSTT is a custom speech to text engine, and the row claims no paid API is required. TextToSpeech gives natural spoken replies. Brain, backed by co_brain.py, holds language model reasoning and conversation memory. Real_Time is live web search, described as the reason answers are not limited to training data. TextToImage generates images from a spoken prompt. Vision does camera capture and image understanding. Automation opens apps, controls the desktop and runs system tasks. Whatsapp_automation sends messages hands free. Weather_Check gives live weather by location, and Time_Operations handles alarms, reminders and scheduling.

That list is also the troubleshooting map. Each directory in the project structure block holds one of those names, so a failure in speech output is a TextToSpeech problem rather than a jarvis.py problem, and swapping a subsystem does not mean rewriting the router.

The sample utterances show the routing in action: a weather question for Bangalore, opening Chrome to search for transformer architecture, generating an image of a red sports car at sunset, sending a WhatsApp message to Rahul, and asking what is happening in the news right now.

## The structure block leaves four directories out

The documented layout lists the entry points and the nine subsystem directories:

```
jarvis.py              entry point, intent routing
ui.py                  desktop interface
co_brain.py            reasoning and conversation memory
NetHyTechSTT/          speech to text engine
TextToSpeech/          voice output
TextToImage/           image generation
Real_Time/             live web search
Vision/                camera and image understanding
Automation/            desktop and app control
Whatsapp_automation/   messaging
Weather_Check/         weather lookups
Time_Operations/       alarms and scheduling
```

The actual top level is longer. Alongside the entry points and the directories above it holds Brain/, Data/, Features/, New/, Alert.py, demo.py, internet_check.py, setup.py and requirements.txt. Four things follow from that gap.

The module table names Brain while the structure block omits the Brain directory and shows only co_brain.py, so the directory and the file are two halves of the same subsystem with no explanation of who calls which. Data/, Features/ and New/ appear nowhere in the documentation at all, and New/ is the kind of name that suggests copied-in material rather than a designed component. Alert.py and internet_check.py are likewise undocumented, which means the router has more entry points than a newcomer can see.

## setup.py repeats the install step instead of replacing it

There is a setup.py in the repository and it does not set up anything. The entire file is an import of subprocess and one call:

```python
import subprocess

subprocess.run("pip install -r requirements.txt")
```

So running it installs the dependency list and nothing else. There is no setuptools metadata, no entry_points, no package declaration and no console script, which means the project has no declared install target even though the page calls it one Python project. Installation goes through pip install -r requirements.txt, exactly as the three-line walkthrough says, and this file is a second, undocumented route to the same result.

It is worth knowing about before you package it yourself. Anything that runs setup.py to build a distribution will get an empty package, and the requirements it pulls in are only the ones already written in requirements.txt.

## The dependency list is Windows shaped and pins one old release

requirements.txt names twenty packages with no version constraints on all but one: requests, winotify, pyautogui, pywhatkit, playsound==1.2.2, psutil, selenium, webdriver_manager, webscout, lxml_html_clean, gradio_client, colorlog, yaspin, opencv-python, pyaudio, scipy, wmi, comtypes and pycaw.

Three observations come out of that list. wmi, comtypes and pycaw are Windows interfaces, so a clean install on Linux or macOS has three packages that cannot resolve to anything useful, and pip has no environment marker in the file to skip them. playsound is pinned to 1.2.2 while every other package floats, which is the kind of pin that outlives the reason for it. And a binary driver, chromedriver.exe, is committed at the top level of the repository, next to webdriver_manager, which is the package whose whole purpose is fetching a driver for you.

There is also a mismatch with the stated tech stack of Python, SpeechRecognition, Selenium, PyWhatKit, OpenCV, Requests and Tkinter. SpeechRecognition is named there and absent from requirements.txt, while gradio_client, webscout and colorlog are required and unnamed.

## Logs, chat history and a captured frame are committed

The top-level listing is the clearest sign of how this repository is worked on. Alongside the code it holds Alam_data.txt, PyWhatKit_DB.txt, chat_hystory.txt, conv_history.txt, history.txt, input.txt, log.txt, schedule.txt and speech_output.txt, then captured_image.png, generated_image.png, a .gif with a hash in its name, and a __pycache__/ directory.

Those files are runtime state. input.txt and speech_output.txt are what the last command was and what was said back; the history files are conversation; log.txt is a run log; PyWhatKit_DB.txt and Alam_data.txt belong to the messaging and automation libraries. A captured camera frame and a generated image are outputs of the Vision and TextToImage modules. Committing them means the repository carries whatever was on the machine of whoever last used it, including any personal content those paths happened to hold.

The filename chat_hystory.txt carries a typo, which suggests these were created by hand and never cleaned up. For a project that reads a microphone and opens a desktop, that is the detail to check before cloning it into anything of your own.

## The licence line and the licence field disagree

There are two licences in play and they are not the same one. The License section of the readme says the project is released under the MIT License, with a badge linking to the LICENSE file near the top of the page. The licence recorded for the repository is GPL-3.0.

That is the kind of conflict worth stopping on rather than working around. Under MIT you can take the code, change it and use it inside a closed product. Under GPL-3.0 a distributed derivative carries source obligations. For anyone copying a module out of this project, the difference decides whether they need to publish what they built, so it is not a formatting question.

Everything else about the project is stated plainly: the default branch is main, the primary language is Python, and the author is Anubhav Chaturvedi, founder of NetHyTech, a developer community of 30,000 plus members, with the recorded homepage being the NetHyTech YouTube channel.

## Offline-friendly sits next to a module that needs the network

The opening description calls this an offline-friendly, voice-activated assistant. Two of the ten modules are plainly not offline: Real_Time performs live web search, and TextToImage generates images from a spoken prompt through gradio_client. Weather_Check is described as live weather by location. So the claim holds for the parts that do not need a service, and nothing on the page says which parts those are.

That matters when you are deciding whether to trust the word offline-friendly with your own data. Speech recognition and spoken replies, the timers and the desktop automation are the parts that plausibly run without a network, and the sample utterances are mostly local actions anyway: open Chrome, set an alarm, send a message.

The contributions path is conventional in the same understated way. Fork the repository, create a feature branch, and open a pull request describing what changed and why. There are no GitHub releases, so nothing to install by version and nothing to compare against, and the last push is dated 12 August 2026.

## Conclusion

This is a project to read and extend rather than to deploy. The subsystem split is a sound way to learn which piece of a voice assistant is hard, and the entry points are simple enough to run on a machine with a microphone and Chrome. What you should settle before using it for anything real is the licence conflict between the stated MIT release and the GPL-3.0 field, the Windows-only packages in the dependency list on a machine that is not running Windows, and the fact that chat and command history files sit in the repository rather than in a directory you control.

## FAQ

### How do I install and run jarvis-ai-assistant?

Clone the repository, cd into jarvis-ai-assistant and run pip install -r requirements.txt. Then start it with python jarvis.py, or python ui.py for the desktop interface. Python 3.10 or newer, a microphone, speakers and Google Chrome are listed as prerequisites.

### Which speech recognition does jarvis-ai-assistant use?

Its own NetHyTechSTT module, described as a custom speech to text engine that needs no paid API. The stated tech stack also names SpeechRecognition, though that package does not appear in requirements.txt.

### Which modules does jarvis-ai-assistant need?

Only the ones you want to use, since it is built as separate swappable subsystems. wmi, comtypes and pycaw in requirements.txt are Windows interfaces, and chromedriver.exe is committed at the top level.

### What licence is jarvis-ai-assistant under?

The readme states MIT in its License section, while the repository's licence field records GPL-3.0. Which one governs a copied module depends on which of the two you believe.

## Sources

- [AnubhavChaturvedi-GitHub/jarvis-ai-assistant on GitHub](https://github.com/AnubhavChaturvedi-GitHub/jarvis-ai-assistant)
- [Issues](https://github.com/AnubhavChaturvedi-GitHub/jarvis-ai-assistant/issues)
- [License: GPL-3.0](https://github.com/AnubhavChaturvedi-GitHub/jarvis-ai-assistant/blob/main/LICENSE)
- [Project website](https://www.youtube.com/@NetHyTech)
- [README](https://github.com/AnubhavChaturvedi-GitHub/jarvis-ai-assistant/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/anubhavchaturvedi-github-jarvis-ai-assistant
