# stepanogil/autonomous-hr-chatbot: a 2026 rewrite in a subfolder behind a 2023 README

> A prototype HR agent that answers policy questions from a Pinecone index and hands a pandas dataframe of employee records to LangChain tools, fronted by Streamlit. The last commit was on 2026-04-29, the documented run command names a file that is not in the tree, and two pinned dependencies are pre-release versions.

**stepanogil/autonomous-hr-chatbot** — An autonomous HR agent that can answer user queries using tools

- Repository: https://github.com/stepanogil/autonomous-hr-chatbot
- Website: https://autonomous-hr-chatbot.vercel.app
- Stars: 460 · Forks: 110
- Language: Python
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/stepanogil-autonomous-hr-chatbot

## The README documents a stack the repository has moved past

The first line of the file is a notice that the rest of the file is history. It says a modernized version using an agent loop through OpenAI's Responses API, with no LangChain and no Pinecone, and with reasoning summaries from gpt-5.2, lives in the `v2/` directory, and that the original 2023 implementation below is preserved as a reference. Everything under that banner is the old path: LangChain agents and tools, Pinecone as the vector database, ChatGPT or gpt-3.5-turbo as the model, and a Streamlit front end using the streamlit_chat component. So the title, the feature description, the setup steps and the tech stack table all describe code the author has already replaced. The v2 directory exists in the top-level listing, confirming the modernised path is in this repository rather than a branch, but nothing beyond the banner describes how to run it.

## The documented run command names a file that is not in the tree

The setup steps tell you to run this in your terminal:

```bash
streamlit run hr_agent_frontent.py
```

There is no such file. The repository contains `hr_agent_frontend.py`, with the letters in the correct order, so the command as printed fails on a fresh clone. A second slip sits one step earlier, where the instructions say to uncomment the Azure backend in the frontend file, referring to it as `frontend.py`, which matches neither the printed command nor the actual filename. These are small typos rather than design errors, but they sit in the two steps a newcomer runs first, and the corrected name is only discoverable by listing the directory. The rest of the setup is more forgiving, since the install step is a conventional requirements file read:

```bash
pip install -r requirements.txt
```

## The setup asks you to paste API keys into a source file

Step four of the setup is where the security posture is decided, and it is decided in a Python file. You are told to enter your own API keys directly in `hr_agent_backend_local.py`, or in `hr_agent_backend_azure.py` if you want the Azure variant, which you then enable by uncommenting it in the front-end file. There is no environment variable, no dotenv file, no secrets manager and no prompt. The same pattern appears in the embedding notebook, where the instruction is to replace the Pinecone and OpenAI API keys with your own inside the notebook itself. The presence of `.gitignore` in the tree does not help here, because the credential-bearing files are tracked source rather than local configuration. Anyone adapting this pattern should take the mechanism and replace the storage location before adding a real key.

## Two pinned pre-release versions block the documented install

The requirements file pins every dependency exactly, more than a hundred of them, and two of those pins carry pre-release suffixes. `azure-storage-blob==12.17.0b1` ends in a beta marker and `pydeck==0.8.1b0` ends in a beta marker. An exact requirement against a pre-release version is not satisfied by a normal resolution, since pip excludes pre-releases unless it is told otherwise, so the documented `pip install -r requirements.txt` is the step most exposed to failure or to resolving differently from what was tested. The Azure storage pin is not incidental: Data Lake is named in the tech stack for landing the employee CSV files, and blob storage is listed as the alternative, so this particular package is on the path the architecture depends on. Both betas also date the dependency set to a moment when Azure SDK packages were still shipping betas as the normal line.

## The dependency set still carries an abandoned telemetry SDK

Pinning everything makes the vintage of the stack unusually legible. The set is uniformly 2023: `langchain==0.0.220`, `openai==0.27.8`, `pinecone-client==2.2.2`, `streamlit==1.24.0`, `pydantic==1.10.10`, `pandas==2.0.3` and `numpy==1.25.0`, with the documented interpreter being Python 3.10. One entry is more pointed than old: `langchainplus-sdk==0.0.19`, the tracing client that was renamed to LangSmith shortly after this generation, and which has not existed under its original name since. It remains in the requirements because LangChain 0.0.220 era code imports it for callback tracing, so removing it breaks the import even though the service it called has a different name. That is the clearest single illustration of what freezing a dependency set does: it preserves the compatibility surface of the moment, including the parts that were about to be renamed.

## Employee-shaped data is committed at the top level

The repository root holds `employee_data.csv` alongside the source, and the description of it is a CSV of dummy employee data including name, supervisor and number of leaves, loaded as a pandas dataframe. Also at the top level is `hr_policy.txt`, the policy document whose embeddings feed the retrieval tool, which is described as a sample HR policy generated by ChatGPT. That combination is what makes the demo work with no infrastructure, and it is also what makes the demo risky to copy: the architecture names Azure Data Lake as the landing zone for employee data files, with blob or S3 as acceptable alternatives, and the CSV that stands in for that lake is checked into version control instead. Two notebooks sit at the root as well, one for embedding storage and one holding the agent code, which means a meaningful share of the project's logic lives in notebooks rather than in the modules.

## One of the three tools is a Python interpreter over employee records

The agent exposes three tools and they differ sharply in what they are allowed to do. The first is retrieval over a Pinecone index built from the policy document using OpenAI's text-embedding-ada-002 model. The third is arithmetic, described as LangChain's calculator chain module, LLMMathChain. The second is the one to look at: employee data is handed to the model as a pandas dataframe and manipulated using LangChain's PythonAstREPLTool, which executes model-generated Python against that data. So the agent does not query a database for leave counts or supervisor names, it writes code that operates on a loaded frame. That is the standard demonstration of tool use in that generation of LangChain, and it is also the component whose safety story has to be supplied by whatever wraps the agent.

## Embedding population is a notebook you edit by hand

Setting up the vector index is a manual, notebook-driven step rather than a command. The instructions are to create a Pinecone account, which has a free tier, and note the Pinecone API and environment values, then run `store_embeddings_in_pinecone.ipynb` and replace the Pinecone and OpenAI API keys for the embedding model with your own. The existence of a separate environment value implies more than one Pinecone deployment topology, but the guide does not say which one the free tier gives you or which the code expects. Nothing in the documented flow removes the index or updates it when the policy document changes, so re-running the notebook is what keeps the stored embeddings in step with the source text. The author is Stephen Bonifacio, and a video walkthrough is linked from the repository.

## Conclusion

This fits a reader studying how an early LangChain agents-and-tools pattern was put together, since the original implementation is deliberately preserved intact. Start from the v2 subdirectory rather than the root if you want current code, fix the front-end filename before running the documented command, and treat the committed employee CSV and the embedded keys workflow as things to replace rather than adopt.

## FAQ

### What is the difference between the root code and the v2 directory in autonomous-hr-chatbot?

The root holds the original 2023 implementation built on LangChain agents and tools, Pinecone and Streamlit, kept as a reference. The v2 directory holds a modernised version using an agent loop through OpenAI's Responses API, with no LangChain, no Pinecone, and reasoning summaries from gpt-5.2. Nothing beyond the banner describes how to run v2.

### How do I run the autonomous-hr-chatbot Streamlit app?

Install dependencies with `pip install -r requirements.txt` and then start Streamlit. The README prints `streamlit run hr_agent_frontent.py`, but the file in the repository is spelled `hr_agent_frontend.py`, so the printed command fails on a fresh clone.

### Where do autonomous-hr-chatbot API keys go?

Into the source files. The setup tells you to enter your own keys in `hr_agent_backend_local.py`, or in `hr_agent_backend_azure.py` for the Azure variant, and the embedding notebook instructs you to replace the Pinecone and OpenAI keys inside the notebook itself. No environment variable or dotenv file is used.

### What tools does the autonomous HR agent expose?

Three: a timekeeping policy search backed by a Pinecone index built with OpenAI's text-embedding-ada-002 embeddings, employee data manipulation using LangChain's PythonAstREPLTool over a pandas dataframe, and a calculator using LangChain's LLMMathChain module.

### Which versions does the autonomous-hr-chatbot requirements file pin?

Every dependency is pinned exactly, uniformly at 2023 versions including `langchain==0.0.220`, `openai==0.27.8`, `pinecone-client==2.2.2` and `streamlit==1.24.0`, alongside `langchainplus-sdk==0.0.19`, the tracing SDK later renamed to LangSmith. Two pins are pre-release versions, `azure-storage-blob==12.17.0b1` and `pydeck==0.8.1b0`, which a default pip resolution will not satisfy.

## Sources

- [Issues](https://github.com/stepanogil/autonomous-hr-chatbot/issues)
- [License: MIT](https://github.com/stepanogil/autonomous-hr-chatbot/blob/main/LICENSE)
- [Project website](https://autonomous-hr-chatbot.vercel.app)
- [README](https://github.com/stepanogil/autonomous-hr-chatbot/blob/main/README.md)
- [stepanogil/autonomous-hr-chatbot on GitHub](https://github.com/stepanogil/autonomous-hr-chatbot)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/stepanogil-autonomous-hr-chatbot
