chatbot_ner: rule and gazetteer entity extraction for Indian-language chatbots
chatbot_ner: Named Entity Recognition for chatbots.
At a glance
- What is it?
- Haptik's chatbot_ner is a Django and Elasticsearch service that pulls dates, times, numbers, phone numbers and dictionary terms out of English, Hindi, Gujarati, Marathi, Bengali and Tamil messages. Its value is in code-mixed Indian text, and its cost is a heavy stack and a detector API that is only partly unified.
- Who is it for?
- Adopt chatbot_ner if your bot receives code-mixed Indian-language messages and you can run Django plus Elasticsearch 5.x in-house; the gazetteer search for text entities is the part that is hard to reproduce elsewhere. Do not adopt it if you need a pip-installable library, if you need transformer-grade accuracy on free-form text, or if you cannot operate a stateful search cluster.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Activity is slowing. The repository last received commits 6 months ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What chatbot_ner extracts that a general NER model does not
General-purpose NER models label people, organizations and locations. A booking or support bot needs different slots: a date, a time, a quantity with a unit, a phone number, a PNR code, a dish or a city name from a fixed catalogue. chatbot_ner is built around that second list. The README describes it as a framework "custom built to supports entity recognition in text messages", and the supported entity table is the clearest statement of scope: Time, Date, Number, Phone number, Email, Text, PNR and regex.
The language coverage is the reason this project exists rather than a wrapper around spaCy. Time, Date, Number, Phone number and Text are listed as supported for 'en', 'hi', 'gu', 'bn', 'mr' and 'ta', and the README adds "and their code mixed form". Code mixing is the normal case for Indian chat traffic: a user types "kal subah 5 baje" or "मुंबई में मौसम कैसा है" without switching keyboards deliberately. The date and time detectors carry CSV training data for exactly those patterns, and the contribution guide points at those CSV files as the intended way to widen coverage.
Who it is for: teams running a conversational product in one or more Indian languages, who want deterministic slot filling they can debug, and who are willing to host a service rather than import a library. Who it is not for: anyone whose text is mostly English free-form prose, where a transformer NER model will beat pattern matching on both accuracy and setup effort.
Detector taxonomy: numeral, pattern, temporal, textual
The repository organises detection by the shape of the problem, not by the entity name. The README splits entities into four types. Numeral covers anything about numbers, listing number, budget and size detection. Pattern covers identification by regular expression or fixed format, listing email, phone_number and pnr. Temporal covers time and date. Textual covers dictionary lookup, described as detecting "entities by looking at the dictionary", with cuisine, dish, restaurant names, cities and user location as examples.
That taxonomy maps onto the directory layout: ner_v2/detectors/numeral, ner_v2/detectors/pattern, ner_v2/detectors/temporal and ner_v1/detectors/textual. The version split is not cosmetic. According to the README, numeral, temporal and pattern detectors "have been moved to ner_v2 for language portability with more flexible detection logic", while in ner_v1 only the text entity has language support, and the project says it will move text to ner_v2 "without any major API changes". Until that happens, a caller touching both a date and a city is calling two different API generations.
The textual path is the one that needs a datastore. Custom entities are found by full text search in the datastore, which is why Elasticsearch is a hard dependency rather than an optional cache. The README also notes a contextual model for text that is supported for 'en' only, so the multilingual story for custom entities is search, not machine learning.
Installing chatbot_ner with Docker and running a first detection
The README does not inline install steps. It states that detailed documentation for setting up chatbot_ner on your system using Docker is available in docs/install.md, so that file is the place to start. What the repository does give you at the top level is the configuration surface: .env.example, config.example, docker/, dev_docker/, datastore_setup.py and manage.py.
The environment file is explicit about how to use it. Its own header says to copy it to docker/.env and fill in the values, and warns against adding spaces around '='. The datastore engine is selected with ENGINE, whose comment says valid values are one of ['elasticsearch']. Connection settings are prefixed ES_, and the comments note that ES_URL, when provided, overrides ES_HOST, ES_PORT, ES_AUTH_NAME and ES_AUTH_PASSWORD. The defaults assume the Elasticsearch instance that comes up with compose: ES_HOST=elasticsearch and ES_PORT=9200, with ES_ALIAS=entity_data and ES_INDEX_1=entity_data_v1.
cp .env.example docker/.env
# then edit docker/.env: set SECRET_KEY, and ES_* if you are not using the compose ESA minimal edited .env for a local run keeps the compose defaults and only replaces the secret:
NAME=chatbot_ner
DJANGODIR=/app
DJANGO_SETTINGS_MODULE=chatbot_ner.settings
SECRET_KEY=replace-this-value
ENGINE=elasticsearch
ES_HOST=elasticsearch
ES_PORT=9200
ES_ALIAS=entity_data
ES_INDEX_1=entity_data_v1
ES_DOC_TYPE=data_dictionary
PORT=8081After the stack is up, the API documentation in docs/api_call.md is the reference for request shapes per entity type. The README says the API structure is "built for ease of accessing it from conversational AI applications", and the repository ships postman_tests/ plus run_postman_tests.sh, which is the fastest way to see working calls without writing a client first. Expect to populate the datastore before text entities return anything: the textual detector searches the dictionary, so an empty index means empty results, not an error.
The dependency pinning is the real operational constraint
requirements.txt pins the whole stack, and several pins are old enough to shape your deployment. Django is pinned at 3.2.19, which is an LTS line. The Elasticsearch client is pinned at elasticsearch==5.5.3, which pairs with Elasticsearch 5.x servers; a modern Elasticsearch 8.x cluster will not accept that client. spaCy is pinned at 2.3.2, with model wheels for en_core_web_sm 2.3.1 and for Dutch, French, German and Spanish at 2.3.0, pulled directly from GitHub release URLs rather than from a package index. numpy is 1.19.2, pandas 1.0.5 and scipy 1.4.1.
The practical consequence is that you are not adding this to an existing Python 3.11 service with a modern dependency set. You are running a container built around an older interpreter and an older search server, and you either keep that isolation or you spend the effort to upgrade the pins yourself. The repository does not document a supported upgrade path for those pins, and the release history shows the cadence: 1.0.36 in October 2023, 1.0.37 in September 2024, 1.0.38 in October 2024. The last push to the repository was on 2026-04-01, so there is recent activity on the default develop branch even though the most recent tagged release predates it.
Second constraint: the datastore is stateful. Elasticsearch holds the gazetteer that makes text detection work, and .env.example includes index, alias, doc type, bulk size and search size settings. That means backups, index mapping changes and reindexing are part of your operating cost. A pure regex library would not have that cost, and it also would not find "pizza" in a user's message.
Custom detectors are a migration risk, not a stable surface
The README is unusually direct about this. It lists city, budget and shopping size as custom detectors derived from the primary detectors, then states they are supported currently in English only and limited to Indian users, that the team is restructuring them to scale across languages and geography, and that their current versions might be deprecated in future. It follows with a recommendation: "for applications already in production, we would recommend you to use only primary detectors mentioned in the table above".
Take that at face value. If your bot needs a city slot, the supported route is the Text entity with your own gazetteer in Elasticsearch, not the city detector. The same reasoning applies to budget and size: they are convenience wrappers whose API may move. Building on them means accepting a rewrite later.
The other limitation is the language boundary of the machine-learned path. Text detection by search covers 'en', 'hi', 'gu', 'bn', 'mr' and 'ta'. The contextual model covers 'en' only. So the moment you want fuzzy, context-dependent custom entity recognition in Hindi, you are outside what the README claims. You get dictionary matching, which is brittle against spelling variants and transliteration unless you add those variants to the index yourself. That is the honest trade-off: deterministic and debuggable versus recall on unseen surface forms.
chatbot_ner versus spaCy or a transformer NER pipeline
The obvious alternative is spaCy, which chatbot_ner itself depends on for English, or a transformer token-classification model. The difference is not accuracy on generic entities; it is what each one is designed to output.
A spaCy pipeline with an NER component gives you statistical labels over a token sequence. You can fine-tune it on your own annotations, and it handles unseen phrasing far better than a pattern list. What it does not give you out of the box is a Hindi date parser that understands "अगले सोमवार", a phone-number normaliser that accepts Devanagari digits such as ९८३३४३०५३५, or a dictionary search over a catalogue of dishes. chatbot_ner ships those as detectors with language-specific data, and the README frames the whole project as filling that gap for conversational AI.
The second difference is where the data lives. A spaCy model is a file you load into the process. chatbot_ner keeps its entity dictionary in Elasticsearch and queries it at request time, which means you can add a restaurant name to the index and have the bot recognise it immediately, with no retraining. That is a genuine operational advantage for catalogue-style entities, and it is also the reason the deployment is heavier. If your entities are stable and your text is English, spaCy alone is the smaller system. If your entities change weekly and your users write Hinglish, the gazetteer approach is doing work a model would need retraining to match.
Licence and contribution model
chatbot_ner is licensed under GPL-3.0. That is a copyleft licence, and it matters more here than for a standalone tool because the project is a service you would deploy alongside your own application code. Whether your use triggers distribution obligations depends on facts a licence file cannot settle, so treat this as a flag to raise with whoever handles licensing rather than a decision. The practical question to answer is whether you are linking the code into a distributed product or running it as an internal service, and that is a legal question, not a technical one.
On contribution, the README is specific about what it accepts today: you can contribute to ner_v2 either by adding training data or by contributing detection patterns in the form of regex. It notes the team will work on removing architectural limitations that currently make it hard to add ML models and new entities. Adding training data is deliberately low-friction; the date, time and number detector directories each hold CSV files, and the README says date detection in Hindi and Hinglish can be improved simply by adding rows there. If you extend those CSVs, you are working with the grain of the project rather than against it.
Editorial conclusion
Adopt chatbot_ner if your bot receives code-mixed Indian-language messages and you can run Django plus Elasticsearch 5.x in-house; the gazetteer search for text entities is the part that is hard to reproduce elsewhere. Do not adopt it if you need a pip-installable library, if you need transformer-grade accuracy on free-form text, or if you cannot operate a stateful search cluster. Before committing, verify that your Elasticsearch version matches the pinned 5.5.3 client, that the Docker install path in docs/install.md works on your host, and that the text entity's language coverage matches the languages you actually receive.
Frequently asked questions
What is NER in AI, and how does chatbot_ner fit that definition?
NER is named entity recognition: finding and classifying spans of text that refer to specific things. chatbot_ner applies it to chat messages, extracting Time, Date, Number, Phone number, Email, Text, PNR and regex-based entities rather than the person and organization labels a general model would produce.
What are the four types of chatbots, and does chatbot_ner suit all of them?
The README does not describe chatbot categories, so this cannot be answered from the repository. What it does say is that the API is designed for conversational AI applications and can be used for other applications too.
Which are the top 5 AI chatbots, and is chatbot_ner one of them?
chatbot_ner is not a chatbot. It is a named entity recognition framework that a chatbot calls to extract entities from user messages, so it does not belong in a list of end-user chatbots.
What exactly is a chatbot, and where does chatbot_ner sit in one?
The repository does not define chatbots. It positions itself as the entity-extraction layer inside a conversational AI application, sitting between the incoming message and whatever logic decides the next reply.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hellohaptik-chatbot-ner)