Library / SDK
winkjs/wink-nlp avatar
winkjs/wink-nlp

wink-nlp needs a second install for its model, publishes no runtime dependencies, and calls a doc script sourcedocs that runs an npm package named docker

Developer friendly Natural Language Processing ✨

1,393 stars65 forksJavaScriptMIT

At a glance

What is it?
A dependency free JavaScript NLP library for Node, browsers, and Deno with a pipeline from tokenization to named entity recognition. Installing it is two npm commands because the language model is a separate package chosen by Node version, the version table on the page stops mid row, and the head to head speed figures all come from one laptop with no published methodology.
Who is it for?
This fits a browser or edge project where bundle size and dependency count are the binding constraints, and where you want POS, NER, and sentiment without shipping a model runtime. Three things to check first.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 130 days ago.
What is it written in?
Mainly JavaScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A model package is a second install, and the version table stops mid row

The headline install is one command:

shell
npm install wink-nlp --save

That is not enough on its own. The page says you also need to install a language model according to the Node version you are on, and gives a table mapping version to package. The first row shown is Node.js 16 or 18 paired with `npm install wink-eng-lite-web-model --save`. The next row begins with a 1 and then stops, so the mapping for every newer Node release is not on the page. That is the single most consequential gap in the documentation, because picking the wrong model package for your runtime is exactly the kind of mistake that produces a confusing failure rather than a clear one. The library's own models directory exists at the repository root, and the sibling packages are separate npm installs, so this is a two package setup rather than an oversight in one file. Read the table on the documentation site rather than guessing, and check that the model you pick is the one the current release expects.

No runtime dependencies, and the models are siblings rather than children

The claim that there is no external dependency is verifiable in the manifest, because there is no `dependencies` key at all. Every entry sits under `devDependencies`: benchmark, chai, coveralls, a package literally named docker, dtslint, eslint, mocha, and nyc. For a library that wants to be dropped into a browser bundle, that is the right shape and it is the reason the minified and gzipped size is quoted as roughly 10Kb. The consequence is that everything the library needs at runtime is either inside it or inside a sibling package you install yourself. Word vectors come from a separate embeddings package covering 100 dimensions and over 350K English words, and the pre-trained language models are separate packages starting at about 1MB minified and gzipped, quoted as cutting model load time to roughly a second on a 4G network. So the dependency count of the core is zero and the download size of a working setup is not, and those are two different numbers that marketing copy tends to blur.

The sourcedocs script runs an npm package named docker

One line in the scripts block is worth reading twice:

code
"sourcedocs": "docker -i src -o ./sourcedocs --sidebar yes"

That is not the container runtime. The `docker` package at version ^1.0.0 is an npm module in devDependencies, and this command is its command line interface, taking `-i` for input, `-o` for output, and a sidebar flag, generating API documentation from the `src` directory. A developer who skims the scripts block, sees the word docker, and reaches for their terminal will find either nothing or a completely different tool. Nothing in the manifest comments the script, so the only way to tell is to look the package up. The naming is a small thing that will cost somebody an afternoon once, and it is a decent argument for reading manifests before assuming a script means what it looks like.

Three speed figures, one laptop, and no published methodology

Three performance claims appear and all three name the same machine, an M1 MacBook Pro, with nothing else specified. The headline is over 650,000 tokens per second for winkNLP in both browser and Node environments. The feature table claims the tokenizer alone runs close to 4 million tokens per second in the M1's browser, roughly six times the headline figure, which is consistent with a tokenizer being much cheaper than a full pipeline but is never explained. The third claim, that it runs smoothly on a low end smartphone's browser, carries no number at all. The tokenizer example given for the multilingual claim, tokenizing `¡Hola! नमस्कार! Hi! Bonjour chéri` into nine tokens with punctuation split out as its own token, is a correctness demonstration rather than a throughput one. The honest half of this is that the benchmark is runnable: `npm run perf` invokes `node benchmark/run.js`, and a browser measurement procedure is published as a notebook. Run it on your own text before quoting any of these numbers.

The coverage script pipes to an external service, and pretest always lints

Two scripts have side effects worth knowing about. The first is coverage:

code
"coverage": "nyc report --reporter=text-lcov | coveralls"

That pipes a coverage report into the coveralls client, which uploads it to an external service. Running `npm run coverage` is therefore a network write, not a local report, and it is not obvious from the name. The second is that `pretest` is set to `npm run lint`, so `npm test` never runs without eslint first, and eslint covers only `./src/*.js` and `./test/*.js` by explicit glob rather than by configuration. Coverage itself comes from nyc wrapping mocha against `./test/`, with an html, lcov, and text reporter, and the nyc settings sit in a `.nycrc.json`. The configuration files also tell a story: `.eslintrc.json` is the pre flat config format, consistent with the eslint 8 in devDependencies, and there is both a `.travis.yml` and a `.github/` directory, so the visible CI history predates the current one.

A multilingual tokenizer paired with English only vectors

The two headline capabilities cover different languages, and the page does not reconcile them. The tokenizer is presented as lossless and multilingual, with the mixed Spanish, Hindi, English, and French string as its demonstration, splitting punctuation into separate tokens so nothing is discarded. The word vectors are 100 dimensional and cover over 350K English words, and the language model packages are separate downloads rather than one multilingual bundle. So a pipeline that classifies English text with embeddings and tokenizes Hindi in the same pass is a supported shape, while a non English pipeline gets the tokenizer and has to find its own model. The utilities follow the same split in practice. BM25, cosine, Tversky, Sørensen-Dice, and Otsuka-Ochiai similarity all operate on whatever the tokenizer produced, and the its and as helpers cover bag of words, frequency tables, lemmas, stems, stop word removal, and negation handling. The pipeline itself runs tokenization, sentence boundary detection, negation handling, sentiment, POS, NER, and custom entity recognition in that order.

Version 2.4.0 added a method called map, and the tags carry no v prefix

The release names are free text rather than conventional summaries: 2.4.0 is `Added map()`, 2.3.2 is `Operational update`, and 2.3.1 is `Fixed some type definitions`. None of the tags carries a v prefix, which is worth noting if you script against the tag list. The cadence has also been uneven. 2.3.1 and 2.3.2 landed six days apart in November 2024, then nothing until 2.4.0 in June 2025, and the branch was last pushed on 2026-05-27 with no release since. The manifest version and the newest tag agree at 2.4.0, so there is no skew between what npm serves and what is tagged. Two design notes sit inside that history. Adding a method named `map` to a fluent chain is a deliberate collision with a name every JavaScript programmer already uses on arrays, so it is worth checking how it reads at a call site. And a release cut specifically to fix type definitions means the `.d.ts` surface is maintained by hand and tested with `dtslint`, which is good practice and also a sign to check the types against your own code after upgrading.

Editorial conclusion

This fits a browser or edge project where bundle size and dependency count are the binding constraints, and where you want POS, NER, and sentiment without shipping a model runtime. Three things to check first. Budget for two installs and pick the model package that matches your Node version, since the table that maps versions to models is incomplete on the page. Treat the speed figures as marketing until you run the bundled benchmark yourself, because a `npm run perf` target exists and takes seconds to invoke. And if you rely on the type definitions, note that the previous release was cut specifically to fix them, so pin a version whose dtslint output you have seen pass.

Frequently asked questions

How do I install wink-nlp?

With npm install wink-nlp --save, and then a second command for the language model, chosen by Node version. The table shown maps Node.js 16 or 18 to npm install wink-eng-lite-web-model --save, and the row for the next version is cut off partway through.

Does wink-nlp have any runtime dependencies?

None. There is no dependencies key in the manifest at all, and every entry sits in devDependencies. Language models and word vectors are separate packages you install alongside it, and the quoted core size is roughly 10Kb minified and gzipped.

What does wink-nlp do in a pipeline?

It covers tokenization, sentence boundary detection, negation handling, sentiment analysis, part-of-speech tagging, named entity recognition, and custom entity recognition. Utilities include a BM25 vectorizer, cosine, Tversky, Sørensen-Dice, and Otsuka-Ochiai similarity, plus its and as helpers.

Which languages does wink-nlp support?

The tokenizer is presented as lossless and multilingual, demonstrated on a string mixing Spanish, Hindi, English, and French. The pre-trained word vectors are narrower, described as 100 dimensional English embeddings for over 350K English words, and language models are separate packages.

How fast is wink-nlp?

The page quotes over 650,000 tokens per second for the library and close to 4 million tokens per second for the tokenizer, both on an M1 MacBook Pro, and says it runs smoothly in a low end smartphone browser. No input size or corpus is given, so the bundled benchmark is the way to check: npm run perf runs node benchmark/run.js.

What is the newest version of wink-nlp?

2.4.0, tagged 2025-06-30 and titled Added map(). The two releases before it were 2.3.2 and 2.3.1 in November 2024, and the branch was last pushed on 2026-05-27. The tags carry no v prefix.

Official sources

  1. License: MIT
  2. Project website
  3. README
  4. Releases
  5. winkjs/wink-nlp on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/winkjs-wink-nlp.svg)](https://hysenlabs.com/projects/winkjs-wink-nlp)