# extractor, a LangChain wrapper whose manifest version trails its newest tag

> Lightfeed Extractor hands HTML, markdown, or plain text to a LangChain chat model and asks for JSON matching a schema you supply with Zod. The interesting engineering is in the edges: recovering malformed output, repairing URLs, counting tokens, and stripping tracking parameters. The manifest is where the looseness shows, with a version behind the newest release, a deprecated linter, a publish hook that runs only the unit tests, and a conversion test whose expected output lives in a submodule you can regenerate.

**lightfeed/extractor** — Use LLMs to robustly extract web data

- Repository: https://github.com/lightfeed/extractor
- Website: https://lightfeed.ai
- Stars: 321 · Forks: 10
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/lightfeed-extractor

## The manifest declares 0.4.1 and the newest release is 0.5.4

The readme states no version anywhere. The manifest in the default branch states one, and it is not the newest. The published releases are three, the newest of them tagged five patch versions after the number in the manifest, and the default branch was pushed about two hours after that tag was cut.

So there are three different answers to the question of which version you are looking at: the number in the manifest, the number on the newest release, and whatever the registry serves when you run the install command the readme opens with. The readme's code samples belong to none of them by name, and nothing in it tells a reader to check.

The install line itself is the first thing in the file:

```bash
npm install @lightfeed/extractor @langchain/core
```

A library whose readme is a set of fragments rather than a versioned document makes this harder than it needs to be. The two flags in the manifest that would fix most of it, an exact version and a changelog link, are absent.

## The publish hook runs the unit tests and nothing else

Four of the manifest's script entries decide what has to pass before anything ships:

```json
"prepare": "npm run clean && npm run build",
"prepublishOnly": "npm run test:unit",
"test": "jest",
"lint": "tslint -p tsconfig.json"
```

The publish hook runs one of the two test directories and stops there. The integration tests do not run, and neither does the linter. The integration suite is not a small remainder either: it holds a dedicated test for the HTML to markdown conversion, plus separate scripts for running that one file, for coverage, and for watching. The conversion is the feature the readme leads with, and it is the part of the suite that publishing ignores.

The linter is a second signal about the project's age. The build is plain TypeScript compilation and the test runner is Jest with its own config file at the root, but style checking is done by a linter that has been deprecated in favour of the other one. Nobody has migrated it, and because the publish hook does not call it, nothing fails while it stays there.

The published file list is a single directory, and the root also carries an ignore file for npm, which is redundant next to a whitelist and quietly becomes wrong the first time someone edits only one of them.

## The conversion test compares against ground truth you can regenerate

The HTML to markdown conversion is checked against expected output that does not live in this repository. It lives in a second repository attached as a submodule, and the manifest carries three scripts for it. One initialises the submodule. One pulls its default branch:

```bash
git submodule update --init --recursive test-data
cd test-data && git pull origin main && cd ..
```

The third runs a script called for regenerating the ground truth. That is the problem. A test that compares a converter's output to a stored expectation is only as good as the review of the change to that expectation, and here the change can be produced by running the converter itself. Regenerate the expectations, and the suite passes without the converter changing by a character.

The second script is the mirror image. Pulling the default branch of the submodule means the expectations can change under a test run because of somebody else's commit, so a green suite on Monday and a red suite on Tuesday with no local edit is a state the manifest makes reachable. Pinning the submodule to a commit is the ordinary fix and the readme does not mention it.

## Token limiting and token tracking have no field in any shown call

The feature list advertises two things that are not settings. Extraction in JSON mode is described as including a token usage limit and token tracking, and the list calls the limit important for production data pipelines. Neither appears in an options object.

Every call the readme shows has the same shape: a model, some content, a format, a schema, and then whichever of three optional fields that example is about. Four examples, four options objects, and not one of them contains a number for tokens, a flag to turn tracking on, or a field to read the count out of.

That is not proof the options do not exist. The readme's option documentation is a set of examples rather than a reference table, and the file is visibly cut off partway down. What it does mean is that a reader trying to control cost per page, which is the use case the feature list names, has to go and read the source, because the one document written for readers does not show the knob.

The tracking claim has a second problem. Even if the count is recorded, the readme shows no place to read it, and the programmatic result in the examples is used only as an answer to an extraction call.

## A custom prompt replaces the default instead of adding to it

There is a default extraction prompt, and supplying the prompt field substitutes for it. The readme is explicit about the substitution and silent about the consequence.

The example that shows a custom prompt narrows the task: extract only products that are on sale or discounted, and include their original prices, discounted prices, and a product address. That is a filter, and a filter is exactly the kind of instruction a default prompt would need to be merged with rather than replaced by. Nothing says the schema instructions survive the swap, and nothing says what the model is told about fields the narrowed prompt stopped mentioning. An optional field in a Zod schema is optional in your types and may or may not survive a prompt that no longer asks for it.

The same example carries a second unexplained field. A source address is passed alongside the content, and the readme's separate section on contextual information tells you to put a website address inside that context object instead. So there are two ways to hand the model a URL, one of them appears exactly once, and the readme never says which one the library uses or what the other is for.

## Contextual information has no declared shape and no precedence rule

One option takes a free form object and asks you to fill it with whatever helps. The example passes three keys: a website address, a country, and a city. The readme then lists four reasons to use it, from filling in fields a page left incomplete to supplying domain knowledge, adding constraints, and merging data from several sources.

The option is documented by example and by list, with no type, no schema, and no statement of how the object reaches the model. Whether it is serialised into the prompt, appended as a system message, or passed as a separate field changes both the token cost and the failure modes, and the readme does not say which happens.

The harder gap is precedence. The readme says the model considers both the content and the context, and that is the whole of the contract. If the page says a product costs one price and your context says the city is in another country, nothing states which one the extracted record reflects. In a pipeline whose output feeds pricing decisions, that is not a documentation nicety; it is the difference between a field you can trust and a field you have to re-check against the source page by hand.

## Both flagship examples are fragments that point at scripts in the source

The two headline examples are both cut off. The e-commerce example stops in the middle of a description string inside a schema definition. The browser agent example stops partway through a navigation call, at a truncated address. Neither is a working sample, and neither is presented as an excerpt.

What the readme offers instead is a pointer. A callout under the first example tells you to run a named script or open a file under the source tree, and the manifest confirms the pattern: five of its scripts exist only to execute example files, covering a local run, a usage sample, a browser extraction, a markdown conversion sample, and a top level example. The readme's code blocks are advertisements for scripts rather than copies of them.

The provider story has the same shape. Four providers are listed in the install section with one line each, and every example uses one of them, the Google one, with a model name and a temperature of zero. The other three are named and never demonstrated, and the readme's own text says any LangChain chat model is accepted, which means the caller owns the configuration that makes extraction repeatable.

## The environment template ends its log level with a trailing space

The repository ships a template environment file with two provider keys, a test timeout, and a log level. The log level value is followed by a space before the line ends. A comparison against the string for that level will fail on the value as written, and whether it fails depends on whether the loader trims, which is a property of the loader rather than of this file.

The template also raises a question it does not answer. It carries keys for two of the four providers the install section lists, so a reader following the Google example needs one file and a reader following the Anthropic example discovers a missing key from a runtime error rather than from the documentation.

One of the readme's three callout boxes is an advertisement for the vendor's hosted product, pitched at teams tracking competitor pricing and promotions at scale. The other two are a tip about running a script and a note about the peer dependency. That ratio is worth knowing before you read the feature list as an assessment of the library rather than as a page on a company's site.

## Conclusion

Use extractor when you already have a LangChain chat model and a schema, and when pages are messy enough that a hand written parser would be worse. Check four things before you build a pipeline on it. That you pin the version yourself, because the manifest in the default branch declares a number five patch releases behind the newest tag and the readme names no version at all. That you keep the conversion ground truth honest, since a script regenerates the expected output and a second script pulls it from a moving branch. That you set the model temperature to zero yourself, because nothing in the library enforces it and one of the two examples omits it. And that you know what happens when your custom prompt and your page disagree, which the readme leaves entirely to the model.

## FAQ

### What does the lightfeed/extractor library do?

It sends HTML, markdown, or plain text to a LangChain chat model together with a Zod schema and returns structured data. It converts HTML into markdown first if you want that, recovers malformed JSON output, validates and repairs URLs, and accepts any LangChain chat model. The manifest requires Node 18 or newer and declares a peer dependency on the LangChain core package at 1.1.31 or above.

### How do I install the lightfeed/extractor and pick a model provider?

Install the extractor together with the LangChain core package, then install one provider package: one for OpenAI, one for Google Gemini, one for Anthropic, and one for locally run models through Ollama. The core package is a required peer dependency that the readme says must be installed explicitly to avoid version conflicts with any provider.

### Can lightfeed/extractor handle pages that need to be clicked through first?

Not by itself. The readme pairs it with a separate browser agent package that navigates pages using natural language commands, covering search, pagination, and dismissing popups, and then hands the resulting page to the extractor for structured output. The main example instead uses Playwright to load a page and pass its HTML to the extractor directly.

### What licence is lightfeed/extractor released under and how often does it change?

The manifest declares the Apache 2.0 licence. Three releases are published, the newest tagged in June 2026, with the two before it both tagged in April 2026 minutes apart. The manifest in the default branch declares an earlier version number than the newest tag, and the readme states no version at all.

## Sources

- [License: Apache-2.0](https://github.com/lightfeed/extractor/blob/main/LICENSE)
- [lightfeed/extractor on GitHub](https://github.com/lightfeed/extractor)
- [Project website](https://lightfeed.ai)
- [README](https://github.com/lightfeed/extractor/blob/main/README.md)
- [Releases](https://github.com/lightfeed/extractor/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/lightfeed-extractor
