Model or dataset
cactus-compute/needle avatar
cactus-compute/needle

Cactus Needle 2: a 14MB tool-calling model you install with pip

14MB foundation model for tiny devices; phones, wearables, smart home, and robots.

11,090 stars714 forksPythonApache-2.0

At a glance

What is it?
Needle 2 packages a 45M-parameter tool-calling and extraction model into a single 14MB engine driven from Python. The contract is clean and the memory bound is the real story, but the repository documents no recovery path when a tool call is wrong.
Who is it for?
Adopt Needle 2 if you are shipping on a device where 28MB of session memory is the budget and your tools fit the documented shapes: closed sets as enums, bounded numbers, five tools or fewer per turn. Do not adopt it if you need a documented rollback path, a published accuracy number for your own tool surface, or a model you can retrain from scratch rather than adapt with LoRA.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 12 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 17, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem Needle 2 targets: tool calls that fit in 28MB of RAM

Most function-calling stacks assume a server. You send a prompt with tool schemas to a hosted model, get JSON back, execute it yourself. That works until the device is a wearable, a robot arm, or a kitchen appliance with no reliable network and a memory budget measured in tens of megabytes. Needle 2 is aimed squarely at that gap. The README describes it as an open 45M-parameter model for tool calling, device use and structured extraction, shipped as a single 14MB binary that runs a full session in about 28MB of RAM.

The audience is embedded and edge engineers, not backend teams. The repository is the Python package: inference, LoRA fine-tuning, and export. The README states that inference does no network, and that the engine itself is fetched once from Hugging Face and cached. If your product already talks to a frontier model over HTTP and has power to spare, this project solves a problem you do not have.

How the Simple Attention Network and the decode grammar work together

Two mechanisms carry the design. The first is the architecture. Needle 2 is a Simple Attention Network: a Hadamard MLP in place of the FFN, GQA attention, engram key-value memory, and multi-lane hyper-connections. The README gives the block update rule in some detail: x-hat is the RMS-normalised flattening of four residual streams, H is the orthonormal Walsh-Hadamard transform applied in n log n time with no weights to read, (k, v) rows are gathered from hashed n-gram tables, and P is a doubly-stochastic normalisation of routing logits computed by Sinkhorn iteration. Weights are compressed to CQ2-bit with Cactus Quants. That fixed Hadamard matrix is the reason the model can be small without the FFN dominating parameter count.

The second mechanism is the one you actually interact with. Tool calls come back as structured data, text in, JSON out, and a byte-level grammar compiled from your schemas constrains every token during decoding. This matters more than the architecture for day-to-day reliability: the model cannot emit a room name outside your enum, because the grammar will not allow the token. A retrieval head renders only the top five tools per turn when you declare a large catalogue, and the grammar is constrained to that subset. Memory stays bounded because a 256-token sliding window pins the tools as KV sinks, so total memory stays near 28MB no matter how long the conversation runs.

Installing cactus-needle and running a first tool call

The package installs from PyPI. The runtime package does not install the training stack, so a plain install is small; add the train extra only when you need fine-tuning or checkpoint export.

bash
pip install cactus-needle
bash
pip install "cactus-needle[train]"

Once installed, the simplest path is to decorate a function. The signature gives the argument types, the docstring is the tool description, and run() completes the loop: the model picks the call, Needle executes your function, feeds the result back, and returns the final response with executed tool results attached as results.

python
import needle

@needle.tool
def get_weather(city: str):
    "Get the current weather for a city."
    return {"city": city, "temp_c": 27, "sky": "clear"}

agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])

The README shows the expected output as a list containing the dict your function returned, with city Lagos, temp_c 27 and sky clear. For extraction, declare a Pydantic model and call extract(); the README's example returns a typed object whose vendor and total fields print as Acme Corp and 1200.0.

python
from pydantic import BaseModel

class Invoice(BaseModel):
    vendor: str
    total: float
    due_date: str

invoice = needle.extract("Invoice from Acme Corp, $1,200.00, due 2026-09-01", Invoice)
print(invoice.vendor, invoice.total)

There is also a browser playground started from the CLI. The README notes the server downloads and initializes the model before serving, so the first query is instant, and that the Finetune on these tools button runs the fine-tuning pipeline from the UI and hands back a downloadable .cact file.

bash
needle playground
needle playground --weights my.cact

Bounded memory is the selling point, and a 256-token window is the cost

The sliding window is the most consequential design choice in the README and the one most likely to surprise you. A 256-token window with tools pinned as KV sinks means memory does not grow with conversation length. It also means the model does not see the whole conversation. A session that runs long enough for the window to slide past an earlier tool result will have that result dropped from context. Nothing in the README describes a summarisation step or a way to widen the window, and doc/apis.md is referenced for the response contract rather than for window tuning.

The confidence head is the second thing to inspect closely. Every response carries a calibrated confidence score from a learned head, and the README says to set a threshold, act above it, escalate below it. That is a sensible pattern, but the README publishes no calibration table per environment, so the threshold you pick is a guess until you measure it against your own acceptance suite. Treating the score as meaningful without that measurement is the most likely way to ship a bad experience quietly.

Where Needle 2 is the wrong tool

The README states the constraint directly: keep shapes as closed sets, bounded numbers, verbatim copy for free text, and five tools or fewer. If your tool surface is open-ended, if arguments are free-form strings the model must compose rather than copy, or if a single turn needs to chain more than a handful of calls, the constrained decoding stops helping and starts hurting. The grammar can guarantee a syntactically valid call; it cannot guarantee the call is the right one.

The second failure mode is operational. The README does not document rollback, version pinning for the fetched engine, or what happens when a tuned .cact is exported against a different base revision. If you need reproducible model artifacts with a documented downgrade path, that gap is a real risk, not a documentation nit. The third case is simply scale: this is a 45M-parameter model that the README says trades wins with FunctionGemma 270M, LFM2.5 230M and Apple FM on the published benchmarks. Trading wins means it loses some. For a task where you need the strongest available reasoning, a larger hosted model is the correct choice.

Needle 2 versus a hosted function-calling API

The obvious alternative is a hosted function-calling endpoint from a major provider. The difference in approach is where the constraints live. With a hosted model, you send the schemas in the prompt and validate the returned JSON yourself; correctness depends on the model following instructions, and you pay per token with a network round trip. With Needle 2, the schema is compiled into a byte-level grammar before decoding starts, so invalid tokens are unreachable, and inference runs locally with no network. The trade is capability for determinism and locality.

A second alternative is a general small language model running under a local runtime such as llama.cpp or an on-device framework, with a separate JSON-schema validator bolted on. That gives you a wider model choice and a familiar toolchain, but you own the grammar plumbing, the memory ceiling, and the retrieval of relevant tools from a large catalogue. Needle 2 bundles all three: the engine, the grammar compiler, and a retrieval head that renders the top five tools per turn.

Licence, packaging and the cost of staying current

The repository LICENSE is MIT, while pyproject.toml declares license text Apache-2.0 for the package. Those two statements do not agree, and anyone shipping Needle 2 inside a product should resolve the discrepancy with the maintainers before relying on either. This is a factual conflict in the repository, not a legal opinion; treat it as an open question rather than a settled term.

Upgrade cost is low by design. The runtime dependency list is a single package, huggingface_hub, and the training stack is an optional extra, so a version bump does not drag in a large dependency tree. The engine is fetched once and cached, which means the artifact you run is not necessarily the artifact you tested unless you pin it. The repository shows no retrieved releases, so there is no changelog to read before upgrading. For a device fleet, that combination argues for pinning the package version and the cached engine together, and for running the environment acceptance suite before rollout rather than after.

Editorial conclusion

Adopt Needle 2 if you are shipping on a device where 28MB of session memory is the budget and your tools fit the documented shapes: closed sets as enums, bounded numbers, five tools or fewer per turn. Do not adopt it if you need a documented rollback path, a published accuracy number for your own tool surface, or a model you can retrain from scratch rather than adapt with LoRA. Before committing, run the frozen acceptance suite for the environment closest to your product and check the confidence threshold against your own escalation policy, because the README describes the score as calibrated but publishes no per-environment calibration table.

Frequently asked questions

What is Needle 2 used for?

It is an open 45M-parameter model for tool calling, device use and structured extraction, packaged as a Python library. The README targets phones, wearables, smart home and robots, and lists ready-made tool surfaces such as smart_home, media_player and wearable.

How do I install Needle 2?

Install the runtime package with pip install cactus-needle. Add the train extra with pip install "cactus-needle[train]" when you need LoRA fine-tuning or checkpoint export, since the runtime package does not install the training stack.

Does Needle 2 run offline?

The README states that inference does no network, and that the engine is fetched once from Hugging Face and cached. It also points to doc/apis.md for offline setup on air gapped devices.

How much memory does a Needle 2 session use?

The README gives the model as a single 14MB binary and a full session at about 28MB of RAM. A 256-token sliding window with tools pinned as KV sinks keeps total memory near that figure regardless of conversation length.

How do I fine-tune Needle 2 on my own tools?

Fine-tuning uses LoRA on the frozen base, with the adapter merged at export so the tuned model is still a single .cact file running on the same engine. The workflow is to synthesize data, LoRA fine-tune, then build a tuned .cact, with dataset sizing and loss-curve reading covered in doc/finetuning.md.

Official sources

  1. Official documentation
  2. Official README
  3. Project repository
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cactus-compute-needle.svg)](https://hysenlabs.com/projects/cactus-compute-needle)