Model or dataset
feder-cr/resume_render_from_job_description avatar
feder-cr/resume_render_from_job_description

jevos: On-Device Yes/No Decisions Without Cloud Dependencies

Resume_Builder_AIHawk is a powerful Python tool that allows you to automatically customize your resume based on a job URL, ensuring it perfectly aligns with the job requirements and skills. With an interactive command-line interface, this tool makes it easy to navigate through options and select from various pre-defined styles

1,074 stars124 forksPythonMIT

At a glance

What is it?
jevos is a local inference server that classifies text against yes/no questions in 50 to 220 milliseconds on a laptop CPU. It runs a quantized GGUF model via llama.cpp and returns probability scores instead of generated text, keeping every decision on the machine that runs it.
Who is it for?
jevos suits teams that route support tickets, triage content, or enforce policy rules and need decisions that are fast, local, and auditable without per-token cloud costs. It is the wrong tool for multiple-choice classification, scoring, or any task that needs generated text.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What jevos Does and Who It Is For

jevos solves one specific problem: answering yes/no questions about text quickly and without a network call. The use case is decision routing. A customer support pipeline checks whether a message is a billing problem, an upset customer, or a misdelivery complaint. A content system checks whether a post triggers a moderation rule. A triage tool checks whether a support ticket belongs to a given queue. In each case, the answer is a probability between 0 and 1, and the decision threshold is yours to set in your application.

The tool is built for engineers who need this kind of classification embedded in a server-side pipeline. It is not a chatbot, a summarizer, or a general-purpose language model wrapper. The output type is called noul: the probability that the answer to your yes/no question is yes. You send a state (any string or JSON object) and one or more named questions; you get back a noul value for each. That is the entire interface, and the simplicity is intentional.

How jevos Works: GGUF Runtime and CPU Inference

The server runs a quantized GGUF model file through llama.cpp. The runtime binary is fetched separately with the download command, which pulls a platform-specific llama.cpp build for your machine. When you start the server, you point it at the model file, and the server listens on port 8017 by default.

The model returns a probability rather than a text sequence. This is what makes the latency achievable: 50 ms for a single question, and roughly 165 ms for three questions on the same state. Multiple questions in one request share the same state read, so batching is efficient. The README comparison table shows context window sizes for local alternatives: jevos uses 8,192 tokens, while Laya, another local option, is capped at 512.

The server speaks TypeSafe Jev's wire format. Code written for Jev's SDK works unchanged for yes/no questions. The served model appears in the model list as both its real name and the alias jev-latest. Every response carries a Server-Timing header with the inference time.

Installing jevos and Making Your First Request

The repository uses uv for dependency management. After cloning, sync the environment, then fetch the llama.cpp runtime for your machine:

bash
uv sync
uv run jev download --only runtime

Download jevos-q4_k_m.gguf from the GitHub release tagged jevos, then start the server:

bash
uv run jev serve --gguf jevos-q4_k_m.gguf --device cpu --threads 16

Set --threads to your CPU core count. Once the health endpoint at GET /health returns ready, send a request:

bash
curl http://127.0.0.1:8017/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "jev-latest",
  "state": "I was charged twice for the same order.",
  "questions": {"billing": {"type": "noul", "instructions": "Is this a billing problem?"}}}'

The response carries a noul value between 0 and 1. The server also exposes a /dino path that shows a visual of the model running.

Batching Multiple Questions on the Same State

When your routing logic needs several decisions about the same piece of text, put them in one request. The state is read once for all questions in a batch, so three questions together take roughly 165 ms against 103 ms for a single one, according to the README. This makes batching meaningfully more efficient than making sequential single-question requests.

From Python, the endpoint is a straightforward call using the requests library:

python
import requests

answer = requests.post("http://127.0.0.1:8017/v1/systemone", json={
    "model": "jev-latest",
    "state": "I was charged twice for the same order.",
    "questions": {"billing": {"type": "noul", "instructions": "Is this a billing problem?"}},
}).json()

if answer["answers"]["billing"]["noul"] > 0.5:
    print("send to billing")

For offline use without a running server, the jev decide subcommand processes a single request file. The --output flag writes the answer to a new file and refuses to overwrite an existing one. A file without a model field is treated as a native engine request and returns full output including prompt hashes and timings, which is useful for debugging inference behaviour.

The 422 Limit: Only Yes/No Questions Are Accepted

jevos handles only questions of type noul. If you send a question with type choice or score, the server responds with HTTP 422 and refuses the request. The README comparison table shows that Jev, the cloud counterpart, already supports multiple choice and scores; jevos marks those as coming soon. Until they ship, any use case that needs ranked choices, probability distributions over several options, or a numeric score on a scale requires a different tool.

CPU inference is also a real constraint. The --device flag defaults to auto, but the README examples all use --device cpu. Running on a machine without a dedicated GPU means inference speed depends directly on thread count and CPU performance. The README's guidance is to match --threads to your core count and reduce it if other heavy processes are competing.

The model is quantized at q4_k_m precision, which reduces memory and speeds inference at the cost of some precision loss compared to full-precision weights. The README does not document accuracy benchmarks for the shipped model, which means you cannot assess fit for your use case without testing against your own data.

Comparing jevos to Jev and Laya

The README positions jevos in a three-way comparison. Jev is the cloud service: it runs on Jev's servers, charges per token, and already supports multiple choice and score question types. Jevos is the local equivalent for yes/no work: free to run, no per-token cost, and limited to noul questions for now. The wire format is shared between them, so switching between jevos and Jev for yes/no use cases requires only a URL change and an API key, not a code rewrite.

Laya is a second local alternative listed in the table. Like jevos, it runs on your machine at no per-token cost and handles yes/no questions. The practical difference is context size: Laya supports 512 tokens per request, while jevos handles up to 8,192. For routing decisions based on long documents, multi-paragraph messages, or ticket threads with history, that difference is the relevant factor when choosing between them.

Licence and the Scope of the jev CLI

jevos is released under the MIT licence, which imposes no restrictions on commercial use or modification. The project's last push was on 2026-09-28 and the jevos release appeared on 2026-09-27, so the project is under active development. The pyproject.toml identifies the package as jev version 0.0.1, which signals early-stage software. The README credits Loris Salsi as a co-author alongside the primary contributor.

The package installs as the jev command-line tool. Beyond serve and download, the decide subcommand runs a single offline decision without a server. The server options table lists the configurable parameters: --gguf for the model file path, --device for cpu or auto, --threads (default 4), and --host and --port for binding (defaulting to 127.0.0.1 and 8017). Two additional endpoints exist: GET /v1/models lists the served model and its jev-latest alias, and GET /health returns a ready status once the model finishes loading.

Editorial conclusion

jevos suits teams that route support tickets, triage content, or enforce policy rules and need decisions that are fast, local, and auditable without per-token cloud costs. It is the wrong tool for multiple-choice classification, scoring, or any task that needs generated text. Before adopting it, verify that jevos-q4_k_m handles the phrasing patterns in your production data: the README does not document accuracy benchmarks, so internal testing against real inputs is the necessary first step.

Frequently asked questions

What model file does jevos use, and where do I download it?

jevos uses jevos-q4_k_m.gguf. The README instructs you to download it from the GitHub release tagged jevos, found at the releases section of the repository.

Can jevos handle multiple yes/no questions in a single request?

Yes. You can include multiple named questions in one POST /v1/systemone call. All questions share the same state read, so three questions together take roughly 165 ms versus 103 ms for one alone, according to the README.

What happens if I send a choice or score question type to jevos?

The server returns HTTP 422 and refuses the request. The README explicitly states that choice and score question types are refused with a 422 status code and are not yet supported.

Official sources

  1. feder-cr/resume_render_from_job_description on GitHub
  2. Issues
  3. License: MIT
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/feder-cr-resume-render-from-job-description.svg)](https://hysenlabs.com/projects/feder-cr-resume-render-from-job-description)