Model or dataset
promptslab/Promptify avatar
promptslab/Promptify

Promptify: A Task-Based Python NLP Library with Pydantic Structured Outputs

Prompt Engineering | Prompt Versioning | Use GPT or other prompt based models to get structured output. Join our discord for Prompt-Engineering, LLMs and other latest research

4,641 stars363 forksPythonApache-2.0

At a glance

What is it?
Promptify is a Python library that wraps LiteLLM to run NLP tasks such as named entity recognition, classification, and question answering against any of 100-plus LLM providers, returning type-safe Pydantic objects instead of raw strings. It is described by its README as a scikit-learn for LLM-powered NLP.
Who is it for?
Python developers building pipelines that apply NLP tasks at scale, such as document classification, entity extraction from medical records, or SQL generation from natural language, will find Promptify reduces the integration work to a few lines per task. The library is classified as Development Status 4: Beta in pyproject.toml, and the last push was on 2026-03-27.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What Promptify Solves and Who It Is For

Building LLM-powered NLP pipelines traditionally requires writing prompt templates, parsing free-text model outputs, handling provider differences, and adding retry logic separately for each task. Promptify addresses this by providing task-specific classes that handle prompting, output parsing, and provider routing through a single call.

The README describes the library as a task-based NLP engine and positions it as "scikit-learn for LLM-powered NLP": a collection of ready-made task objects with a consistent fit-predict style interface, but for language tasks rather than machine learning models. The target audience is developers and researchers who need to apply structured NLP operations, such as named entity recognition or text classification, to text at scale, without training their own models. Because Promptify routes through LiteLLM, the same task object can be pointed at different providers by changing the model string.

The library requires no training data. Tasks are defined by the Pydantic schema of the expected output, and the prompt is constructed by the task class internally. This makes it practical for rapid prototyping, though it also means output quality depends on the underlying model and prompt quality rather than a fine-tuned classifier. The README shows a medical NER example where the domain parameter is set to 'medical', adjusting the internal prompt for clinical entity recognition.

The Task Architecture and LiteLLM Backend

Each NLP operation in Promptify is a class that maps text input to a Pydantic model output. The NER class returns NERResult (a list of Entity objects with text and label fields). The Classify class returns Classification (label and confidence). The QA class returns Answer (answer, evidence, and confidence). The Task class accepts any Pydantic BaseModel as the output_schema, enabling arbitrary structured outputs.

All classes delegate the actual model call to LiteLLM, which provides a unified interface across OpenAI, Anthropic, Google, Ollama, Azure, and other providers. Switching providers requires only changing the model string in the class constructor. The pyproject.toml pins LiteLLM to a range of >=1.50.0,<=1.82.6, which means very new LiteLLM versions are excluded until the dependency range is updated.

The library includes a safe parser that falls back to JSON completion for providers that do not natively support structured outputs, avoiding the use of eval() on model responses. Retry logic is handled by the tenacity dependency, which wraps LiteLLM calls with configurable wait and stop strategies.

The README lists the full table of supported tasks and their output schemas. Multilabel classification uses Classify(multi_label=True) and returns MultiLabelResult. GenerateQuestions returns a list of GeneratedQuestion objects. GenerateSQL returns an SQLQuery object. ExtractRelations and ExtractTable both return ExtractionResult. These output types are Pydantic models, so they can be serialised to JSON, validated, and integrated into data pipelines without additional parsing code.

Installation and a First NER Task

The library requires Python 3.9 or later. Install the base package with:

bash
pip install promptify

To include evaluation metrics (precision, recall, F1, ROUGE), install the eval extra:

bash
pip install promptify[eval]

Running a named entity recognition task on a medical text takes three lines:

python
from promptify import NER

ner = NER(model="gpt-4o-mini", domain="medical")
result = ner("The patient is a 93-year-old female with a medical history of chronic right hip pain, osteoporosis, hypertension, depression, and chronic atrial fibrillation admitted for evaluation and management of severe nausea and vomiting and urinary tract infection")

The result is a NERResult containing Entity objects with text and label fields. The domain parameter is optional and adds context to the prompt. Changing the model string to claude-sonnet-4-20250514 or ollama/llama3 routes the same task to a different provider without any other change.

Supported Tasks and Custom Schemas

The README lists twelve built-in task types: named entity recognition, binary classification, multiclass classification, multilabel classification, question answering, summarisation, relation extraction, tabular extraction, question generation, SQL generation, text normalisation, and topic modelling. Each has a corresponding Pydantic output schema that is returned directly from the task call.

For tasks outside that list, the Task class accepts any Pydantic BaseModel as the output_schema parameter. The README shows an example where a MovieReview schema with sentiment, rating, and key_themes fields is passed directly to Task, and the model returns a typed MovieReview object rather than a raw string response.

Few-shot examples can be added to any task to improve accuracy on domain-specific inputs. The domain parameter accepts any string and adjusts the internal prompt for domain-specific context, such as medical, legal, or financial text. These two levers, few-shot examples and domain strings, are the main ways to tune task performance without modifying the library itself.

Batch Processing, Async, and Cost Tracking

Batch processing is available through the batch method on any task object. The max_concurrent parameter controls how many requests run in parallel:

python
results = ner.batch(["text1", "text2", "text3"], max_concurrent=10)

Async support uses native Python coroutines through the acall method:

python
result = await ner.acall("Patient has diabetes")

Cost tracking is built into every session. The get_cost_summary() method returns token usage and cost data via LiteLLM's built-in cost monitoring. This is useful for estimating running costs when processing large datasets.

The evaluation framework, available with the promptify[eval] extra, supports precision, recall, F1, accuracy, exact match, and ROUGE metrics on a labelled dataset through the evaluate function in promptify.eval.

Limitations and Cases Where Promptify Is the Wrong Choice

Promptify is classified as Development Status 4: Beta in pyproject.toml. The API may change between versions. The litellm version pin (<=1.82.6) means the library does not automatically pick up new LiteLLM provider support until the upper bound is raised.

All task outputs depend on the quality of the underlying model and the internal prompt template. There is no mechanism for fine-tuning the prompts for a specific domain beyond passing the domain string and few-shot examples. Engineers who need deterministic, auditable NLP pipelines will find that the output quality varies with model and task phrasing in ways that cannot be fully predicted without testing.

The library does not include data ingestion, storage, or pipeline orchestration. For workflows that need to pull text from APIs, store results in a database, and retry failed records, Promptify needs to be embedded in a larger system that provides those capabilities.

The test suite uses asyncio_mode = 'auto' as configured in pyproject.toml. Contributors running tests need pytest-asyncio installed from the dev extras. The README does not document a CI pipeline or automated test workflow, so the test coverage of production use cases is not visible from the repository.

How Promptify Compares to LangChain

LangChain is a widely used Python framework for building LLM-powered applications. It provides chains, agents, memory, retrieval, and a large library of integrations for databases, tools, and data sources. Its scope is much wider than Promptify's.

Promptify is focused on applying NLP task types to text and returning structured outputs. It does not include document loaders, vector stores, or agent loops. For a developer who needs to run a single NER or classification task on a batch of texts, Promptify's three-line API is more direct than constructing a LangChain chain with a custom output parser.

For workflows that require retrieval-augmented generation, conversational memory, multi-step reasoning, or integrations with tools and databases, LangChain's broader feature set is more appropriate. The two libraries can also be used together: LangChain handles orchestration and retrieval while Promptify's typed task objects handle the NLP steps that need Pydantic-validated outputs.

Another relevant difference is that LangChain's structured output support is general-purpose and relies on model-specific tool-calling APIs. Promptify's safe parser explicitly handles the case where a provider does not support native structured output, adding a fallback path without requiring the developer to configure it.

Maintenance Status and Licence

The last push to the repository was on 2026-03-27. The project carries no GitHub releases. The pyproject.toml lists the development status as 4: Beta. The licence is Apache-2.0.

The README links to a Discord server at discord.gg/m88xfYMbK6 for the PromptsLab community. Contributions go through GitHub pull requests as described in contribute.md. There is no automated release pipeline documented in the repository.

The repository includes a benchmarks/ directory and a notebooks/ directory alongside the main promptify/ package. The examples/ directory contains at least one worked example (examples/medical_ner.py), giving contributors a reference implementation to test against. The mypy and ruff configurations in pyproject.toml indicate that the project enforces type annotations and style guidelines, though contributors need the dev extras installed to run those checks locally.

Editorial conclusion

Python developers building pipelines that apply NLP tasks at scale, such as document classification, entity extraction from medical records, or SQL generation from natural language, will find Promptify reduces the integration work to a few lines per task. The library is classified as Development Status 4: Beta in pyproject.toml, and the last push was on 2026-03-27. Engineers who need a production-grade, actively maintained LLM framework with a large ecosystem of integrations should evaluate LangChain or a similar alternative before committing to Promptify.

Frequently asked questions

What Python version does Promptify require?

Promptify requires Python 3.9 or later, as declared in pyproject.toml. It supports and is tested against Python 3.9 through 3.13.

Can Promptify work with local models like Ollama?

Yes. Because Promptify uses LiteLLM as its backend, passing a model string like ollama/llama3 routes the task to a locally running Ollama instance. The README shows this as one of three model string examples alongside OpenAI and Anthropic models, with no other code changes required.

Does Promptify include evaluation metrics?

Yes, with the eval optional extra. Install it with pip install promptify[eval], which adds precision, recall, F1, accuracy, exact match, and ROUGE metric support through the evaluate function in promptify.eval.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. Project website
  4. promptslab/Promptify on GitHub
  5. README
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/promptslab-promptify.svg)](https://hysenlabs.com/projects/promptslab-promptify)
Community notes

Community notes