Library / SDK
tavily-ai/tavily-python avatar
tavily-ai/tavily-python

tavily-python: A Python SDK for Search, Extract, Crawl, Map and Research

The Tavily Python SDK allows for easy interaction with the Tavily API, offering the full range of our search, extract, crawl, map, and research functionalities directly from your Python programs. Easily integrate smart search, content extraction, and research capabilities into your applications, harnessing Tavily's powerful features.

1,409 stars190 forksPythonMIT

At a glance

What is it?
The tavily-python package wraps the Tavily API in a small Python client with a keyless trial mode. It is aimed at RAG pipelines and agents that need web content, and its crawl endpoint is still invite-only.
Who is it for?
Adopt tavily-python if you are building a Python RAG pipeline or agent and want web search, page extraction and site traversal behind one client object. Do not adopt it if you need a self-hosted stack or cannot accept per-call API costs, because every method except the rate-limited keyless search and extract path talks to Tavily's servers.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 14 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What tavily-python actually removes from your code

The package is a thin wrapper, not a search engine. Tavily runs the retrieval service; the SDK gives Python callers named methods instead of raw HTTP. The README describes it as offering "the full range of our search, extract, crawl, map, and research functionalities directly from your Python programs".

The intended reader is someone building a retrieval-augmented generation pipeline or an agent loop. The README's own examples point that way: get_search_context() returns a string meant to be fed into a RAG application, and qna_search() returns an answer the README says is "Perfect for usage by LLMs". If your job is to hand a language model current web text, this library is the adapter. If your job is to rank documents yourself or run your own crawler, it is not.

The practical value is small and specific: one client object, one API key, and JSON responses you can index into. Nothing here manages embeddings, chunking or vector storage.

Client, API key, and the keyless path

Every method hangs off a TavilyClient instance. With an API key, the constructor takes TavilyClient(api_key="tvly-YOUR_API_KEY") and all endpoints are available. Without one, TavilyClient() runs in keyless mode against the public Tavily API.

Keyless mode is the interesting design choice. The README states it supports search() and extract() only, and that other methods raise an error explaining an API key is required. Keyless usage is also rate-limited, and the SDK exposes a dedicated exception, TavilyKeylessLimitError, for when the cap is reached. That exception carries structured fields the README lists as code, window, retry_after_seconds and next_actions.

That is a more honest trial path than a free tier you have to sign up for, but it is a trial, not a deployment mode. Any code path that reaches crawl, map or research will fail without a key, so keyless mode is useful for a first script and misleading as a foundation.

The four endpoints and what each returns

search() takes a query and returns a response object. Two conveniences sit on top of it. get_search_context() returns a single context string for RAG, and qna_search() returns a short answer. The README also documents exact_match=True, which restricts results to pages containing the exact phrase inside quotes, and gives the example query '"John Smith" CEO Acme Corp'.

extract() takes a list of URLs and returns raw page content. The README says you can supply up to 20 URLs simultaneously, and that pages which could not be extracted land in response["failed_results"] rather than disappearing. That failure bucket matters: partial success is normal when scraping the open web, and you need to handle it explicitly.

crawl() traverses a site from a starting URL. The README example passes url, max_depth=3, limit=50 and an instructions string, and it notes crawl is "currently available on an invite-only basis" with a pointer to crawl.tavily.com. map() discovers site structure from a base URL. The README section for map is truncated in the repository copy, so the exact parameters are not verifiable from what is published there.

The response shapes are dictionaries, not typed models. You index into response["results"], response["failed_results"] and per-result keys like url, raw_content and images. There is no schema object to validate against, which is a real ergonomic gap for larger codebases.

Installing tavily-python and running a first search

The README gives one install command. Note the package name uses a hyphen while the import name uses an underscore.

bash
pip install tavily-python

After installation, the shortest working script uses keyless mode, so no signup is needed to see whether the response shape suits you. Run this and you should see a dictionary printed with a results list.

python
from tavily import TavilyClient, TavilyKeylessLimitError

client = TavilyClient()

try:
    response = client.search("Who is Leo Messi?")
    print(response)
except TavilyKeylessLimitError as e:
    print(e)
    print("retry after:", e.retry_after_seconds, "seconds")

When you want crawl, map or research, or simply higher limits, the README directs you to sign up at tavily.com and pass the key into the constructor.

python
from tavily import TavilyClient

tavily_client = TavilyClient(api_key="tvly-YOUR_API_KEY")
response = tavily_client.search("Who is Leo Messi?")
print(response)

For a RAG-shaped first use, get_search_context() collapses search plus formatting into one call and returns a string you can drop straight into a prompt.

python
from tavily import TavilyClient

client = TavilyClient(api_key="tvly-YOUR_API_KEY")
context = client.get_search_context(query="What happened during the Burning Man floods?")
print(context)

setup.py declares python_requires='>=3.8' and pulls in requests, httpx and tiktoken>=0.5.1, so those arrive automatically. The examples directory ships company_information.py, hybrid_rag.py and openai_assistant.py if you want runnable starting points.

Where the SDK stops being the right tool

The library is a client for a hosted service, so the usual constraints of that arrangement apply. There is no offline mode, no local index, and no way to point the client at your own backend. If your application cannot send queries to a third party, the SDK is the wrong layer entirely, no matter how clean the interface is.

Keyless mode has a hard ceiling. Search and extract work; crawl, map and research raise. A prototype built on keyless mode will need a key before it can do site traversal at all, and the rate limit means even search-heavy testing will hit TavilyKeylessLimitError.

Crawl is gated. The README states it is invite-only and points to crawl.tavily.com, so you cannot assume it is available to you on day one. Building a pipeline whose core step is crawl() means depending on access you may not have.

Extraction is best-effort by design. The failed_results bucket exists because some URLs will not extract, and the README does not describe retry behaviour, per-URL error codes, or timeouts in the section that is published. Plan for partial results rather than complete ones. The README also does not document rollback, versioning policy or deprecation handling for the SDK itself, so pinning a version in your own requirements file is the only control you have over upgrades.

How this differs from wiring up a search API yourself

The obvious alternative is calling a general web search API directly, for example a search engine's own HTTP endpoint, and then fetching and parsing pages with requests and BeautifulSoup. The difference is where the work lives. With a raw search API you get links and snippets, and you own fetching, HTML stripping, boilerplate removal and the decision about what counts as page content. With tavily-python, extract() returns raw_content already pulled from the page, and get_search_context() returns an assembled context string.

That trade is real in both directions. You give up control over chunking and cleaning, and you inherit the service's extraction quality and its failed_results. You gain a much shorter path from query to prompt-ready text, which is the whole point for an agent loop.

A second alternative is a self-hosted retrieval stack built on your own crawler and an open source search index. That keeps data in your infrastructure and costs nothing per call, but you take on crawling politeness, JavaScript rendering, freshness and index maintenance. For a small team shipping an agent, that is usually a much larger project than the feature warrants. The README's hybrid_rag.py example suggests the authors expect the SDK to sit alongside a vector store rather than replace one.

Licence, maintenance and upgrade cost

tavily-python is MIT licensed, and setup.py carries the matching classifier. MIT is permissive: you can use, modify and redistribute the SDK, including in closed products, provided the copyright notice and licence text travel with it. That covers the wrapper code only. The Tavily API behind it is a separate commercial service with its own terms, and nothing in the repository speaks to pricing, quotas or data handling. Treat the licence question and the service contract as two separate reviews.

The repository is not archived, and the last push was on 2026-09-09, which is recent. setup.py declares version 0.8.3, and the leading zero is worth noticing: the project does not present itself as API-stable. The README does not describe a deprecation policy, so a minor bump could change response handling without a major version signal. Pinning tavily-python in your requirements file is cheap insurance.

Upgrade cost is mostly in the response dictionaries. Because results are plain dicts with keys like raw_content and failed_results rather than typed objects, a renamed key shows up as a KeyError at runtime, not as a type error at import. Wrapping the fields you read in small accessor functions keeps that blast radius contained.

Editorial conclusion

Adopt tavily-python if you are building a Python RAG pipeline or agent and want web search, page extraction and site traversal behind one client object. Do not adopt it if you need a self-hosted stack or cannot accept per-call API costs, because every method except the rate-limited keyless search and extract path talks to Tavily's servers. Before writing it into a production path, verify your current rate limits and whether crawl access has been granted at crawl.tavily.com, since the README states that endpoint is invite-only.

Frequently asked questions

What is tavily-python?

It is a Python wrapper around the Tavily API that exposes search, extract, crawl, map and research as methods on a TavilyClient object. The README describes it as offering those functionalities directly from your Python programs.

Can I use tavily-python without an API key?

Yes, in keyless mode. Instantiating TavilyClient() with no arguments runs against the public Tavily API, but the README states keyless mode supports search() and extract() only, that other methods raise an error saying an API key is required, and that keyless usage is rate-limited.

How do I install tavily-python?

The README gives a single command, pip install tavily-python. The package name uses a hyphen while the import name is tavily, and setup.py requires Python 3.8 or later.

Which endpoints are available in tavily-python?

The README documents search, extract, crawl, map and research. Search and extract work in keyless mode, while crawl, map and research need an API key, and crawl is described as invite-only.

Official sources

  1. Issues
  2. License: MIT
  3. Project website
  4. README
  5. tavily-ai/tavily-python on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/tavily-ai-tavily-python.svg)](https://hysenlabs.com/projects/tavily-ai-tavily-python)