# Podcastfy: an open source Python alternative to NotebookLM's podcast feature

> Podcastfy turns websites, PDFs, images, YouTube videos and plain topics into two-voice audio conversations through GenAI. It is a library and CLI first, not a hosted product, and the quality of the output depends almost entirely on the API keys you bring.

**souzatharsis/podcastfy** — An Open Source Python alternative to NotebookLM's podcast feature: Transforming Multimodal Content into Captivating Multilingual Audio Conversations with GenAI

- Repository: https://github.com/souzatharsis/podcastfy
- Website: https://www.podcastfy.ai
- Stars: 6,573 · Forks: 762
- Language: Python
- License: Apache-2.0
- Published: 2026-09-14 · Updated: 2026-09-14 · Language: en
- Canonical page: https://hysenlabs.com/projects/souzatharsis-podcastfy

## What Podcastfy actually solves, and for whom

NotebookLM's podcast feature is a button inside a closed product. You paste sources, you get a two-host conversation, and you have no way to call that from a script, batch it over a hundred documents, or swap the voice provider. Podcastfy exists to fill that gap. The README frames the project as an "Open Source API alternative to NotebookLM's podcast feature", and the emphasis is on programmatic and bespoke generation rather than research synthesis.

The target user is a Python developer who already has content somewhere: a folder of PDFs, a list of URLs, a YouTube link, a set of images, or just a topic string. The package accepts all of those as input and returns a path to an audio file. That is the whole contract. If you need a drag-and-drop interface, the README points to a web app at openpod.fly.dev and to a Colab notebook, but the primary surface is the library.

The second audience is anyone who wants the conversation in a language other than English. The README lists sample audio in French and Portuguese-BR, generated from a website and a news article respectively, which suggests the multilingual path is a first-class feature rather than an afterthought.

## How the pipeline turns a URL into a two-voice conversation

The dependency list in pyproject.toml is the clearest description of the architecture. Content acquisition uses beautifulsoup4 for web pages, PyMuPDF for PDFs, and youtube-transcript-api for YouTube. Images are handled through the same multimodal path as text, since the project description covers text and images together. Those inputs are normalised into text that an LLM can read.

Generation runs through LangChain. The dependencies include langchain, langchain-google-genai, langchain-google-vertexai, langchain-community and litellm, plus the direct SDKs openai and google-generativeai. So the script generation step is provider-agnostic in principle: you can point it at Gemini, at OpenAI, or at Vertex AI depending on which keys you set. The conversation structure itself is not hardcoded in Python. Two YAML files, podcastfy/config.yaml and podcastfy/conversation_config.yaml, are shipped inside the package, and pyproject.toml explicitly includes them in the distribution.

Speech synthesis is the last stage. The dependencies cover ElevenLabs, edge-tts, and google-cloud-texttospeech. Audio assembly uses pydub and the ffmpeg package, which is why the prerequisites list ffmpeg separately. The output is a single audio file whose path generate_podcast returns.

One design consequence worth naming: because the LLM writes the dialogue and a separate TTS engine voices it, the two halves can be tuned independently, but they can also fail independently. A transcript that reads well can still sound wrong if the voice mapping in conversation_config.yaml does not match the speakers the LLM invented.

## Installing Podcastfy and generating a first audio file

The README states two prerequisites: Python 3.11 or higher, and ffmpeg for audio processing. The Python constraint is enforced in pyproject.toml as python = "^3.11".

Install the package from PyPI:

```bash
pip install podcastfy
```

Before generating anything you need API keys. The README links to usage/config.md for that step and does not inline the key names, but docker-compose.yml shows which environment variables the container expects: GEMINI_API_KEY and OPENAI_API_KEY. Set the ones your chosen providers need.

The shortest working call is the Python API. The README gives this example, which takes a list of URLs and returns the path to the generated audio file:

```python
from podcastfy.client import generate_podcast

audio_file = generate_podcast(urls=["<url1>", "<url2>"])
```

There is an equivalent command-line entry point. The README shows the module invocation with repeated --url flags:

```bash
python -m podcastfy.client --url <url1> --url <url2>
```

If you would rather not manage the Python environment, the repository ships a Dockerfile and a docker-compose.yml. The compose file defines a podcastfy service that builds from that Dockerfile, exposes port 8000, and runs a healthcheck that simply imports the package. The Dockerfile itself installs podcastfy from PyPI inside an Ubuntu 24.04 image with ffmpeg present. Note that the compose file reads the API keys from your shell environment, so they must be exported before you run it.

## Where Podcastfy breaks down or is the wrong choice

The honest limitation is that Podcastfy is a generator, not a studio. It produces a file. The README does not document any editing step, any transcript review workflow, or any way to regenerate a single segment without rerunning the whole pipeline. If your requirement is fine control over pacing, pronunciation of proper nouns, or a specific host personality, you will be fighting the prompt and the voice config rather than editing an output.

Cost and latency are externalised to you. Every run spends LLM tokens and TTS characters on providers you configure. Nothing in the repository meters that, and the README does not describe a caching layer for previously fetched URLs. A batch job over a large corpus will re-fetch and re-generate unless you build that yourself.

The FastAPI path is explicitly labelled beta in the README, with the note that you should look at the notebook for a clear example of making requests. Treat the HTTP surface as unfinished. The web app at openpod.fly.dev is a separate hosted instance, not something you run.

There is also a maintenance signal to read carefully. The most recent release listed is v0.4.0 from 2024-11-16, while pyproject.toml declares version 0.4.3 and the last push to the repository was on 2026-05-04. The repository is not archived, but the release list stops well before that push, so if you depend on tagged releases rather than the main branch you may be installing something older than the source tree.

## Podcastfy compared with NotebookLM and with a hosted TTS API

The obvious comparison is NotebookLM, and the README makes it directly. The difference is not output quality, which the README does not claim to beat. The difference is control and reach. NotebookLM gives you a fixed interface over sources you upload there. Podcastfy gives you a function call you can put inside an existing ingestion pipeline, and it accepts inputs NotebookLM does not, including images and arbitrary user-provided topics.

The second comparison is against calling a text-to-speech API directly, such as ElevenLabs. That approach gives you exact control over the script, because you write it. Podcastfy inserts an LLM between your source material and the script, which means you get a conversational structure for free and you give up determinism. If your content is already written as dialogue, or if you need the audio to match a script word for word, the direct API is the better fit and Podcastfy adds a layer you do not want.

A third option is to use the project's own FastAPI container as a service. The README describes it as beta for URLs, so it is closer to a demo of the library than a product. Anyone who needs a stable HTTP contract should wrap generate_podcast themselves rather than depend on that layer.

## Licence, dependencies and the real upgrade cost

Podcastfy is licensed Apache-2.0, and the LICENSE and NOTICE files are both present at the repository root. Apache-2.0 is permissive and includes an explicit patent grant, which matters if you are embedding the package in a commercial product. It does not give you any rights to the model providers you call. Your ElevenLabs, OpenAI, Gemini or Google Cloud terms apply separately to the audio and text those services produce, and nothing in this repository changes that.

The upgrade cost is dominated by the YAML configuration rather than by the Python API. Because config.yaml and conversation_config.yaml ship inside the package and define the conversation behaviour, a version bump can change defaults without any change to your code. Pin the version, and diff those two files when you move.

The dependency surface is wide. LangChain, litellm, the OpenAI SDK, the Google GenAI SDK, google-cloud-texttospeech, ElevenLabs, edge-tts, PyMuPDF and pandas all sit in the same environment. Installing Podcastfy into an existing application means resolving against all of them. The Dockerfile avoids this by building a clean virtual environment at /opt/venv and installing only the package, which is the safer route if your main environment already has opinions about LangChain versions.

## Conclusion

Adopt Podcastfy if you want podcast generation inside your own Python pipeline and are willing to manage LLM and TTS keys yourself. Skip it if you need a polished hosted editor, since the web app and FastAPI layer are described as beta. Before committing, verify which LLM and TTS providers your chosen config points to, and check the conversation_config.yaml defaults in the installed package, because that file decides the voices, the language and the conversation style.

## FAQ

### What is the best free AI podcast generator?

Podcastfy itself is Apache-2.0 licensed and free to install from PyPI, but it is not free to run at scale: every generation calls an LLM and a text-to-speech provider that you configure and pay for separately. The README lists ElevenLabs, edge-tts and Google Cloud Text-to-Speech among the speech options, so the running cost depends on which one you pick.

### Is it too late to start a podcast in 2026?

Podcastfy does not address podcast audience growth or market timing. It is a tool for generating audio conversations from source content, and the README presents it as a programmatic alternative to NotebookLM's podcast feature rather than as advice on starting a show.

### What is the best free software for podcasting?

Podcastfy covers one part of podcasting: turning text, images, PDFs, YouTube videos or a topic into a two-voice audio conversation. It is not a recording, editing or publishing tool, and the README does not document any editing workflow after the audio file is generated.

## Sources

- [License: Apache-2.0](https://github.com/souzatharsis/podcastfy/blob/main/LICENSE)
- [Project website](https://www.podcastfy.ai)
- [README](https://github.com/souzatharsis/podcastfy/blob/main/README.md)
- [Releases](https://github.com/souzatharsis/podcastfy/releases)
- [souzatharsis/podcastfy on GitHub](https://github.com/souzatharsis/podcastfy)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/souzatharsis-podcastfy
