# youtube-transcript-api: pull YouTube subtitles in Python without an API key

> The library fetches human and auto-generated captions for a video ID and returns timestamped snippets. It is MIT-licensed, installs from PyPI, and its main weakness is that YouTube can block the request.

**jdepoix/youtube-transcript-api** — This is a python API which allows you to get the transcript/subtitles for a given YouTube video. It also works for automatically generated subtitles and it does not require an API key nor a headless browser, like other selenium based solutions do!

- Repository: https://github.com/jdepoix/youtube-transcript-api
- Stars: 8,399 · Forks: 841
- Language: Python
- License: MIT
- Published: 2026-09-22 · Updated: 2026-09-22 · Language: en
- Canonical page: https://hysenlabs.com/projects/jdepoix-youtube-transcript-api

## What youtube-transcript-api actually solves

YouTube does not offer a public caption endpoint through the Data API that returns the transcript text for an arbitrary video. The captions a viewer sees in the player come from a separate track that the page requests. youtube-transcript-api targets that track directly. You give it a video ID, it returns the caption lines with start times and durations, and it works for both human-written subtitles and automatically generated ones. The README states the library "does not require a headless browser, like other selenium based solutions do", which is the design claim that matters: no Chromium process, no page render, no browser driver to keep updated.

The audience is narrow but real. Anyone building a search index over video, a summarizer, a subtitle QA tool, or a dataset pipeline needs caption text in a structured form. The library returns a FetchedTranscript object whose snippets carry text, start and duration, plus video_id, language, language_code and is_generated. That last field lets you separate machine captions from human ones, which matters for any downstream quality threshold. If you only ever need one transcript by hand, a browser extension is less work. This is for code that does it at some volume.

## How the fetch and translation path is structured

The API surface is small. You instantiate YouTubeTranscriptApi and call fetch on it with a video ID. The README is explicit that you pass the ID, not the URL: for https://www.youtube.com/watch?v=12345 the ID is 12345. Getting that wrong is the most common first mistake.

Language selection is a priority list, not a filter. Passing languages=['de', 'en'] means the library tries German first and falls back to English. The README notes the default is English, so a video with only Spanish captions will fail unless you ask for it. The companion list() call reports which languages exist for a video before you fetch, which is the correct order of operations when you do not control the input set.

The returned object behaves like a sequence. The README gives three examples: iterating over it yields snippets, indexing with fetched_transcript[-1] returns the last one, and len() gives the snippet count. When you want plain data, to_raw_data() returns a list of dictionaries with text, start and duration keys. That is the shape you would write to JSONL for a pipeline. Translation is also supported according to the README description, though the excerpt here shows the fetch and list paths in more detail than the translation path.

## Install and first fetch: pip, video ID, output

The README recommends installing from PyPI with pip. There is no API key to configure and no environment variable to set for the basic path.

```bash
pip install youtube-transcript-api
```

After that, the smallest working program constructs the API object and fetches a transcript. Note that the ID is the eleven-character video identifier, not the full watch URL.

```python
from youtube_transcript_api import YouTubeTranscriptApi

ytt_api = YouTubeTranscriptApi()
fetched_transcript = ytt_api.fetch("12345")

for snippet in fetched_transcript:
    print(snippet.start, snippet.text)
```

You should see lines pairing a floating point start time with the caption text. If you would rather not write Python, the package also installs a console script named youtube_transcript_api, declared in pyproject.toml as pointing at youtube_transcript_api.__main__:main, so the same fetch is available from a shell.

To get the data out in a serializable form, call to_raw_data() on the fetched transcript. The result is a list of dictionaries with text, start and duration, which is what you would hand to json.dump or append to a JSONL file.

```python
raw = fetched_transcript.to_raw_data()
print(raw[0])
```

The README shows that first element looking like {'text': 'Hey there', 'start': 0.0, 'duration': 1.54}. If your video is not in English, add the languages argument before you assume the fetch failed.

## The failure mode the README does not solve for you

This is an unofficial client. It reads a caption track that YouTube serves to the web player, and YouTube is under no obligation to keep that path open or stable. The practical consequence is that a pipeline built on it can start returning errors without any change on your side. The related search phrases people use around this project include "youtube transcript api not working" and "youtube transcript api ip blocked", which reflects a real operational pattern: request volume from a single address is the usual trigger, and shared or datacenter IP ranges are the usual victims.

There is no documented retry policy, no documented backoff, and no documented proxy configuration in the excerpt available here. If you are fetching thousands of videos, you are responsible for pacing the requests yourself, and you should treat a failure as a per-video event rather than aborting a batch. The library also cannot help when a video simply has no caption track. Some videos have captions disabled by the uploader, some are music or silent content, and some are region-restricted. list() tells you what is available before you fetch, which is why calling it first is worth the extra request.

Finally, the default English behavior is a silent trap. A batch job that never passes languages will quietly skip every non-English video rather than error in a way that is obvious from the aggregate count.

## Where a hosted transcript service differs

The README points to sponsored hosted alternatives, including SerpApi, TranscriptAPI.com, supadata and Dumpling AI. The difference in approach is not just price. A hosted service runs the fetching from its own infrastructure and exposes an HTTP endpoint, so the IP reputation problem moves off your machines. You get a stable contract, a status page and someone to email when YouTube changes something. You also get a per-request cost and a dependency on a third party seeing every video ID you process.

A second alternative is a Selenium or Playwright based scraper. That approach drives a real browser, which means it can survive some changes that break a direct caption request, but it costs a browser process per worker, is slower, and needs its own maintenance for driver and browser versions. The library's stated advantage is precisely the absence of that layer.

A third option is the official YouTube Data API. It is a documented, keyed API, but it is not a drop-in replacement for this task, because the caption download it exposes is restricted to videos you own or have permission for. If your use case is your own channel, that is the correct tool. If your use case is other people's videos, it is not.

## Maintenance, licence and upgrade cost

The repository is not archived, and the last push was on 2026-09-10. The most recent release listed is v1.2.4 on 2026-01-29, following v1.2.3 on 2025-10-13 and v1.2.2 on 2025-08-04. That is a release cadence measured in months, with commits landing between releases. Nothing in the README describes a deprecation policy or a support window, so pinning a version is the only predictable option.

The licence is MIT, declared both in pyproject.toml and in the LICENSE file at the repository root, and the PyPI classifier is "License :: OSI Approved :: MIT License". MIT is permissive: you can use it commercially and in closed-source products, provided you keep the copyright notice and permission notice. That covers the library itself. It does not cover the caption text you retrieve, which is YouTube's and the uploader's content, and the README says nothing about what you may do with the output. That is a separate question and not one the licence answers.

Upgrade cost is low on the surface. The API is a handful of methods and the object model is stable across recent releases. The real cost is the test story: pyproject.toml defines a coverage task with --fail-under=100, so the project holds itself to full coverage. That is a good signal for contributors and a mild warning for consumers, because it means the maintainers care about behavior changes being caught. Read the release notes before bumping, since a change in how fetch handles missing languages would show up as a behavior difference rather than a signature change.

## Conclusion

Adopt youtube-transcript-api if you need timestamped captions for a known video ID inside a Python process and you accept that YouTube can refuse the request from a given IP. Do not adopt it if you need an official, contract-backed data source, or if you need transcripts from a platform other than YouTube; the library only speaks to YouTube's caption endpoints. Before committing, verify three things: that your target videos have captions at all, that your network path is not already blocked, and which language codes you actually need, since fetch() defaults to English. Also check whether the 100 percent coverage gate in pyproject.toml matches the code you are pinning to, because a version you cannot run the test suite against is a version you cannot patch quickly.

## FAQ

### Is youtube-transcript-api free to use?

Yes. The package is published on PyPI and licensed under MIT, and the README states it does not require an API key. You pay nothing to the project, though the README points to sponsored hosted alternatives if you would rather not run the fetching yourself.

### How do I install youtube-transcript-api?

The README recommends installing the module with pip, using the command pip install youtube-transcript-api. That installs both the importable Python module and the console script named youtube_transcript_api.

### How do I use youtube-transcript-api to get a transcript?

Instantiate YouTubeTranscriptApi and call fetch with the video ID, not the full URL. The returned FetchedTranscript is iterable, indexable and has a length, and to_raw_data() converts it to a list of dictionaries with text, start and duration.

### What is youtube-transcript-api?

It is a Python API that retrieves the transcript or subtitles for a given YouTube video. It works with automatically generated subtitles, supports translating subtitles, and does not need a headless browser.

### Is there a free way to get a YouTube transcript?

Yes. The README states that youtube-transcript-api requires no API key and no headless browser, so the only cost of the fetch path it describes is running the Python process yourself.

### How do I get youtube-transcript-api running on my machine?

Install it from PyPI with pip install youtube-transcript-api, then either import YouTubeTranscriptApi in a Python script or call the console script youtube_transcript_api that the package installs. No key or browser setup is involved.

## Sources

- [Issues](https://github.com/jdepoix/youtube-transcript-api/issues)
- [jdepoix/youtube-transcript-api on GitHub](https://github.com/jdepoix/youtube-transcript-api)
- [License: MIT](https://github.com/jdepoix/youtube-transcript-api/blob/master/LICENSE)
- [README](https://github.com/jdepoix/youtube-transcript-api/blob/master/README.md)
- [Releases](https://github.com/jdepoix/youtube-transcript-api/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jdepoix-youtube-transcript-api
