# Paper2Video and PaperTalker: turning a LaTeX paper into a narrated presentation video

> Paper2Video from Show Lab, NUS ships two things: PaperTalker, an agent pipeline that renders slides, subtitles, speech, cursor motion and an optional talking head, and a benchmark for scoring the result. It expects LaTeX sources, API keys and a 48 GB GPU.

**showlab/Paper2Video** — Automatic Video Generation from Scientific Papers

- Repository: https://github.com/showlab/Paper2Video
- Website: https://showlab.github.io/Paper2Video/
- Stars: 2,381 · Forks: 326
- Language: Python
- License: MIT
- Published: 2026-09-28 · Updated: 2026-09-28 · Language: en
- Canonical page: https://hysenlabs.com/projects/showlab-paper2video

## The paper-to-talk-video problem Paper2Video attacks

Recording a conference talk is mostly mechanical work. Someone builds slides from the paper, writes a script, records narration, moves a cursor to the figure being discussed, and optionally films a presenter. Paper2Video splits that chain into named stages and automates each one. The README frames it as two problems: on the left, how to create a presentation video from a paper, which is the PaperTalker agent; on the right, how to evaluate a presentation video, which is the Paper2Video benchmark with its own metrics.

The intended user is a researcher or engineer who already has a paper in LaTeX form and wants a video draft without a studio session. The input contract is explicit: a paper, an image, and an audio sample. The README's example pairs an arXiv paper on character-level convolutional networks with a photo of Hinton and a reference audio clip. The output is a presentation video.

That three-part input is the first real constraint. The image and the audio are not optional decoration in the full pipeline: they feed the talking-head stage, which needs a face and a voice sample to work from. If you have neither, the project's own answer is the light pipeline, which the 2025-10-15 update describes as a version without talking-head for fast generation.

## Inside the PaperTalker pipeline: slides, subtitles, speech, cursor, talking head

The README describes pipeline.py as an automated pipeline that takes LaTeX paper sources plus a reference image and audio, then runs through sub-modules in a stated order: Slides, Subtitles, Speech, Cursor, Talking Head. Each stage consumes the previous stage's artifact, and the result directory is where those intermediate artifacts land, since the argument table describes --result_dir as holding slides, subtitles, videos and similar outputs.

Two model slots drive the language and vision work. --model_name_t is the LLM and --model_name_v is the VLM, both defaulting to gpt-4.1 in the table. The README's stated best practice is GPT4.1 or Gemini2.5-Pro for both roles, and it notes that locally deployed open-source models such as Qwen are supported through the Paper2Poster project.

The cursor stage is the part that distinguishes this from a slideshow generator. Cursor grounding means the video points at the region of the slide under discussion, which is why the pipeline needs the rendered slide and the narration text together rather than either alone. The talking-head stage is last and is the only stage with a separate environment, because it delegates to Hallo2.

One design choice worth naming: the default --model_name_talking value is hallo2, and the argument table says hallo2 is currently the only supported talking-head model. That is a narrow integration surface, and it is the main reason the setup instructions split into two conda environments.

## Installing Paper2Video and running a first fast generation

Installation happens inside the src directory. The README creates a conda environment named p2v on Python 3.10, installs the requirements file, and adds tectonic through conda-forge. Tectonic is the LaTeX engine, which is consistent with the pipeline taking LaTeX sources rather than a compiled PDF.

```bash
cd src
conda create -n p2v python=3.10
conda activate p2v
pip install -r requirements.txt
conda install -c conda-forge tectonic
```

Before inference you export API credentials. The README shows two variables, and the pipeline reads whichever provider your chosen model names point at.

```bash
export GEMINI_API_KEY="your_gemini_key_here"
export OPENAI_API_KEY="your_openai_key_here"
```

If you skip the human presenter, the README explicitly says you can skip the Hallo2 section and go straight to configuring LLMs. The light pipeline then runs with a single command. Note that the GPU list is passed as a bracketed list, exactly as written in the README.

```bash
python pipeline_light.py \
    --model_name_t gpt-4.1 \
    --model_name_v gpt-4.1 \
    --result_dir /path/to/output \
    --paper_latex_root /path/to/latex_proj \
    --ref_img /path/to/ref_img.png \
    --ref_audio /path/to/ref_audio.wav \
    --gpu_list [0,1,2,3,4,5,6,7]
```

What you should see is a populated result directory containing the generated slides, subtitles and video artifacts for your paper. The README does not publish an expected runtime, so treat the first run as a smoke test rather than a benchmark.

## Adding the talking head means a second environment and a 48 GB GPU

The full pipeline adds --model_name_talking hallo2 and --talking_head_env, pointing at the Python environment where Hallo2 lives. The README is direct about why this is separate: you need to prepare the environment separately for talking-head generation to avoid package conflicts. That is a real operational cost. Two conda environments, two requirements installs, and a path that has to stay correct between them.

```bash
cd hallo2
conda create -n hallo python=3.10
conda activate hallo
pip install -r requirements.txt
```

Hallo2 itself is cloned from the fudan-generative-vision repository, and the README says to follow that project's instructions to download the model weights. After installing, it suggests running which python to capture the environment path that --talking_head_env expects.

The hardware line is the other hard constraint. The README states that the minimum recommended GPU for running this pipeline is an NVIDIA A6000 with 48G. That is a minimum recommendation for the pipeline as a whole, and the example commands pass an eight-element GPU list, which tells you the authors expect multi-GPU machines. A single consumer card is not the target configuration.

The argument table also lists --ref_img as a reference image with a note that it must be something the README truncates before finishing. Because that requirement is cut off in the README text, check the source file before assuming any portrait will do.

## Where Paper2Video is the wrong tool

The clearest failure mode is input format. The pipeline takes --paper_latex_root, the root directory of a LaTeX paper project, and the overview describes the input as LaTeX paper sources. If your paper exists only as a compiled PDF, or as a Word document from a journal template, there is no documented path in. Converting a PDF back to LaTeX is not part of this project.

Cost and dependency are the second limit. The README recommends GPT4.1 or Gemini2.5-Pro for both the LLM and the VLM roles, so a run spends API credits on text and vision calls, and the full variant adds Hallo2 weights on top. Nothing in the README describes an offline mode that avoids hosted APIs entirely, only that locally deployed open-source models are supported through Paper2Poster.

The environment split is a third friction point, and it is not cosmetic. The README's own warning about package conflicts means a working light setup does not imply a working talking-head setup. Budget the Hallo2 install as a separate task with its own debugging.

Finally, the evaluation half is a benchmark, not a quality guarantee. Paper2Video is described as a benchmark with well-designed metrics to evaluate presentation quality. A benchmark score tells you how a video compares against those metrics; it does not tell you whether the narration is accurate about your specific results. Read the generated subtitles before publishing anything.

## How Paper2Video differs from Paper2Poster and Code2Video

Paper2Poster is the closest relative, and the README treats it as such: it is the route for running locally deployed open-source models instead of hosted APIs. The difference in approach is the output artifact. Paper2Poster targets a poster, a single static layout, while PaperTalker targets a timed video with narration, cursor movement and an optional presenter. A poster has no timeline, so none of the subtitle, speech or cursor stages have an analogue there.

Code2Video is a different framing of the same broad idea, and the search data around this project pairs the two names. The distinction that matters here is the input contract: Paper2Video's pipeline is built around LaTeX paper sources and an academic presentation structure, not around source code repositories. If your material is a codebase rather than a paper, the slide and subtitle stages here have nothing to consume.

The talking-head comparison is internal rather than external. Paper2Video offers two modes: pipeline_light.py without a presenter, and pipeline.py with Hallo2. The light mode exists precisely because the talking-head stage is the expensive, conflict-prone part. Choosing between them is the main architectural decision a user makes, and the README's 2025-10-15 note frames the light version as the fast path.

## Maintenance, licence and what an upgrade costs

The repository is not archived, and the last push was on 2026-03-05. The README's update log runs from 2025-09-28, when the work was accepted to the Scaling Environments for Agents Workshop at NeurIPS 2025, through 2025-10-15, when the no-talking-head version was added. There are no retrieved releases, so installation tracks the main branch rather than a tagged version. That matters for reproducibility: pinning to a commit is the only way to get the same code twice.

Upgrade cost concentrates in the two environments. A change to requirements.txt in src and a change to Hallo2's requirements are separate maintenance events, and the README's instruction to install Hallo2 separately means an upstream Hallo2 change can break the full pipeline without touching this repository. The --model_name_talking argument currently accepts hallo2 only, so there is no fallback model to switch to when that integration breaks.

The licence is MIT, which permits commercial and private use and modification, subject to the usual condition that the copyright notice and permission notice travel with copies. That covers this repository. It does not automatically cover Hallo2's weights or the hosted model APIs you call, which carry their own terms, and the README does not discuss those. Check the Hallo2 repository and your API provider's terms separately; this is a description of the stated licence, not legal advice.

The project also invites contributions, which is worth weighing if you plan to depend on it: an open contribution model with no tagged releases means you should expect to read commits rather than changelogs.

## Conclusion

Adopt Paper2Video if you already keep a paper as a LaTeX project and want a draft talk video without recording yourself; the light pipeline is the cheaper entry point, and the README's own command shows how to start it. Do not adopt it if you only have a PDF, since the pipeline takes a LaTeX root, or if you have no 48 GB GPU class card, since the minimum recommendation is an NVIDIA A6000. Before committing, verify three things: that your paper compiles under tectonic, that your reference image and audio meet the pipeline's expectations, and that Hallo2 installs cleanly in its own conda environment, because the README warns about package conflicts between the two.

## FAQ

### What input does Paper2Video need to generate a presentation video?

The README states the input is a paper plus an image plus an audio clip. In practice the pipeline takes a LaTeX paper root through --paper_latex_root, a reference image through --ref_img and reference audio through --ref_audio.

### Can I run Paper2Video without generating a talking head?

Yes. The README marks the Hallo2 setup as optional if you do not need a human presenter, and the 2025-10-15 update added a version without talking-head for fast generation, run through pipeline_light.py.

### What GPU does Paper2Video require?

The README states that the minimum recommended GPU for running this pipeline is an NVIDIA A6000 with 48G. The example commands also pass an eight-element --gpu_list, so the authors expect multi-GPU machines.

### Which talking-head model does Paper2Video support?

The argument table lists --model_name_talking with a default of hallo2 and notes that hallo2 is currently the only supported talking-head model. Hallo2 is cloned from the fudan-generative-vision repository and installed in its own conda environment.

### What licence does Paper2Video use?

The repository is licensed under MIT. That covers this codebase, but the README does not address the licence terms of the Hallo2 model weights or of the GPT4.1 and Gemini2.5-Pro APIs the pipeline can call.

## Sources

- [Issues](https://github.com/showlab/Paper2Video/issues)
- [License: MIT](https://github.com/showlab/Paper2Video/blob/main/LICENSE)
- [Project website](https://showlab.github.io/Paper2Video/)
- [README](https://github.com/showlab/Paper2Video/blob/main/README.md)
- [showlab/Paper2Video on GitHub](https://github.com/showlab/Paper2Video)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/showlab-paper2video
