# hacker-news-digest: An AI Pipeline for Summarizing Hacker News

> hacker-news-digest is a self-hosted Python tool that fetches Hacker News articles, extracts their main content using a machine learning algorithm, and deploys LLM-generated summaries as a static site. It suits developers who want to scan HN quickly without relying on a third-party service.

**polyrabbit/hacker-news-digest** — :newspaper: Let ChatGPT Summarize Hacker News for You

- Repository: https://github.com/polyrabbit/hacker-news-digest
- Website: http://hackernews.betacat.io/
- Stars: 755 · Forks: 94
- Language: Python
- License: LGPL-3.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/polyrabbit-hacker-news-digest

## What hacker-news-digest Does and Who Uses It

Every day, Hacker News surfaces dozens of articles that each take time to read in full. hacker-news-digest solves the triage problem: it fetches the front page, extracts the main content from each linked article, passes that content to a language model, and publishes the resulting summaries as a static site. The site is already hosted at hackernews.betacat.io for anyone who does not want to run the software themselves.

The project targets two groups. The first is developers who read HN daily and want to decide which articles deserve a full read without opening each one. The second is engineers who prefer to self-host their reading infrastructure and avoid routing their habits through a third-party summary service. Because the project is licensed under LGPL-3.0, anyone can deploy it commercially as long as modifications to the library code itself are released under the same terms.

The project also supports Chinese translation, added as a single extra prompt step. The translated version is available at hackernews.betacat.io/zh.html.

## The Five-Step Pipeline from Parsing to Deployment

The full pipeline runs in five steps, each handled by a separate part of the codebase:

1. The project parses the Hacker News front page to collect article URLs, point counts, submission times, and comment links.
2. A machine learning score algorithm, implemented as a Jupyter notebook included in the repository (`[tutorial] How-to-extract-main-content-from-web-pages-using-Machine-Learning.ipynb`), scores each HTML text block by relevance and removes boilerplate to produce a clean extract.
3. A suitable illustration is found for each article and cached locally.
4. The clean extract is sent to a language model to generate a summary. If the remote LLM service is unavailable, the project falls back to a locally running Qwen model via llama-cpp-python.
5. A Jinja-style template combines the summary, illustration, and metadata (points, submitter, submission time, comment count, and source domain) and the output is deployed to GitHub Pages.

The README documents a set of emoji icons that represent each metadata field in the rendered page: a heart for point count, a person icon for the submitter, a clock for submission time, a chat bubble for comment count, a link for the article source, and a newspaper icon for the summary model. These icons give the page a consistent visual shorthand for scanning metadata at a glance.

The README describes two deployment targets in the Makefile: `gh_daily_page` generates per-day archives, and `gh_home_page` rebuilds the main index, copies static assets, creates a symlink for backward compatibility, and runs HTML, CSS, and JavaScript minification through the `minify` tool. Articles can be sorted by points, comment count, or publication time, and the topN filter limits the output to the highest-scoring articles above a point threshold.

## Installing and Running the Project Locally

The project requires Python, a PostgreSQL database, and either an OpenAI API key or a locally downloaded Qwen model. After cloning the repository, the first step is to create the database schema:

```bash
make initdb
```

This runs `python -c 'from db import init_db; init_db()'` to create the tables. Once the schema is in place, the Flask development server starts with:

```bash
make run
```

For a production-grade deployment, the project can run under Gunicorn inside Docker:

```bash
make run-in-docker
```

Dependencies are pinned in requirements.txt. The key packages are Flask 2.3.3, SQLAlchemy 2.0.21, psycopg2-binary 2.9.9, openai 0.28.1, llama-cpp-python 0.3.34, Pillow 10.2.0, and huggingface-hub 0.36.2. The `setcron` Makefile target starts a background loop that polls the local server every ten minutes to trigger content updates, using an optional `HN_UPDATE_KEY` environment variable for authenticated requests.

## The OpenAI SDK Version Pin and the Local Qwen Fallback

requirements.txt pins the OpenAI Python client at version 0.28.1. This version predates the OpenAI SDK v1 rewrite that changed the client's class structure and method signatures. Code written against 0.28.1 will not run against a v1 or v2 client without modification. Developers who already have a newer version of the OpenAI SDK installed should treat this as a migration task before running the project.

The local fallback relies on llama-cpp-python 0.3.34, which compiles a C++ binary during installation. Build times can be significant on slower machines. The huggingface-hub dependency (0.36.2) handles downloading the Qwen model weights from Hugging Face. The README does not document how large the downloaded model is or how much memory it requires at inference time, so capacity planning requires consulting the Qwen model card separately.

For token accounting, the project includes tiktoken 0.5.2, which counts tokens before sending requests to the remote API.

## Four Limitations That Affect Real Deployments

The README's TODO list documents comment summarization as a feature that has not yet been implemented. The project summarizes the linked article, not the HN discussion thread. If the discussion is the main reason someone reads a story, hacker-news-digest does not address that use case.

The ML content extractor is based on a scoring algorithm for visible HTML text. Pages that render their content entirely through JavaScript at runtime, or that sit behind a paywall requiring login, will yield incomplete or empty extracts. The README acknowledges this limitation and lists PhantomJS or Selenium integration as a potential future improvement.

Only Chinese is listed as a supported non-English locale. Developers who want summaries in other languages would need to modify the prompt themselves, as that path is not documented.

The openai==0.28.1 pin is also a maintenance cost. Each time dependencies are updated, this version constraint will conflict with newer packages that have dropped backward compatibility with the old client, requiring a project-specific migration.

## Comparing to the Official HN Front Page and RSS Readers

The closest alternative for consuming Hacker News without summaries is the official front page at news.ycombinator.com, which shows titles and point counts but requires opening each link to assess the content. The Algolia-powered HN search at hn.algolia.com adds full-text search and filters by date, type, and author, but generates no summaries.

General RSS readers such as Feedly or Inoreader can subscribe to hacker-news-digest's feed.xml output once it is deployed, but those services summarize nothing on their own. The RSS feed from hacker-news-digest is listed as a completed feature in the README and is generated as part of the `gh_home_page` Makefile target.

The main difference between the self-hosted version and the already-running instance at hackernews.betacat.io is control. The hosted version requires no setup but cannot be customized. The self-hosted version allows changing the number of articles retained, the summary prompt, the sorted order, and which language the translation produces, but it requires maintaining the database and the LLM dependency.

## Conclusion

hacker-news-digest fits developers who already manage a Python-plus-PostgreSQL stack and want full control over which articles get summarized and how. It is a poor fit for anyone whose primary interest is reading HN discussion threads, since comment summarization is listed as a TODO item and has not been implemented. Before deploying, confirm that your OpenAI API key is compatible with the openai==0.28.1 SDK pinned in requirements.txt, or that you have the disk space and RAM available to run a local Qwen model via llama-cpp-python.

## FAQ

### How does hacker-news-digest extract the main content from an article?

The project uses a machine learning score algorithm implemented as a Jupyter notebook included in the repository. The algorithm scores each HTML text block by its likely relevance and strips boilerplate such as navigation and ads to produce a clean text extract.

### Does hacker-news-digest support an RSS feed?

Yes. The README lists RSS as a completed feature. The Makefile includes RSS generation as part of the gh_home_page target, and the output is written to feed.xml.

### Can hacker-news-digest run without an OpenAI API key?

Yes. When the remote LLM service is unavailable, the project falls back to a locally running Qwen model using llama-cpp-python. However, the README does not document the model size or memory requirements for the local fallback.

## Sources

- [Issues](https://github.com/polyrabbit/hacker-news-digest/issues)
- [License: LGPL-3.0](https://github.com/polyrabbit/hacker-news-digest/blob/master/LICENSE)
- [polyrabbit/hacker-news-digest on GitHub](https://github.com/polyrabbit/hacker-news-digest)
- [Project website](http://hackernews.betacat.io/)
- [README](https://github.com/polyrabbit/hacker-news-digest/blob/master/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/polyrabbit-hacker-news-digest
