# Doragd/Algorithm-Practice-in-Industry: A Curated Search, Recommendation and Ads Reading Pipeline

> The repository collects industry practice articles from Zhihu, Datafuntalk and technical WeChat accounts, then adds a paper bot that ranks, translates and pushes arXiv and conference papers. It is a reading and aggregation project, not a library you import.

**Doragd/Algorithm-Practice-in-Industry** — 搜索、推荐、广告、用增等工业界实践文章收集（来源：知乎、Datafuntalk、技术公众号）

- Repository: https://github.com/Doragd/Algorithm-Practice-in-Industry
- Stars: 4,602 · Forks: 489
- Language: HTML
- License: BSD-2-Clause
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/doragd-algorithm-practice-in-industry

## What Doragd/Algorithm-Practice-in-Industry actually collects

The README opens with a narrow statement of purpose: the repository gathers industry practice articles about search, recommendation, advertising and user growth, sourced from Zhihu, Datafuntalk and technical WeChat accounts. That is the original scope, and it is still the part most readers will use. The articles themselves are not reproduced. The README says the repository only collects resources and does not quote specific content, and it points to source.xlsx as the source file, which can be sorted with a custom order.

Over time the scope widened. The README lists four additions: a paper push bot for search, recommendation and advertising, conference paper lists, a collection of articles from well-known bloggers, and a series of algorithm walkthroughs. The conference list covers ACL, CIKM, ECIR, EMNLP, ICLR, ICML, KDD, NAACL, NIPS, RecSys, SIGIR, WSDM and WWW, with per-conference, per-year Markdown pages under paperBotV2/conf_summary/data/papers/. The audience is practitioners in that field: the README describes the author as a 2023 graduate who moved from NLP to search, recommendation and advertising and now works on recall. There is no library API here. The deliverable is a reading list plus a scheduled digest.

## How the paper bot ranks, translates and pushes arXiv papers

The mainline code lives in paperBotV2/. The README describes the arXiv flow as fetching new papers in cs.IR and related directions, then using a large model for a coarse ranking, a fine ranking and abstract translation, producing daily JSON and a web page, with optional push to a Feishu group. The web version is published at doragd.github.io/Algorithm-Practice-in-Industry/arxiv_daily.

The conference side works differently. The daily conference push selects recommendation, search and advertising papers from paperBotV2/conf_summary/data/results.json, completes the abstracts, translates them, and sends them through a Feishu bot. Translation for that path uses Caiyun Xiaoyi rather than the large model used on the arXiv path, so the two pipelines do not share a translation backend. Conference Markdown is generated separately, and README updates for conference papers are handled by their own script.

Automation runs through GitHub Actions. The README says workflows trigger arXiv updates, Feishu notifications, conference updates, conference daily pushes and industry practice page deployment, either on a schedule or from an Issue label. The old scripts and the old arXiv workflow were moved to legacy/ and the README states they are kept only for historical reference, compatibility checks and rollback, with paperBotV2/ as the daily entry point.

## Installing the repository and running the arXiv daily job

There is no published package and no install command in the README. You clone the repository and install the dependencies listed in requirements.txt, which pins urllib3 below version 2 alongside feedparser, openai, tenacity, tqdm, beautifulsoup4, aiohttp, requests and jinja2.

```bash
pip install -r requirements.txt
```

After that, the README lists the arXiv main flow as a module invocation. The README does not document which environment variables the flow reads, so check paperBotV2/arxiv_daily/arxiv.py and the workflow files before running it against a live OpenAI key.

```bash
python -m paperBotV2.arxiv_daily.arxiv
```

The Feishu notification step is a separate script, run as a plain file rather than a module. The README gives no sample invocation with arguments, so expect to read the script for the webhook it expects.

```bash
python paperBotV2/arxiv_daily/arxiv_feishu_msg.py
```

The remaining entry points follow the same pattern: update_from_issue.py for conference Issue updates, conf_daily.py for the conference daily push, convert_to_md.py for Markdown generation, update_readme_papers.py for README updates, and industry_practice/maintain.py for the industry practice pages. Contributing a new article means opening an Issue; the README says a GitHub Action then updates the README and source.xlsx, and an Issue template is provided.

## Where the pipeline breaks or does not fit

The arXiv path depends on a large model for ranking and translation, so every daily run costs tokens and fails when the API is unavailable or rate limited. The tenacity dependency suggests retries are handled somewhere in the code, but the README does not describe backoff behaviour, partial failure handling or what happens to a day's JSON when the model call fails halfway through. That is a real gap for anyone planning to rely on the digest.

The two push paths also use different translation services, one a large model and one Caiyun Xiaoyi. If either key expires, only part of the digest degrades, which is easy to miss unless you watch the output.

This is the wrong tool if you want a searchable corpus with full text. The repository collects links and metadata, and the README explicitly says it does not quote the articles, so you still read the originals on Zhihu or a WeChat account. It is also the wrong tool outside the Chinese-language search, recommendation and advertising community. The conference lists cover international venues, but the practice articles, the blogger collection and the translated abstracts are aimed at Chinese-speaking practitioners.

## Compared with running your own arXiv digest with arxiv-sanity or a feed reader

A feed reader or an arXiv alert gives you everything in a category, unfiltered, in English, with no ranking and no translation. That approach costs nothing to run and never breaks on an API key, but it puts the filtering work on you every morning.

This project takes the opposite position: it spends model calls to cut the list down before you see it, then translates the abstracts so a commute-sized read is possible. The README is candid about the trade: it says the papers are not very useful but are worth skimming to widen your thinking, and suggests reading the push during a morning commute since the abstracts are already translated. A general-purpose summarizer would give you the same translation, but it would not maintain the conference paper lists by conference and year, and it would not accept new industry articles through an Issue template that rewrites the README and spreadsheet automatically. The value is in the maintained index, not in the summarization technique.

## Licence and the cost of keeping this running

The repository is under BSD-2-Clause, which permits reuse and modification with the copyright notice and disclaimer retained. That covers the code and the repository's own files. It does not cover the linked articles, which remain with their original authors, and the README asks that anyone with a copyright concern make contact so the link can be removed. If you fork this and republish the digest, the licence question about the collected links is separate from the licence on the scripts.

The maintenance cost is mostly external. The repository's last push was on 2026-09-22, so it is current. Keeping a fork alive means keeping an OpenAI key, a Caiyun Xiaoyi key and a Feishu webhook working, plus watching upstream changes to the arXiv and conference pipelines. The conference lists need periodic additions as new years are published, and the Markdown generation and README update scripts must be re-run for those to appear. The legacy/ directory exists for rollback, but the README does not document a rollback procedure beyond keeping those files.

## Conclusion

Adopt it if you work on search, recommendation or advertising and want a maintained index of industry write-ups plus a daily arXiv and conference digest with Chinese summaries. Do not adopt it if you need a library to call from your own code, or if you cannot run GitHub Actions with an OpenAI key and a Feishu webhook, because the automation is the product. Before committing, open paperBotV2/arxiv_daily/arxiv.py and requirements.txt to confirm the dependencies match your environment, and check the GitHub Actions workflow files under .github/workflows/ to see exactly which secrets the scheduled jobs expect.

## FAQ

### Is Doragd/Algorithm-Practice-in-Industry a Python library I can import?

No. The repository is a collection of curated industry articles plus a set of scripts under paperBotV2/ that fetch, rank, translate and push papers through GitHub Actions.

### How do I install Doragd/Algorithm-Practice-in-Industry?

There is no published package. Clone the repository and run pip install -r requirements.txt, which installs feedparser, openai, tenacity, tqdm, beautifulsoup4, aiohttp, requests, urllib3 below version 2 and jinja2.

### How do I add a new industry practice article to Doragd/Algorithm-Practice-in-Industry?

Submit an Issue. The README states that a GitHub Action then updates the README and the source.xlsx content, and an Issue template is provided for this.

### Where can I read the daily arXiv papers without running the code?

The README links a web version of the arXiv daily digest at doragd.github.io/Algorithm-Practice-in-Industry/arxiv_daily, generated by the same pipeline.

### What licence does Doragd/Algorithm-Practice-in-Industry use?

The repository is under BSD-2-Clause. The linked articles themselves remain with their original authors, and the README asks that copyright concerns be raised so the link can be removed.

## Sources

- [Doragd/Algorithm-Practice-in-Industry on GitHub](https://github.com/Doragd/Algorithm-Practice-in-Industry)
- [Issues](https://github.com/Doragd/Algorithm-Practice-in-Industry/issues)
- [License: BSD-2-Clause](https://github.com/Doragd/Algorithm-Practice-in-Industry/blob/main/LICENSE)
- [README](https://github.com/Doragd/Algorithm-Practice-in-Industry/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/doragd-algorithm-practice-in-industry
