Open-source project
Peldom/papers_for_protein_design_using_DL avatar
Peldom/papers_for_protein_design_using_DL

A protein design paper list whose taxonomy explanation is commented out of its own page

List of papers about Proteins Design using Deep Learning

1,980 stars220 forksUnknownGPL-3.0

At a glance

What is it?
Peldom/papers_for_protein_design_using_DL is a curated reading list for deep learning in protein design, organised into sixty-six subsections across eight groups. The note explaining why the groups divide where they do sits in the source inside an HTML comment, so nobody reading the rendered page sees it.
Who is it for?
This suits a researcher entering the field who needs a map of which architecture families have been applied to which design task, rather than a search index, because the value here is the cross-tabulation. Two things to know before you rely on it.
Can I use it commercially?
Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
Is it still maintained?
Yes. The repository last received commits 51 days ago.
What is it written in?
GitHub does not report a main language for this repository.

Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Six files at the root, and one of them is the entire project

The repository is a readme, a licence, a contributing guide, a code of conduct, a git ignore file and one image that serves as the cover. There is no code, no script that fetches citations, no bibliography file and no release, so the project's language is recorded as unknown and the only way to consume it is to read the page. That is a legitimate shape for a curated list, and it has one advantage worth naming: a link rot checker or a dependency bot has nothing to break. It also means there is no way to ask what changed, since a list that lives in one file has no history a reader can browse without opening the commit log.

Sixty-six subsections across eight groups

The index is a cross-tabulation of two things: what you are trying to do, and what kind of model does it. The first group is not design at all but the measuring apparatus, five subsections covering sequence datasets and benchmarks, structure datasets and benchmarks, public databases, similar lists and guides. The second is reviews and surveys, five again, split by the object being designed. Then come five design groups and a catch-all. Counting the entries in the index gives five, five, five, eight, fourteen, fourteen, ten and five, which is sixty-six subsections in total. The two largest are the ones that map a target onto a sequence, in both directions, and they are also the two with the longest tails of architectures.

The note explaining the taxonomy is commented out

Directly above the index there is a paragraph, wrapped in an HTML comment, that says which grouping paradigm each section follows. It states that one section follows a generator, predictor and optimizer paradigm; that three sections follow the inside-out paradigm of function, scaffold and sequence taken from the RosettaCommons naming; and that two follow other machine learning strategies. That paragraph is the only place the reader learns why the groups divide where they do, and it does not render. Two other comment blocks sit in the same region: one listing the design categories the repository intends to include and pointing at a second, larger de novo list, and another pointing at two non-English columns where the maintainers' notes on these papers live. All three are in the source and none is visible to a reader.

One bucket ends in a technique and another in a Boltzmann machine

The index looks like a clean two-axis taxonomy and mostly is one, with two exceptions that show it is a filing convention rather than a taxonomy. The largest scaffold-to-sequence group has fourteen entries, thirteen of which are architecture families, and the fourteenth is a method of training. So that list mixes what to build with how to fit it. The function-to-sequence group, also fourteen entries, includes a Boltzmann machine among its neural architectures. Neither is wrong, and both are useful, but a reader counting architectures in a bucket will overcount by one in the first case and be surprised in the second.

The weekly section is two papers, and one of them ships code

Above the index sits a running additions section with a date stamp on it, the most recent being August 2026, and two entries. One is a review in a chemistry-biology journal on generative methods for enzyme design. The other is a preprint of what the authors call a proteome-scale atlas of designed binder candidates, and its entry carries two extra links beyond the citation, one to a code repository and one to a hosted site. That convention is the useful signal in the section: an entry with more than one link usually has something you can run or look at, while a bare citation usually does not.

The credit chain runs through three repositories and two columns

The header credits a specific person and a specific repository as the inspiration, naming it as the source of the idea for a curated machine learning list. The visible text then invites contributions through a contributing document and an issue tracker, and links a set of community values, guiding principles and commitments for responsible development in this field, hosted on its own site. That last link is unusual on a paper list and says something about how the maintainers see the field. The one place the list is described as narrower than a general protein-design reading list is itself inside a comment.

A software licence on a repository with no software

The repository is licensed under the GNU General Public License at version three, the same terms you would put on a codebase, on a repository whose entire content is a Markdown file. That is not wrong and it is arguably the right choice, since it prevents someone vendoring the list into a closed product and relicensing the curation away. It does mean the terms describe something that is not code, and the page carries no statement about the papers themselves, which remain under their publishers' licences whatever the repository says. The cover image is the only other asset, and the rest of the root is the two policy files any community project would have.

Editorial conclusion

This suits a researcher entering the field who needs a map of which architecture families have been applied to which design task, rather than a search index, because the value here is the cross-tabulation. Two things to know before you rely on it. The taxonomy is not documented on the page that uses it, so the grouping has to be inferred from the section names. And the list is maintained as one large readme with no releases, no per-paper metadata beyond a citation and an occasional code link, so it will not tell you which entries have been superseded.

Frequently asked questions

What is Peldom/papers_for_protein_design_using_DL?

It is a curated list of papers on protein design using deep learning, created to provide a focused platform for that field rather than a general reading list. It adds a running weekly section of new papers with links to code and hosted sites where those exist, and it credits a specific earlier repository as its inspiration.

How is the paper list organised?

Into eight top-level groups cross-tabulating design task against model family. The first group is benchmarks, datasets, public databases, similar lists and guides; the second is reviews and surveys split by the object being designed; five groups then cover design tasks from scaffold generation through to structure prediction; and the last is a catch-all for mutation effects, representation learning, molecular design, frameworks and unclassified entries. Counting the index gives sixty-six subsections.

What files does the papers_for_protein_design_using_DL repository contain?

Six entries: the readme that is the whole project, a licence, a contributing guide, a code of conduct, a git ignore file and a cover image. There is no code, no releases and no bibliography file, so the list is consumed by reading the page and its history.

Where are the notes on these papers?

The maintainers' notes are published in two non-English columns, one bilingual, and both are pointed at from inside an HTML comment in the readme rather than from the rendered page. The same comment also points at a second, larger list of de novo protein design papers maintained by someone else.

What licence does the repository use?

The GNU General Public License at version three, applied to a repository whose content is a Markdown reading list rather than code. The page says nothing about the licences of the papers themselves, which remain under their publishers' terms.

Official sources

  1. Issues
  2. License: GPL-3.0
  3. Peldom/papers_for_protein_design_using_DL on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/peldom-papers-for-protein-design-using-dl.svg)](https://hysenlabs.com/projects/peldom-papers-for-protein-design-using-dl)