masamasa59/ai-agent-papers: A Curated Four-Layer Reading List for Agent Research
A collection of AI Agents papers (Updated biweekly)
At a glance
- What is it?
- This repository is a hand-picked, biweekly-updated index of AI agent papers, filed into capabilities, architecture, operations and applications. It is a reading list, not a toolkit, and its value depends on how closely its taxonomy matches the way you think about agents.
- Who is it for?
- Adopt this list if you are doing a literature review, writing a survey, or trying to find where a specific subfield such as failure attribution or self-evolution currently stands, and you want a maintainer to have done the first pass. Do not adopt it if you need runnable code, reproducible benchmark numbers, or a stable identifier for each paper, because the repository is a set of Markdown indexes with links and nothing else.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 12 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What ai-agent-papers actually is, and who it is not for
The README describes the repository as a curation of "the latest research papers on the applications and architectural technologies of AI agents." The maintainer runs weekly arXiv searches using specific keywords and picks only papers that seem particularly interesting. The stated selection rule is explicit: the project does not try to be comprehensive, and papers are added when they introduce a distinctively new approach or novel concept that stands out from existing methods.
That single sentence defines the audience. This is for an engineer or researcher who already knows the agent literature well enough to want a filtered view, and who would rather read twenty selected papers than two hundred search results. It is not for someone who needs a complete bibliography, and it is not for someone who wants code. There is no package, no importable module and no CLI in the repository layout; the top-level entries are directories of Markdown files plus scripts/ and assets/.
The selection bias is the product. A list that admits it skips papers is more useful than one that claims completeness, but it also means absence from this repository tells you nothing about a paper's quality.
The four-layer taxonomy and how papers get filed
Papers are filed in four layers: capabilities (what an agent can do), architecture (how it is built), operations (how it is run) and applications (where it is used). Each layer breaks into clusters, and each cluster links to a date-ordered Markdown file. TAXONOMY.md is described in the README as the full directory map and the rules for where each paper is filed, which matters because the filing decision is the only editorial judgement a reader can audit.
The structure is uneven by design. Under capabilities, Core Cognition holds reasoning, planning, ideation and perception, while Adaptation and Self-Improvement holds exploration, experience and trajectory learning, failure attribution, self-correction, verification, self-evolution and agent tuning. Trust and Measurement holds safety and agent evaluation. The architecture layer has only three files: agent design and frameworks, multi-agent systems, and harness. Applications splits three ways, by interface (embodied, computer-use, web, mobile), by domain (finance, enterprise, AI scientist, vertical), and by system pattern (coding, data, deep research, world simulation).
A reader looking for retrieval-augmented generation will not find a top-level home for it. The nearest cluster is Knowledge and Context, which covers memory, context engineering and knowledge graphs. That is a defensible grouping, but it forces you to learn the maintainer's vocabulary before you can search the list, and the README does not offer a mapping from common terms to clusters.
Badges, recency markers and the update script
The README carries a badge convention: a flame marks recommended papers, a book marks surveys, and scales mark benchmarks. Cluster headings show a count of recent additions in the form (+N), and the README states that badges show papers added in the last two months, with the example window given as Jul to Aug 2026. The flame on a cluster heading is described as high activity.
Those counts are regenerated by a script, and the README names it: python scripts/update_readme_badges.py. That is the one piece of automation visible in the repository, and it tells you something about the maintenance model. The paper lists themselves are edited by hand; only the aggregate numbers in the README are produced by code. If you fork the repository and add entries, the counts will drift from the lists until you run the script.
The two-month window is a design choice with a cost. A cluster that was busy last year but quiet this quarter shows no badge, which makes it look dormant even though the underlying list may be long and still worth reading. The badge measures recent activity, not accumulated depth, and the README does not label it that way.
Installing nothing: a first real use of the reading list
There is no install step, because there is nothing to install. The README gives no pip command, no clone instructions and no environment setup, and the repository has no homepage. The way to get it is to clone the Git repository from GitHub, or to read the Markdown files directly in the browser.
If you want the files locally, the standard clone is the only step the repository supports:
git clone https://github.com/masamasa59/ai-agent-papers.git
cd ai-agent-papersAfter that you are in a directory of Markdown. To find the reading list for a specific topic, follow the paths the README prints. For example, the self-evolution list is at capabilities/adaptation/self-evolution.md, and the harness list is at architecture/harness.md. Opening one gives you a date-ordered set of paper entries with links.
The one script the README documents regenerates the badge counts in the README after you add papers:
python scripts/update_readme_badges.pyRun it from the repository root. What you should see is the README's (+N) counts and flame markers rewritten to match the current contents of the lists. If you have not added anything, expect no meaningful change. The README does not document what the script does when a list file is missing or malformed, so treat a silent run as unverified rather than as success.
Where this list breaks down
The first limitation is that nothing here is verified by execution. An entry is a title and a link. There is no summary of results, no note on whether the code was released, and no indication of whether the reported numbers were reproduced. If you are choosing between two methods for a production system, this repository can tell you that both exist and roughly where they sit in the taxonomy. It cannot tell you which one works.
The second limitation is the update cadence against the volume of the field. The README says weekly arXiv searches, and the repository description says updated biweekly. Either way, the selection filter is novelty, and novelty is judged by one maintainer. Topics with a lot of near-duplicate submissions will be under-represented by construction, because a paper that improves a known method by a few points does not meet the stated bar.
The third limitation is language. The trend newsletters section of the README is written in Japanese, with headings such as 研究トレンド and per-month files under newsletters/. The paper lists are in English. A reader who wants the monthly deep dives needs to read Japanese, and the README does not say whether translations exist. The newsletters also describe a different method from the lists: from 2026-06 onward they are said to involve close reading of the arXiv HTML full text and citation of figures, which is a heavier process than the list entries get.
Finally, the licence is not stated in the repository. That is a real obstacle if you intend to mirror the content or reuse the taxonomy in your own documentation.
How it compares with zjunlp/LLMAgentPapers and the awesome-* lists
The README points to three other collections: zjunlp/LLMAgentPapers, hyp1231/awesome-llm-powered-agent, and kaushikb11/awesome-llm-agents. The difference is in the filing scheme rather than the subject matter. The awesome-* convention is a flat or lightly grouped list of links, which is easy to scan and easy to contribute to, but it gives you no way to ask where a subfield sits relative to the rest of the field.
This repository's four layers are an argument about how agent research decomposes. Putting harness next to multi-agent systems under architecture, and putting observability and governance under operations rather than architecture, is a claim that runtime concerns are separate from design concerns. You may disagree, and if you do, the flat lists will suit you better because they make no such claim.
The other difference is the editorial filter. A flat awesome list accepts what contributors send. This one applies a novelty test and admits that it skips papers. If you want breadth, the flat lists win. If you want a smaller set that someone has already argued is distinctive, this is the one to open first, with the caveat that you are trusting one person's judgement.
Maintenance cost, licence and what to check before you depend on it
The repository is not archived, and the last push was on 2026-08-29. That is recent enough that the lists are likely to reflect the current state of the field, but the repository gives no release history and no changelog, so there is no way to see how the taxonomy has shifted over time or whether entries are ever removed.
Upgrade cost is close to zero in the software sense, because there is no dependency to update. The cost is editorial: if you fork and add papers, you inherit the filing rules in TAXONOMY.md and the badge script, and you have to keep both consistent. The README does not document rollback, conflict resolution or a contribution process, so a fork is effectively your own project from the first commit.
On licensing, the repository's licence is not identified, and the README does not state terms for reuse. Paper titles and abstracts belong to their authors and publishers, and the links point at arXiv and similar hosts. If you plan to republish the curated lists, resolve the licence question at the source before you do, and treat the absence of a stated licence as a reason to ask rather than an implied permission. This is not legal advice.
Editorial conclusion
Adopt this list if you are doing a literature review, writing a survey, or trying to find where a specific subfield such as failure attribution or self-evolution currently stands, and you want a maintainer to have done the first pass. Do not adopt it if you need runnable code, reproducible benchmark numbers, or a stable identifier for each paper, because the repository is a set of Markdown indexes with links and nothing else. Before relying on it, open TAXONOMY.md and check whether its four-layer scheme matches your own mental model, then open capabilities/adaptation/self-evolution.md and see whether the entries there are the ones you would have chosen. If the filing rules put your topic in a layer you would not have picked, the list will cost you more time than it saves.
Frequently asked questions
Does masamasa59/ai-agent-papers include the actual PDFs of the papers?
No. The repository contains Markdown reading lists with links to papers, plus newsletters and a taxonomy file. The README describes it as a curation of papers, and the directory listing shows no paper files, only index files under capabilities/, architecture/, operations/ and applications/.
How often is masamasa59/ai-agent-papers updated?
The README says the maintainer performs weekly arXiv searches, while the repository description says it is updated biweekly. The last push was on 2026-08-29, so both statements describe a list that changes on a short cycle.
What do the flame, book and scales badges mean in masamasa59/ai-agent-papers?
The README defines a flame as recommended papers, a book as survey papers, and scales as benchmark papers. Cluster headings also carry a (+N) count of recent additions and a flame for high activity, and those counts are regenerated by python scripts/update_readme_badges.py.
Is there a licence for masamasa59/ai-agent-papers?
The licence is not identified in the repository, and the README does not state reuse terms. If you intend to republish the lists or reuse the taxonomy, resolve that question at the repository before doing so.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/masamasa59-ai-agent-papers)