RLCF turns citation counts into a reward model for research ideas
We propose Reinforcement Learning from Community Feedback (RLCF), a training paradigm that uses large-scale community signals as supervision, and formulate scientific taste learning as a preference modeling and alignment problem.
At a glance
- What is it?
- A paper and model release that reframes scientific taste as preference modelling: match papers by field and publication period, use which one got cited as the preference label, train a judge with GRPO, then train an ideator against that judge. The repository holds no code at all, only a PDF, a licence and a README.
- Who is it for?
- Use this as a reference for a method rather than as a tool. There is no code in the repository, so nothing here can be run or extended, and the trained checkpoints live on a Hugging Face collection rather than in the tree.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 18 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The repository is a PDF, a licence and a README, with no code at all
The top level holds .gitignore, an assets directory, an English README, a Chinese README, the licence, and one PDF whose filename carries a typo, Scientifc rather than Scientific. No source directory, no requirements file, no build script, and no detected primary language.
Everything that can be run was released somewhere else. The paper is on arXiv as 2603.14473, the models are in a Hugging Face collection under the OpenMOSS-Team organisation, and the interactive demo is hosted on a separate domain.
So this repository is a citation for a paper. That is a legitimate way to release work, particularly where the artefact is a model rather than a library, but it does mean the claim has to be evaluated from the numbers in the README and from the paper, since there is nothing here to inspect.
The licence is Apache 2.0, which is worth noting because the weights and the paper sit under different terms and only the paper-side terms are stated here.
Citations become labels by matching papers inside a field and a publication period
The first of three stages is where the whole method lives, and it is a data construction rather than a model.
Citations are converted into pairwise preference signals by matching papers within the same field and the same publication period. A pair where one paper accumulated high citation and the other did not becomes a preference example, with the highly cited paper as the winner.
The matching constraints are what stop the signal from being noise. Without a field match, a chemistry paper is being compared to a machine learning paper and the citation difference says nothing about taste. Without a period match, a 2019 paper has had longer to accumulate citations than a 2024 one, and the label would encode recency rather than impact.
This is the step to copy if you want the same supervision for a different field, because it needs a citation graph, a field classification and a date, and nothing else.
Two models, one as judge and one as ideator, trained in sequence
Scientific Judge is a generative reward model that reasons over a pair of paper abstracts and predicts which is more likely to have higher impact. It is trained with GRPO on 720K field-matched and time-matched citation-based preference pairs, and it does two jobs: it evaluates research ideas, and it serves as the reward model for the second model.
Scientific Thinker is a scientific ideation policy trained with the judge as its reward. It takes a paper title and abstract as input and proposes follow-up research ideas, optimised with comparison-based GRPO for open-ended generation.
So the sequence is: build preferences from citations, fit a reward model on them, then optimise a policy against that reward model. Everything is comparison based, which is what lets it work on open-ended generation where a reference answer does not exist.
The reward model is doing the alignment work that a human reviewer would otherwise do, and the quality of the whole system is bounded by how good the citation-derived labels are.
The in-domain margin over GPT-5.4 Thinking is 1.1 points
The headline result is Scientific Judge-Qwen3-30B reaching 82.7% in-domain accuracy, surpassing all listed strong LLM baselines including GPT-5.4 Thinking at 81.6%.
Read that margin carefully. 1.1 points on a paired preference task where the strongest general model is already above 80% is not a decisive separation. It says the specialised model edges out a frontier model on this particular dataset, not that general models cannot do the task.
The more informative claim from the same section is the scaling one: scientific judgement is said to scale with both data size and model size. If that holds, the interesting version of this result is not 30B beating a large general model but small specialised models improving steeply with more citation pairs, since that is the regime where a domain-specific approach can win outright.
The Thinker numbers are reported separately and use a different metric, a win rate rather than accuracy.
The temporal split is the result that carries the argument
Everything else on this page is in-domain or near-domain. One number is not, and it is the one to look at.
On the refreshed 904-pair test set built from papers published in 2025, Qwen3-4B improves from 64.7% to 80.9%, a gain of 16.2 points, and Qwen3-30B-A3B improves from 71.7% to 83.1%, a gain of 11.4 points.
A 16 point jump on papers the training could not have contained is the evidence that something transferable was learned rather than a period artefact. It is also the number that would decay fastest if citation counts for 2025 papers keep rising as they accumulate, since the test set was refreshed in July 2026 against papers that were still young.
That refresh is worth watching. The same reasoning applies to any citation-derived label: the ground truth is not fixed, it moves upward as time passes, so a benchmark built from it needs periodic re-scoring.
720,341 pairs and 1,440,682 records are not 1.4 million papers
The benchmark section is unusually careful about one number and it is worth repeating in the same spirit.
SciJudgeBench contains 720,341 preference pairs and 1,440,682 pair-level paper records, and the page states directly that papers may recur across pairs, so this is not a unique-paper count.
That distinction is the difference between a benchmark over half a million distinct works and one over roughly a million rows of data about a smaller set. Both numbers are correct; only the second is a data volume.
The corpus is arXiv papers across Computer Science, Mathematics, Physics, and a category named Other, so three named domains and a catch-all. Evaluation runs in-domain and then across temporal out-of-distribution with the 904 future-year pairs, metric out-of-distribution using ICLR peer review and Altmetric attention, field transfer, and controlled comparison, with bioRxiv as an additional biology evaluation.
The Thinker's 54.2% is measured against its own base policy at 30.3%
The ideator result needs its denominator, and the denominator is in the same sentence.
Scientific Thinker achieves a 54.2% average win rate against three strong LLM baselines in both in-domain and out-of-domain settings. Its base policy scores 30.3% and 27.8% respectively.
So 54.2% is a majority but a thin one, against baselines that win 46% of comparisons. Whether that is a meaningful gain depends entirely on who the three baselines are and how the wins are judged, neither of which is broken out.
The structural point is worth more than the number. The policy is being trained against a reward model that was itself trained on citations, so the ceiling on the ideator is the ceiling on the judge. A better judge would produce a better thinker, and a judge with a citation bias will pass that bias to every idea it scores as high impact.
That coupling is the main thing to keep in mind when reading both halves of this work.
Metric out-of-distribution asks whether citation count is the only taste signal
The strongest part of the evaluation design is the metric out-of-distribution split, and it is the part most likely to be skipped by anyone reproducing this.
Training labels come from citations. Evaluation on ICLR peer review and on Altmetric attention asks whether the learned judgement predicts preferences that were never expressed as a citation at all. If a model only learned citation behaviour, those two evaluations should not move.
The reported result is that judgement generalises across fields and community metrics, including bioRxiv biology transfer, ICLR peer-review preferences and Altmetric attention, and that the gains persist under author and institution controls and topic controls.
Those controls are what make the claim load bearing. Without them, a model could appear to judge impact well simply because it had learned which authors and which topics accumulate citations, which is a very different capability from judging whether an idea is good.
Editorial conclusion
Use this as a reference for a method rather than as a tool. There is no code in the repository, so nothing here can be run or extended, and the trained checkpoints live on a Hugging Face collection rather than in the tree. What the work contributes is a construction: if you have a citation graph and a way to match papers within a field and period, you can build the same preference set, and the temporal split is the part worth copying. Judge it on two numbers. The in-domain figure, 82.7% against GPT-5.4 Thinking at 81.6%, is a narrow margin on a task where strong general models are already close to the ceiling. The temporal figure is the interesting one: a 4B model going from 64.7% to 80.9% on papers from 2025 is evidence that the signal transfers forward rather than memorising a period. Also note the Thinker's 54.2% win rate is against a base policy scoring 30.3%, so judge the margin, not the number.
Frequently asked questions
Can AI develop taste?
The paper argues it can, framing scientific taste as the ability to judge and propose research ideas with potential for long-term impact. The evidence offered is Reinforcement Learning from Community Feedback: citation-derived preference pairs train a judge, and that judge then trains a model to propose research ideas.
Can AI accelerate scientific discovery?
The paper describes its results as marking an important step towards AI systems that could help accelerate scientific discovery, and does not claim to have done so. The demonstrated capability is a model that ranks paired abstracts by likely impact and proposes follow-up ideas that win 54.2% of comparisons against three baselines.
Is there code in the AI Can Learn Scientific Taste repository?
No. The top level holds a README in English and Chinese, a licence, an assets directory and a PDF of the paper. The primary language is not even detected, and the models were released separately in a Hugging Face collection under the OpenMOSS-Team organisation.
How big is the SciJudgeBench dataset?
720,341 preference pairs and 1,440,682 pair-level paper records, built from arXiv papers across Computer Science, Mathematics, Physics and an Other category. The page notes that papers may recur across pairs, so the record count is not a unique-paper count.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/tongjingqi-ai-can-learn-scientific-taste)