The [arXiv](https://arxiv.org/),
[bioRxiv](https://www.biorxiv.org/), and
[medRxiv](https://www.medrxiv.org/)
are free-to-use "preprint servers" where scientists share and discover science. They connect us to other scientists, to relevant results, and ultimately
[move science forward (faster than journals)](https://doi.org/10.1371/journal.pbio.3000959).
But with the rise of AI-written papers, discovering relevant work on preprint servers is increasingly challenging. arXiv
[reported 40,363 submissions in September 2026](https://blog.arxiv.org/2026/10/01/updated-rate-limit-policy/),
nearly twice the number it received two years earlier, warning that AI tools are making it easier to flood preprint servers. This means finding papers worth reading becomes more time-consuming, leaving less time to read and understand them. How can scientists stay afloat without spending all their time treading water?
One solution is to limit the number of papers through submission caps, which is what arXiv chose to do. On October 1, it
[announced a limit of two submissions per calendar month](https://blog.arxiv.org/2026/10/01/updated-rate-limit-policy/)
per submitter to reduce the burden on its moderators. The decision was met with criticism. Sabine Hossenfelder
[predicted that papers would get longer](https://x.com/skdh/status/2105865508163658240),
while others
[argued that Euler would not have been able to post his work](https://x.com/predict_addict/status/2106320049422184888)
were he alive today. Even if submissions slow, longer papers could leave scientists with just as much to search through.
A second solution is to give scientists ways to find papers relevant to their research interests. Tools such as
[arXiv Sanity](https://github.com/karpathy/arxiv-sanity-lite),
[Scholar Inbox](https://arxiv.org/abs/2504.08385), and
[Prompts & Papers](https://promptsandpapers.com/)
already offer personalized paper recommendations with techniques like TF-IDF and learned paper embeddings. I hypothesized that decision language models could match papers directly to a scientist's stated interests with comparable or better accuracy, while remaining fast and inexpensive. I built
[author.link](https://author.link/)
to test this idea and help scientists prioritize what to read.
## [author.link](https://author.link/) helps you find relevant papers
You describe your research interests (the more descriptive, the better!), and
[author.link](https://author.link/)
ranks new papers from bioRxiv, medRxiv, and arXiv by their estimated relevance to you. It checks for new papers every four hours and scores them as they arrive. (You can also opt in to a daily email with a short list of relevant papers.)
[author.link](https://author.link/)
uses a decision language model (LM) to rank papers. Rather than relying on an LLM to summarize papers and leave you to filter the summaries, a decision LM scores their relevance directly. Given your research interests and a paper's title and abstract, the model returns an estimated probability that the paper is relevant to you.
[Typesafe's Jev](https://typesafe.ai/blog/introducing-system-one-models-and-jev)
provides this interface, as do
[OpenAI's Decisions API](https://developers.openai.com/api/docs/guides/decisions)
and open-weight models such as
[Laya](https://huggingface.co/convaiinnovations/laya)
and
[Cloudflare's Clef](https://huggingface.co/Cloudflare/clef).
[author.link](https://author.link/)
intentionally does not evaluate scientific quality; it's simply a filter for relevance to your interests. A paper can be excellent and completely irrelevant to your work. Another can introduce a method or dataset that is useful to you without being important to everyone else. As I argued in
[Does your paper really suck?](https://sina.bio/posts/does-your-paper-really-suck.html),
assigning a universal score for scientific quality is a much stronger claim than helping someone decide what to read next—[author.link](https://author.link/)
does the latter and leaves it up to you, the scientist, to evaluate the work.
Decision language models rank relevance #
To test how well decision LMs rank relevant papers, I compared Jev (`jev-1.13.0`) with OpenAI's Decisions API (` gpt-6-luna`) on two benchmarks.
The first benchmark consists of 66 papers that I selected to test relevance of them to a fixed description of my research interests. It contains 22 papers I coauthored, 22 keyword-matched controls, and 22 controls without those keyword matches. I use these groups as proxies for high, medium, and low relevance to my research interests, respectively. I call this Sina's Science Relevancy Benchmark (SinSciRelBench). The second benchmark is
[RELISH](https://doi.org/10.1093/database/baz085),
a dataset of papers judged by humans to be relevant, partially relevant, or irrelevant to other papers.
For each benchmark, both models received the same inputs and relevance criteria. I used Jev's `noul` option and OpenAI's `predicate` option to score relevance. (Note: a score of 0.8 should not be read as an established 80% chance that a reader will find the paper useful.)
Jev and OpenAI's Decisions API largely agreed on the relevance ordering of papers in SinSciRelBench (Spearman correlation 0.945). For separating my papers from the keyword-unmatched controls, Jev achieved an ROC AUC of 0.981 and OpenAI achieved 0.915. Including both control groups made the comparison slightly harder (the AUCs were 0.938 and 0.893, respectively). Interestingly, both models also ranked a keyword-matched control paper,
[Gravlax](https://doi.org/10.64898/2026.09.18.752708),
highly (0.92 for Jev and 0.98 for OpenAI). Turns out, this paper was authored by Professor Rob Patro,
[whose work greatly overlaps with mine](https://www.biorxiv.org/content/10.1101/2021.01.25.428188v2)!
On RELISH, both decision models outperformed TF-IDF, a baseline that measures similarity using shared words. For separating relevant from irrelevant papers, Jev achieved an average ROC AUC of 0.893 and OpenAI achieved 0.872, compared with 0.721 for TF-IDF. Separating relevant papers from both partially relevant and irrelevant papers was harder (the AUCs were 0.799, 0.785, and 0.653, respectively).
Both models were fast and inexpensive. OpenAI had the faster median response, while Jev had fewer slow responses and lower cost. At input token rates of $0.042 per million for Jev and $0.10 for OpenAI, OpenAI cost about 1.87 times as much as Jev on SinSciRelBench (output tokens are free).
SinSciRelBench is, of course, a small, single-person benchmark, and RELISH tests a related but different problem (finding papers relevant to other papers rather than to a scientist's stated interests). The real test is whether
[author.link](https://author.link/)
helps other scientists find papers they actually want to read.
Give scientists a lifejacket and let them write #
Limiting scientific production to combat the explosion of submitted papers is like
[bailing buckets of water out of a sinking ship](https://youtu.be/Z_Z6RAd3-ZE?t=823).
Instead, we need tools that help scientists discover work relevant to them and make it easier to evaluate that work.
[author.link](https://author.link/)
is my attempt at this solution—a lifejacket to keep scientists afloat, if you will—made possible by fast and cheap decision language models. Scientists can continue to share their work, and readers can find the papers relevant to them.
Please [try author.link](https://author.link/) and
[send me feedback](mailto:feedback@author.link).
Does it surface papers you want to read? What does it miss? This is an actively developing project, and your feedback would be very helpful.