cd /news/ai-safety/science-or-slop · home › topics › ai-safety › article
[ARTICLE · art-143737] src=yerimoh.github.io ↗ pub= topic=ai-safety verified=true sentiment=· neutral

Science or Slop?

The Science Slop Index, a tool from the project Science or Slop?, scores papers on six measures of scientific reasoning rather than token-level AI detection, reporting a result out of 100. The project reports pair accuracy on SciSlopBench, where 390 AI-generated papers were each matched with a human-written paper on the same problem, with 0.5 as chance. SciSlop v2 is being written with community contributions, and users can submit a link or upload a PDF or LaTeX .zip, receiving a private report key; uploaded files are deleted after analysis while only the report and rendered figures remain.

read4 min views2 publishedOct 2, 2026
Science or Slop?
Image: source

A paper can read well sentence by sentence while its science does not connect. Unlike token-level AI detectors, the Science Slop Index reads the scientific reasoning: whether sections build on one another, claims and citations are argued, and the method and evidence can be inspected.

Paste a link or upload a PDF / LaTeX .zip. Your report is private: you get a key, and you decide on the report whether to list it.

SciSlop v2 is being written with the community. Flag slop on any paper and become a co-author. How ↓

Papers analyzed so far #

· Leaderboard · Gallery Click the center paper to open its full report; click a side one to bring it forward. Highlights on each first page are the findings.

Team #

The people behind Science or Slop? and this site.

How it works

Each measure counts the share of a paper's units that show one pattern. The index averages the measures within each plane, then averages the planes, and reports the result out of 100. Higher means more of the patterns that set AI-generated papers apart from matched human papers.

Reads the paper, not the prose #

AI text detectors score sentences: how predictable each token is, how a model would have phrased it. A paper written or revised by a language model can pass that test and still fail as science. Each paragraph looks fine on its own while the connections between paragraphs do not hold: a figure that no section ever refers to, a claim that arrives before anything leads up to it, a related-work paragraph that lists citations without relating them, a method diagram crowded with hyperparameters, a results table with no concrete example anywhere in the paper.

The Science Slop Index counts these breakdowns. It reads the whole paper, groups six patterns into three planes, and asks three questions. Below, each plane is illustrated with real findings from the papers analyzed on this site; use the arrows to browse them, and click a finding to open it on the paper.

What happens to a paper #

Every report receives a key that reopens it later. Uploaded files are deleted after analysis; only the report and rendered figures stay.

1Read

LaTeX source is read directly; arXiv links fetch the source and the PDF. PDFs are rebuilt into sections, captions, references, and citations.

2Measure

Four measures are exact counting rules. Argument graph and Figure exposition ask a language model to label sentences and read the method figure.

3Highlight

Every flagged unit is placed back on the PDF, colored by plane, with a graph per measure. Export as highlighted PDF, JSON, or CSV.

The six measures #

After Table 1 of the paper. Each score is a share of units: the numerator over the denominator. Pair accuracy is how often that measure alone ranks the human-written paper above its AI-generated counterpart on SciSlopBench.

One number #

The index averages the measures within each plane, then averages the three planes. A measure that does not apply to a paper is skipped, not counted as zero.

The four bands are descriptive cut-points on the share of flagged units, not a probability of AI authorship.

Reported pair accuracy #

Table 2 of the paper: on SciSlopBench, 390 AI-generated papers each matched with a human-written paper on the same problem. 0.5 is chance.

Papers analyzed on this site, ranked by Science Slop Index. Higher means more scientific slop. Click a row to see its six measures; open the report for every finding on the paper.

+ Add a paper Gallery

Papers analyzed from public links, and uploads whose owners chose to list them. Ranked by Science Slop Index.

View report

Enter the report key you received when you submitted your paper.

Propose a pattern

Seen a recurring way AI-generated papers break as science that the six measures miss? Name it and say what a reader notices. We do the rest: implement it, run it on SciSlopBench, and adopt it if it separates AI from human papers, with your name on it.

── more in #ai-safety 4 stories · sorted by recency
── more on @science or slop? 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/science-or-slop] indexed:0 read:4min 2026-10-02 · —