{"slug": "sloptotal-a-self-hosted-ai-text-detector-that-runs-23-open-models", "title": "SlopTotal, a Self-hosted AI text detector that runs 23 open models", "summary": "SlopTotal, a free self-hosted open-source AI text detector from developer pablocaeg, runs 23 independent detection engines in parallel on a user's own CPU and reports an overall AUC of 0.974 on a 110-text RAID corpus, with 90% of AI text reaching \"Suspicious\" or above and 1 of 66 human texts wrongly labeled \"Likely AI.\" A September 2026 re-measurement on a fresh 180-text RAID sample with upgraded dependencies produced AUC 0.979, 1 of 40 human texts called \"Likely AI,\" and 0 of 26 literary passages flagged. The tool ships as a Docker image (ghcr.io/pablocaeg/sloptotal), accepts text, URLs, PDF, DOCX, TXT and MD input, and publishes its evaluation harness, raw per-sample scores and failure cases, including three engines found scoring backwards and two loading a randomly initialized network while carrying real ensemble weight.", "body_md": "**VirusTotal for AI-generated text.** Paste text, drop in a PDF or Word file, or\ngive it a URL. Twenty-three independent AI detectors (neural classifiers,\nstatistical tests and linguistic heuristics) score it in parallel, and a\ncalibrated ensemble turns their votes into one verdict you can inspect engine by\nengine. It runs on your own CPU, so nothing you scan leaves your machine.\n\nIt is a free, self-hosted, open-source alternative to hosted AI content detectors such as GPTZero, Originality.ai, Copyleaks, ZeroGPT and Humalingo. Instead of one number from one model, it shows you every model's opinion, and it publishes how accurate that is, failures included.\n\n**Try it:** [sloptotal.com](https://sloptotal.com) · **Run it:** `docker run -p 8000:8000 ghcr.io/pablocaeg/sloptotal`\n\n- **23 detection engines, one calibrated score.** DeBERTa and RoBERTa\nclassifiers, Binoculars, Fast-DetectGPT, GLTR, perplexity and burstiness\ntests, and stock-phrase heuristics. Results stream in as each engine finishes.\n- **Text, URLs and documents.** Paste text, scan a web page (main content is\nextracted automatically), or upload`.pdf` ,`.docx` ,`.txt` or`.md` .\n- **Site check: was this website vibe-coded?** Finds the fingerprints that\nLovable, v0, Bolt, Base44, Replit and Same leave in the sites they deploy,\nand shows the evidence for each one.[How it works](#site-check-detect-sites-built-with-ai-app-builders)\n- **Per-paragraph heat map** through the API, to see which parts read as AI.\n- **Measured, not claimed.** Every accuracy number below comes with the corpus,\nthe harness and the raw per-sample scores.\n- **Private by default.** Self-hosted, no third-party AI APIs, no tracking,\nreports deleted after 30 days.\n- **CPU-only is fine.** Auto-detects your hardware; 4 GB RAM is enough for the\nlite profile, a GPU is optional.\n- **JSON API and a [Chrome extension](https://github.com/pablocaeg/sloptotal-extension)** that marks AI-looking results in Google Search and LinkedIn.\n\nMost detectors publish an accuracy figure without saying what it was measured on.\nThese numbers, the harness that produced them and the raw per-sample results are\nall in [tests/eval/](https://github.com/pablocaeg/sloptotal/blob/master/tests/eval).\n\nTwo corpora, deliberately:\n\n| Corpus | What | Size | \n|---|---|---|\n| Multi-domain | RAID: news, book prose, poetry, academic abstracts. AI from GPT-4, ChatGPT, Llama, Mistral, Cohere, GPT-3 | 110 (40 human, 70 AI) | \n| Literary control | Project Gutenberg prose published 1532-1915 -- Machiavelli, Austen, Melville, Kafka | 26 (all human) | \n\nThe second exists because a high score there cannot be anything but an error: the writing predates language models by a century or more. Optimising on the first corpus alone produces a threshold that mislabels literature.\n\n|  | Result | \n|---|---|\n| Overall AUC | 0.974 | \n| AI reaching \"Suspicious\" or above | 90% | \n| Human text wrongly called \"Likely AI\" | 1 of 66 | \n| Literary passages flagged | **0 of 26** | \n\nRe-measured in September 2026 on a fresh RAID sample (180 texts) with upgraded dependencies: AUC 0.979, 1 of 40 human texts called \"Likely AI\", 0 of 26 literary passages flagged.\n\n**What does not work.** Short text is unreliable below roughly 80 words and\nsettles from about 200. Hand-edited AI loses fingerprints with every rewriting\npass. Source code is outside what these engines do: in testing they never falsely\naccused human code, and never caught machine-written code either -- so we do not\nclaim they can.\n\nThe failures are published too, including three engines found scoring backwards\nand two loading a randomly initialised network while carrying real ensemble\nweight. Read them at\n[sloptotal.com/detect/ai-detector-benchmark/](https://sloptotal.com/detect/ai-detector-benchmark/)\nand [sloptotal.com/detect/ai-detector-false-positives/](https://sloptotal.com/detect/ai-detector-false-positives/).\n\n```\ndocker run -p 8000:8000 -v sloptotal-models:/app/models ghcr.io/pablocaeg/sloptotal\n```\n\nOpen [http://localhost:8000](http://localhost:8000). The first scan downloads about 2 GB of models into\nthe `sloptotal-models` volume, so later starts are quick. To build from source\ninstead, run `docker compose up`.\n\nRequires **Python 3.10+** (macOS ships 3.9, which is too old).\n\n```\ngit clone https://github.com/pablocaeg/sloptotal.git\ncd sloptotal\npython3.11 -m venv venv && source venv/bin/activate\npip install -r requirements.txt\n./start.sh            # or: uvicorn app.main:app --port 8000\n```\n\nCheck that every engine loads and scores, end to end:\n\n```\npython scripts/smoke_test.py          # against http://localhost:8000\n```\n\n\"Is this website vibe-coded?\" checkers mostly score style (Tailwind class counts, missing security headers, buzzwords) and turn it into a percentage. Hand-written sites share all of those traits. SlopTotal looks only for markers the builders themselves leave in what they deploy, each one confirmed on live sites or in the builders' own templates:\n\n| Builder | Fingerprints | \n|---|---|\n| Lovable | `gptengineer.js` runtime,`/lovable-uploads/` assets, the Lovable badge,`/~flock.js` ,`*.lovable.app` | \n| v0 (Vercel) | `<meta name=\"generator\" content=\"v0.app\">` from v0's layout template,`*.vusercontent.net` | \n| Bolt | `X-Powered-By: Bolt.new` header,`bolt.new/badge.js` ,`*.bolt.host` | \n| Base44 | `app.base44.com` platform calls,`base44_access_token` ,`*.base44.app` | \n| Replit | Replit Agent dev banner, Replit badge, `*.replit.app` | \n| Same | assets served from `same-assets.com` | \n\nA site with no marker may still have been written with AI: code exported from these tools and hosted elsewhere, or written in an AI editor, carries no fingerprint. So the result is evidence, not a probability. The page's copy is scored separately by the text engines.\n\n```\ncurl -X POST http://localhost:8000/api/scan/site \\\n  -H \"Content-Type: application/json\" -d '{\"url\": \"example.com\"}'\n```\n\n| Endpoint | Method | What it does | Typical latency (CPU) | \n|---|---|---|---|\n| `/api/analyze` | POST | Full 23-engine report for `text` or`url` | 2-8 s | \n| `/api/quick-score` | POST | 4 classifiers plus heuristics | 0.1-0.5 s | \n| `/api/paragraph-score` | POST | Score per paragraph (heat map) | 1-3 s | \n| `/api/scan/site` | POST | AI app builder fingerprints plus a copy score | 1-3 s | \n| `/api/extract` | POST | Text from an uploaded `.pdf` /`.docx` /`.txt` (multipart`file` ) | < 1 s | \n| `/api/scan/snippets` | POST | Batch of 1-30 short snippets | ~0.5 s | \n| `/api/scan/urls` | POST | Batch of 1-10 URLs, page-type aware | 1-5 s | \n| `/api/engines` | GET | Engine metadata | instant | \n| `/api/report/{id}` | GET | A stored report | instant | \n| `/api/queue/status` | GET | Queue capacity | instant | \n\n```\ncurl -X POST http://localhost:8000/api/analyze \\\n  -H \"Content-Type: application/json\" \\\n  -d '{\"text\": \"Your text to analyze here...\"}'\n```\n\nThe response lists every engine with its score, verdict and a plain-language\ndetail line, plus `overall_score` (0-100) and `overall_verdict`.\n\nEvery engine links to its page on sloptotal.com, which carries its measured scores against both corpora. AUC below is the probability the engine ranks a random AI passage above a random human one: 1.0 is perfect, 0.5 is a coin flip.\n\n| Engine | Model | AUC | Notes | \n|---|---|---|---|\n| [Desklib DeBERTa](https://sloptotal.com/engines/desklib-deberta/) | DeBERTa-v3-large (435M) | 1.000 | Strongest separation in our own tests | \n| [SuperAnnotate](https://sloptotal.com/engines/superannotate/) | RoBERTa-large (355M) | 0.989 | No measurable bias against archaic prose | \n| [E5-Small](https://sloptotal.com/engines/e5-small/) | E5 + LoRA (33M) | 0.999 | Matches far larger models at 33M params | \n| [TMR Detector](https://sloptotal.com/engines/tmr-detector/) | RoBERTa-base (125M) | 1.000 | RAID-trained, so RAID scores flatter it | \n| [BERT-tiny RAID](https://sloptotal.com/engines/bert-tiny-raid/) | BERT-tiny (4.4M) | 1.000 | Answers in milliseconds | \n| [ReMoDetect](https://sloptotal.com/engines/remodetect/) | DeBERTa (184M) | 0.941 | Targets RLHF-aligned LLMs | \n| [ChatGPT Detector](https://sloptotal.com/engines/chatgpt-detector/) | RoBERTa-base (125M) | 0.829 | ChatGPT-specific | \n| [Fakespot](https://sloptotal.com/engines/fakespot/) | RoBERTa-base (125M) | 0.999 | Accurate on modern text, but +0.533 bias on pre-1920 prose | \n| [OpenAI Detector](https://sloptotal.com/engines/openai-detector/) | RoBERTa-base (125M) | 0.771 | The 2019 GPT-2 detector; weaker on modern LLMs | \n\n| Engine | Method | AUC | \n|---|---|---|\n| [Log-Rank](https://sloptotal.com/engines/log-rank/) | Average log-rank under GPT-2 | 0.909 | \n| [GLTR](https://sloptotal.com/engines/gltr/) | Token rank distribution | 0.904 | \n| [Perplexity](https://sloptotal.com/engines/perplexity/) | GPT-2 perplexity scoring | 0.901 | \n| [Cross-Perplexity](https://sloptotal.com/engines/cross-perplexity/) | Two-model perplexity comparison | 0.891 | \n| [Fast-DetectGPT](https://sloptotal.com/engines/fast-detectgpt/) | Conditional probability curvature | 0.890 | \n| [Binoculars](https://sloptotal.com/engines/binoculars/) | Cross-entropy ratio between two LMs | 0.836 | \n| [DivEye](https://sloptotal.com/engines/diveye/) | Surprisal diversity | 0.730 | \n\n| Engine | Signal | AUC | \n|---|---|---|\n| [Structural Analysis](https://sloptotal.com/engines/structural-analysis/) | Em-dash usage, sentence uniformity | 0.836 | \n| [Linguistic Markers](https://sloptotal.com/engines/linguistic-markers/) | AI-preferred phrases (\"delve\", \"tapestry\"...) | 0.713 | \n| [Formulaic Patterns](https://sloptotal.com/engines/formulaic-patterns/) | Cliche openings and closings | 0.698 | \n| [Vocabulary Richness](https://sloptotal.com/engines/vocabulary-richness/) | Type-token ratio, hapax legomena | 0.583 | \n| [Readability Uniformity](https://sloptotal.com/engines/readability-uniformity/) | Cross-paragraph consistency | 0.581 | \n| [Burstiness](https://sloptotal.com/engines/burstiness/) | Per-sentence perplexity variance | 0.582 | \n| [Sentiment & Hedging](https://sloptotal.com/engines/sentiment-and-hedging/) | Hedging and forced balance | 0.522 | \n\nThe linguistic heuristics are weak on their own. They are kept because they fail *independently* of the neural classifiers, which is what makes them useful as tiebreakers rather than as evidence.\n\nThe final score is **calibrated**, not a simple average, and every weight is\nderived from measurement rather than intuition. See\n[tests/eval/FINDINGS.md](https://github.com/pablocaeg/sloptotal/blob/master/tests/eval/FINDINGS.md) and\n[sloptotal.com/detect/ai-detector-ensemble/](https://sloptotal.com/detect/ai-detector-ensemble/).\n\n1. **Anchored on the unbiased classifiers** -- Desklib, SuperAnnotate, E5 and\nReMoDetect all score high AUC with no measurable bias against older prose.\nTheir consensus is blended 60/40 with the full weighted set.\n2. **Weights from measurement** -- each engine's share is proportional to\nSomers' D (2*AUC - 1), scaled down by any bias it shows against archaic\nwriting. RAID-trained engines are damped because our corpus is RAID.\n3. **Confidence from agreement** -- a tight cluster across independent engine\nfamilies is trustworthy; one confident engine is not.\n4. **Skepticism, but only when earned** -- unanimous high classifier scores are\ndamped*only* when the text itself carries human markers (contractions,\nfirst-person, slang). Applied unconditionally it fired on 69 of 70 AI samples\nand 0 of 66 human ones, suppressing correct detections.\n\nFakespot was previously the anchor, weighted 0.13. It is accurate on modern text (AUC 0.999) but scored pre-1920 human prose at 0.645 against 0.112 for modern human writing -- the largest bias of any engine -- and anchoring amplified it. Machiavelli scored 62.5. After demotion to 0.033, literary passages average 10.2 and none is flagged.\n\nSlopTotal detects CPU, RAM and GPU at startup and picks a profile. Everything\ncan be overridden with environment variables; see [`.env.example`](https://github.com/pablocaeg/sloptotal/blob/master/.env.example).\n\n| Profile | RAM | CPU | GPU | Notes | \n|---|---|---|---|---|\n| Lite | 4 GB | 2 cores | None | All engines, slower | \n| Standard | 8 GB | 4 cores | None | Default for most laptops | \n| Performance | 16 GB+ | 6+ cores | CUDA optional | Pool replicas, max throughput | \n\n**High-RAM CPU servers (e.g. 64 GB, no GPU):** you automatically get the `performance` profile. With no CUDA, all inference stays on CPU but you can run more concurrent workers and model pool replicas:\n\n```\n# Tune for a 64 GB CPU-only server\nexport SLOPTOTAL_PROFILE=performance\nexport SLOPTOTAL_TORCH_THREADS=8\nexport SLOPTOTAL_FULL_WORKERS=8\nexport SLOPTOTAL_SNIPPET_WORKERS=6\nexport SLOPTOTAL_MAX_CONCURRENT_FULL=4\nexport SLOPTOTAL_POOL_FAKESPOT=2\nexport SLOPTOTAL_POOL_TMR=2\n./start.sh\n```\n\n| Variable | Default | Purpose | \n|---|---|---|\n| `SLOPTOTAL_PROFILE` | auto | `lite` ,`standard` or`performance` | \n| `SLOPTOTAL_RETENTION_DAYS` | `30` | Delete reports after N days ( `0` keeps them) | \n| `SLOPTOTAL_ALLOW_PRIVATE_URLS` | off | Let URL scans reach private or intranet hosts (blocked by default) | \n| `HF_HOME` | `./models` | Where model weights are cached | \n\n**Can AI detectors be trusted?** Not blindly. No detector, this one included,\nshould be the only evidence for an accusation. That is why SlopTotal shows\nall 23 votes, how much they agree, and its measured false-positive rate. Short\ntext (under about 80 words) and heavily edited AI text are unreliable for every\ndetector.\n\n**Does it detect ChatGPT, Claude, Gemini, Llama and Mistral?** The evaluation\ncorpus includes GPT-4, ChatGPT, Llama and Mistral output. The classifiers were\ntrained on a wider mix. Newer models are covered as far as they share those\nfingerprints; the [evaluation harness](https://github.com/pablocaeg/sloptotal/blob/master/tests/eval) lets you measure any model\nyou care about.\n\n**Will it flag classic literature or formal writing?** Not in our tests: none\nof the 26 passages from Austen, Melville, Kafka, Machiavelli and others is\nflagged. Pre-1920 prose is part of the evaluation precisely because naive\ndetectors fail on it.\n\n**Is my text stored or shared?** It is processed on the server you run. Reports\nare kept for 30 days (configurable) so report links work, and nothing is sent\nto an outside service.\n\n**Can it detect AI-generated code?** No, and we do not claim it can: in testing\nthe engines never flagged human code but never caught machine-written code\neither. The Site check reports which AI app builder produced a website, which is\na different question.\n\n| Symptom | Fix | \n|---|---|\n| `TypeError: unsupported operand type(s) for \\|` at startup | Python 3.9 or older; use 3.10+ | \n| An engine reports `Model loading failed` | Check disk space and network for the first model download, then restart; `python scripts/smoke_test.py` shows which engine fails | \n| First scan is slow | Models are loading; later scans take seconds | \n| A URL scan says \"private network address\" | Intended; set `SLOPTOTAL_ALLOW_PRIVATE_URLS=1` to scan intranet pages | \n\n```\napp/            FastAPI backend: engines, ensemble, site fingerprints, API\nweb/            The web UI (Jinja2 templates, vanilla JS, no build step)\ntests/          Unit tests (seconds, no downloads) and tests/eval/ accuracy harness\nscripts/        smoke_test.py (end-to-end) and the model drift check\nbenchmarks/     Speed and load scripts\n```\n\n[ARCHITECTURE.md](https://github.com/pablocaeg/sloptotal/blob/master/ARCHITECTURE.md) covers the internals. [AGENTS.md](https://github.com/pablocaeg/sloptotal/blob/master/AGENTS.md)\nis a short brief for contributors and AI coding assistants.\n\nDetector models keep appearing on Hugging Face, each with its own accuracy claim. Before adding any, we score them on the same two corpora. September 2026, standalone, 180 RAID texts plus the 26 literary passages:\n\n| Model | RAID AUC | Literary bias (lower is better) | Status | \n|---|---|---|---|\n| [Gradient](https://huggingface.co/ShantanuT01/gradient-ai-text-detector) (DeBERTa-v3-large) | 0.998 | 0.033 | Next engine to add | \n| [Vanguard](https://huggingface.co/ShantanuT01/vanguard-ai-text-detector) (ModernBERT-large) | 0.998 | 0.037 | Candidate; poorly calibrated at 0.5 | \n| [Earlybird-fast](https://huggingface.co/noumenon-labs/Earlybird-fast) (82M) | 0.913 | 0.040 | Candidate for fast snippet scans | \n| [rasbt ModernBERT](https://huggingface.co/rasbt/ai-text-detector-modernbert) | 0.864 | 0.000 | Not added | \n\nRAID-trained models score near 1.0 on RAID by construction and need a different\ntest set first. The full table, the models we excluded and why, and the raw\nscores are in [tests/eval/FINDINGS.md](https://github.com/pablocaeg/sloptotal/blob/master/tests/eval/FINDINGS.md#newer-open-detectors-measured-standalone).\nThe roadmap is in [TODO.md](https://github.com/pablocaeg/sloptotal/blob/master/TODO.md).\n\n- [RAID benchmark](https://arxiv.org/abs/2405.07940) (ACL 2024): adversarial AI text detection dataset used in our evaluation\n- [Detecting the Machine (2026)](https://arxiv.org/pdf/2603.17522) : cross-architecture detector benchmark; ensembles beat single detectors\n- [EditLens](https://arxiv.org/abs/2510.03154) : human vs AI-edited vs AI-generated classification\n- [GLTR](http://gltr.io/) : visual token-rank inspection, the inspiration for our GLTR engine\n- [distil-labs/distil-ai-slop-detector](https://github.com/distil-labs/distil-ai-slop-detector) : a 270M Gemma detector that runs in the browser\n- [sloptotal-extension](https://github.com/pablocaeg/sloptotal-extension) : the Chrome extension\n\nContributions are welcome, especially new engines with measurements. Start with\n[CONTRIBUTING.md](https://github.com/pablocaeg/sloptotal/blob/master/CONTRIBUTING.md).\n\nIf SlopTotal is useful to you, a star helps other people find it.\n\nMIT. Model weights keep their own licenses; see\n[THIRD_PARTY_LICENSES.md](https://github.com/pablocaeg/sloptotal/blob/master/THIRD_PARTY_LICENSES.md).", "url": "https://wpnews.pro/news/sloptotal-a-self-hosted-ai-text-detector-that-runs-23-open-models", "canonical_source": "https://github.com/pablocaeg/sloptotal", "published_at": "2026-09-28 12:00:05+00:00", "updated_at": "2026-09-28 12:18:05.661465+00:00", "lang": "en", "topics": ["ai-tools", "ai-products", "machine-learning", "natural-language-processing", "developer-tools"], "entities": ["SlopTotal", "pablocaeg", "GPTZero", "Originality.ai", "Copyleaks", "ZeroGPT", "Humalingo", "RAID"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/sloptotal-a-self-hosted-ai-text-detector-that-runs-23-open-models", "markdown": "https://wpnews.pro/news/sloptotal-a-self-hosted-ai-text-detector-that-runs-23-open-models.md", "text": "https://wpnews.pro/news/sloptotal-a-self-hosted-ai-text-detector-that-runs-23-open-models.txt", "jsonld": "https://wpnews.pro/news/sloptotal-a-self-hosted-ai-text-detector-that-runs-23-open-models.jsonld"}}