SlopTotal 开源自托管 AI 文本检测器,23 个引擎在 CPU 上并行打分
Self-host your own AI text detector on CPU to filter out slop
SlopTotal 是一个可自托管的开源 AI 文本检测器,用 23 个独立检测引擎(DeBERTa、RoBERTa 分类器、Binoculars、Fast-DetectGPT、GLTR、困惑度与语言启发式)并行打分,再由校准后的集成模型给出单一判定,支持文本、URL 和 .pdf/.docx/.txt/.md 文件,全部在本机 CPU 上运行。
VirusTotal for AI-generated text. Paste text, drop in a PDF or Word file, or give it a URL. Twenty-three independent AI detectors (neural classifiers, statistical tests and linguistic heuristics) score it in parallel, and a calibrated ensemble turns their votes into one verdict you can inspect engine by engine. It runs on your own CPU, so nothing you scan leaves your machine.
It is a free, self-hosted, open-source alternative to hosted AI content detectors such as GPTZero, Originality.ai, Copyleaks, ZeroGPT and Humalingo. Instead of one number from one model, it shows you every model's opinion, and it publishes how accurate that is, failures included.
Try it: sloptotal.com · Run it: docker run -p 8000:8000 ghcr.io/pablocaeg/sloptotal
Features
- 23 detection engines, one calibrated score. DeBERTa and RoBERTa classifiers, Binoculars, Fast-DetectGPT, GLTR, perplexity and burstiness tests, and stock-phrase heuristics. Results stream in as each engine finishes.
- Text, URLs and documents. Paste text, scan a web page (main content is
extracted automatically), or upload
.pdf,.docx,.txtor.md. - Site check: was this website vibe-coded? Finds the fingerprints that Lovable, v0, Bolt, Base44, Replit and Same leave in the sites they deploy, and shows the evidence for each one. How it works
- Per-paragraph heat map through the API, to see which parts read as AI.
- Measured, not claimed. Every accuracy number below comes with the corpus, the harness and the raw per-sample scores.
- Private by default. Self-hosted, no third-party AI APIs, no tracking, reports deleted after 30 days.
- CPU-only is fine. Auto-detects your hardware; 4 GB RAM is enough for the lite profile, a GPU is optional.
- JSON API and a Chrome extension that marks AI-looking results in Google Search and LinkedIn.
Measured accuracy
Most detectors publish an accuracy figure without saying what it was measured on. SlopTotal is measured on SlopBench: 1,626 human texts, every one written before ChatGPT, and 1,626 AI texts on the same topics and at the same lengths from 14 current models, across 15 kinds of writing (news, Wikipedia, arXiv, Stack Exchange, Reddit, reviews, student essays, non-native English, fiction and literature published 1532-1915). Every number below is measured on kinds of writing the model was not tuned on.
| AUC | 0.942 |
| AI texts called "Likely AI" (55+) | 67% |
| Human texts called "Likely AI" (55+) | 1.7% |
| Human texts flagged at all (45+) | 3.9% |
| Literature published 1532-1915 flagged | 0 of 87 |
The score bands are anchored on that human text: 45 is where the top 5% of human writing begins, 55 the top 2%, 80 the top 0.5%. So "Likely AI" means fewer than 2 in 100 human texts score this high.
Other languages. Spanish, French, German, Italian, Portuguese, Dutch, Polish, Russian and Japanese are supported (AUC 0.91 to 0.997); Arabic and Korean are experimental; Hindi, Turkish and Chinese are not reliable yet, and the report says so.
What does not work. Essays by non-native English writers are still flagged more than native ones (15% of TOEFL essays called Likely AI, against none of 88 US school essays). Under about 80 words a score is a weak signal. AI text run through a "humanizer" is caught about half the time. Source code is outside what these engines do. All of it, per source, per model and per engine, is in the findings, including that the classifiers which top the RAID benchmark drop to AUC 0.75 on current models.
Quick start
Docker (fastest)
docker run -p 8000:8000 -v sloptotal-models:/app/models ghcr.io/pablocaeg/sloptotal
Open http://localhost:8000. The first scan downloads about 2 GB of models into
the sloptotal-models volume, so later starts are quick. To build from source
instead, run docker compose -f docker/docker-compose.yml up.
From source
Requires Python 3.10+ (macOS ships 3.9, which is too old).
git clone https://github.com/pablocaeg/sloptotal.git cd sloptotal python3.11 -m venv venv && source venv/bin/activate pip install -r requirements.txt ./scripts/start.sh # or: uvicorn app.main:app --port 8000
Check that every engine loads and scores, end to end:
python scripts/smoke_test.py # against http://localhost:8000Site check: detect sites built with AI app builders
"Is this website vibe-coded?" checkers mostly score style (Tailwind class counts, missing security headers, buzzwords) and turn it into a percentage. Hand-written sites share all of those traits. SlopTotal looks only for markers the builders themselves leave in what they deploy, each one confirmed on live sites or in the builders' own templates:
| Builder | Fingerprints |
|---|---|
| Lovable | gptengineer.js runtime, /lovable-uploads/ assets, the Lovable badge, /~flock.js, *.lovable.app |
| v0 (Vercel) | <meta name="generator" content="v0.app"> from v0's layout template, *.vusercontent.net |
| Bolt | X-Powered-By: Bolt.new header, bolt.new/badge.js, *.bolt.host |
| Base44 | app.base44.com platform calls, base44_access_token, *.base44.app |
| Replit | Replit Agent dev banner, Replit badge, *.replit.app |
| Same | assets served from same-assets.com |
A site with no marker may still have been written with AI: code exported from these tools and hosted elsewhere, or written in an AI editor, carries no fingerprint. So the result is evidence, not a probability. The page's copy is scored separately by the text engines.
curl -X POST http://localhost:8000/api/scan/site \ -H "Content-Type: application/json" -d '{"url": "example.com"}'
API
| Endpoint | Method | What it does | Typical latency (CPU) |
|---|---|---|---|
/api/analyze |
POST | Full 23-engine report for text or url |
2-8 s |
/api/quick-score |
POST | 4 classifiers plus heuristics | 0.1-0.5 s |
/api/paragraph-score |
POST | Score per paragraph (heat map) | 1-3 s |
/api/scan/site |
POST | AI app builder fingerprints plus a copy score | 1-3 s |
/api/extract |
POST | Text from an uploaded .pdf / .docx / .txt (multipart file) |
< 1 s |
/api/scan/snippets |
POST | Batch of 1-30 short snippets | ~0.5 s |
/api/scan/urls |
POST | Batch of 1-10 URLs, page-type aware | 1-5 s |
/api/engines |
GET | Engine metadata | instant |
/api/report/{id} |
GET | A stored report | instant |
/api/report/{id}/feedback |
POST | Record who actually wrote the text: {"label": "human" | "ai" | "mixed" | "unsure"} |
instant |
/api/queue/status |
GET | Queue capacity | instant |
curl -X POST http://localhost:8000/api/analyze \ -H "Content-Type: application/json" \ -d '{"text": "Your text to analyze here..."}'
From Python, examples/python_client.py analyses a
text, prints the five engines scoring highest and runs a site check, waiting
in the queue when the server is busy:
python examples/python_client.py "Paste at least 50 characters of text here..." example.comThe response lists every engine with its score, verdict and a plain-language
detail line, plus overall_score (0-100) and overall_verdict.
Detection Engines
Every engine links to its page on sloptotal.com, which carries its measured scores against both corpora. AUC below is the probability the engine ranks a random AI passage above a random human one: 1.0 is perfect, 0.5 is a coin flip.
Neural Classifiers
| Engine | Model | AUC | Notes |
|---|---|---|---|
| Desklib DeBERTa | DeBERTa-v3-large (435M) | 1.000 | Strongest separation in our own tests |
| SuperAnnotate | RoBERTa-large (355M) | 0.989 | No measurable bias against archaic prose |
| E5-Small | E5 + LoRA (33M) | 0.999 | Matches far larger models at 33M params |
| TMR Detector | RoBERTa-base (125M) | 1.000 | RAID-trained, so RAID scores flatter it |
| BERT-tiny RAID | BERT-tiny (4.4M) | 1.000 | Answers in milliseconds |
| ReMoDetect | DeBERTa (184M) | 0.941 | Targets RLHF-aligned LLMs |
| ChatGPT Detector | RoBERTa-base (125M) | 0.829 | ChatGPT-specific |
| Fakespot | RoBERTa-base (125M) | 0.999 | Accurate on modern text, but +0.533 bias on pre-1920 prose |
| OpenAI Detector | RoBERTa-base (125M) | 0.771 | The 2019 GPT-2 detector; weaker on modern LLMs |
Statistical Methods
| Engine | Method | AUC |
|---|---|---|
| Log-Rank | Average log-rank under GPT-2 | 0.909 |
| GLTR | Token rank distribution | 0.904 |
| Perplexity | GPT-2 perplexity scoring | 0.901 |
| Cross-Perplexity | Two-model perplexity comparison | 0.891 |
| Fast-DetectGPT | Conditional probability curvature | 0.890 |
| Binoculars | Cross-entropy ratio between two LMs | 0.836 |
| DivEye | Surprisal diversity | 0.730 |
Linguistic Heuristics
| Engine | Signal | AUC |
|---|---|---|
| Structural Analysis | Em-dash usage, sentence uniformity | 0.836 |
| Linguistic Markers | AI-preferred phrases ("delve", "tapestry"...) | 0.713 |
| Formulaic Patterns | Cliche openings and closings | 0.698 |
| Vocabulary Richness | Type-token ratio, hapax legomena | 0.583 |
| Readability Uniformity | Cross-paragraph consistency | 0.581 |
| Burstiness | Per-sentence perplexity variance | 0.582 |
| Sentiment & Hedging | Hedging and forced balance | 0.522 |
The linguistic heuristics are weak on their own. They are kept because they fail independently of the neural classifiers, which is what makes them useful as tiebreakers rather than as evidence.
Scoring
The final score is calibrated, not a simple average, and every weight is derived from measurement rather than intuition. See tests/eval/FINDINGS.md and sloptotal.com/detect/ai-detector-ensemble/.
- Anchored on the unbiased classifiers -- Desklib, SuperAnnotate, E5 and ReMoDetect all score high AUC with no measurable bias against older prose. Their consensus is blended 60/40 with the full weighted set.
- Weights from measurement -- each engine's share is proportional to Somers' D (2*AUC - 1), scaled down by any bias it shows against archaic writing. RAID-trained engines are damped because our corpus is RAID.
- Confidence from agreement -- a tight cluster across independent engine families is trustworthy; one confident engine is not.
- Skepticism, but only when earned -- unanimous high classifier scores are damped only when the text itself carries human markers (contractions, first-person, slang). Applied unconditionally it fired on 69 of 70 AI samples and 0 of 66 human ones, suppressing correct detections.
Fakespot was previously the anchor, weighted 0.13. It is accurate on modern text (AUC 0.999) but scored pre-1920 human prose at 0.645 against 0.112 for modern human writing -- the largest bias of any engine -- and anchoring amplified it. Machiavelli scored 62.5. After demotion to 0.033, literary passages average 10.2 and none is flagged.
Configuration
SlopTotal detects CPU, RAM and GPU at startup and picks a profile. Everything
can be overridden with environment variables; see .env.example.
| Profile | RAM | CPU | GPU | Notes |
|---|---|---|---|---|
| Lite | 4 GB | 2 cores | None | All engines, slower |
| Standard | 8 GB | 4 cores | None | Default for most laptops |
| Performance | 16 GB+ | 6+ cores | CUDA optional | Pool replicas, max throughput |
High-RAM CPU servers (e.g. 64 GB, no GPU): you automatically get the performance profile. With no CUDA, all inference stays on CPU but you can run more concurrent workers and model pool replicas:
# Tune for a 64 GB CPU-only server export SLOPTOTAL_PROFILE=performance export SLOPTOTAL_TORCH_THREADS=8 export SLOPTOTAL_FULL_WORKERS=8 export SLOPTOTAL_SNIPPET_WORKERS=6 export SLOPTOTAL_MAX_CONCURRENT_FULL=4 export SLOPTOTAL_POOL_FAKESPOT=2 export SLOPTOTAL_POOL_TMR=2 ./scripts/start.sh
| Variable | Default | Purpose |
|---|---|---|
SLOPTOTAL_PROFILE |
auto | lite, standard or performance |
SLOPTOTAL_RETENTION_DAYS |
30 |
Delete reports after N days (0 keeps them) |
SLOPTOTAL_ALLOW_PRIVATE_URLS |
off | Let URL scans reach private or intranet hosts (blocked by default) |
HF_HOME |
./models |
Where model weights are cached |
FAQ
Can AI detectors be trusted? Not blindly. No detector, this one included, should be the only evidence for an accusation. That is why SlopTotal shows all 23 votes, how much they agree, and its measured false-positive rate. Short text (under about 80 words) and heavily edited AI text are unreliable for every detector.
Does it detect ChatGPT, Claude, Gemini, Llama and Mistral? The evaluation corpus includes GPT-4, ChatGPT, Llama and Mistral output. The classifiers were trained on a wider mix. Newer models are covered as far as they share those fingerprints; the evaluation harness lets you measure any model you care about.
Will it flag classic literature or formal writing? Not in our tests: none of the 26 passages from Austen, Melville, Kafka, Machiavelli and others is flagged. Pre-1920 prose is part of the evaluation precisely because naive detectors fail on it.
Is my text stored or shared? It is processed on the server you run. Reports are kept for 30 days (configurable) so report links work, and nothing is sent to an outside service.
Can it detect AI-generated code? No, and we do not claim it can: in testing the engines never flagged human code but never caught machine-written code either. The Site check reports which AI app builder produced a website, which is a different question.
Troubleshooting
| Symptom | Fix |
|---|---|
TypeError: unsupported operand type(s) for | at startup |
Python 3.9 or older; use 3.10+ |
An engine reports Model loading failed |
Check disk space and network for the first model download, then restart; python scripts/smoke_test.py shows which engine fails |
| First scan is slow | Models are loading; later scans take seconds |
| A URL scan says "private network address" | Intended; set SLOPTOTAL_ALLOW_PRIVATE_URLS=1 to scan intranet pages |
Project layout
app/ FastAPI backend: engines, ensemble, site fingerprints, API
web/ The web UI (Jinja2 templates, vanilla JS, no build step)
tests/ Unit tests (seconds, no downloads) and tests/eval/ accuracy harness
scripts/ smoke_test.py (end-to-end) and the model drift check
benchmarks/ Speed and load scripts
ARCHITECTURE.md covers the internals. AGENTS.md is a short brief for contributors and AI coding assistants.
Newer open detectors we measured
Detector models keep appearing on Hugging Face, each with its own accuracy claim. Before adding any, we score them on the same two corpora. September 2026, standalone, 180 RAID texts plus the 26 literary passages:
| Model | RAID AUC | Literary bias (lower is better) | Status |
|---|---|---|---|
| Gradient (DeBERTa-v3-large) | 0.998 | 0.033 | Next engine to add |
| Vanguard (ModernBERT-large) | 0.998 | 0.037 | Candidate; poorly calibrated at 0.5 |
| Earlybird-fast (82M) | 0.913 | 0.040 | Candidate for fast snippet scans |
| rasbt ModernBERT | 0.864 | 0.000 | Not added |
RAID-trained models score near 1.0 on RAID by construction and need a different test set first. The full table, the models we excluded and why, and the raw scores are in tests/eval/FINDINGS.md. The roadmap is in TODO.md.
Related projects and reading
- RAID benchmark (ACL 2024): adversarial AI text detection dataset used in our evaluation
- Detecting the Machine (2026): cross-architecture detector benchmark; ensembles beat single detectors
- EditLens: human vs AI-edited vs AI-generated classification
- GLTR: visual token-rank inspection, the inspiration for our GLTR engine
- distil-labs/distil-ai-slop-detector: a 270M Gemma detector that runs in the browser
- sloptotal-extension: the Chrome extension
Contributing
Contributions are welcome, especially new engines with measurements. Start with CONTRIBUTING.md.
If SlopTotal is useful to you, a star helps other people find it.
License
MIT. Model weights keep their own licenses; see THIRD_PARTY_LICENSES.md.
来源:Hacker News · AI · github.com


