cd /news/artificial-intelligence/can-we-still-trust-disaster-social-s… · home › topics › artificial-intelligence › article
[ARTICLE · art-142251] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↓ negative

Can We Still Trust Disaster Social Sensing? Empirical Evidence on Detecting AI-Generated Social Media Posts

A study of 12,000 texts from nine disasters found that text-based AI detectors cannot reliably separate human-authored from AI-generated disaster social-media posts, with AUROC between 0.402 and 0.517 across fourteen frozen cross-family configurations and best prospective recall of just 3.6% at a calibration-derived low-false-positive operating point. OSM-Det reached AUROC 0.521 and 10.4% recall at a 6.7% false-positive rate, while a disaster-trained linear head hit AUROC 0.817 but a seven-feature surface classifier reached 0.784 on the H0-versus-A0 contrast, and neutralising surface asymmetries cut the head from 0.733 to 0.594. The authors conclude text-based detection is not reliable enough to serve as an operational trust gate and recommend multimodal claims, accountable sources, and other contextual evidence to safeguard trust in disaster social sensing.

by read1 min views1 publishedSep 30, 2026

arXiv:2609.35821v1 Announce Type: new Abstract: Disaster social sensing converts public social-media posts into evidence for situational awareness and humanitarian needs, but generative artificial intelligence (AI) can produce plausible messages that resemble eyewitness reports. This study investigates whether text-based AI detectors can reliably distinguish human-authored from AI-generated disaster posts. We construct a dataset of 12,000 texts organised into 3,000 matched semantic units from nine disasters: original human posts (H0), minimally LLM-proofread human posts (H1), factual AI-generated posts based on the same verified facts (A0), and affectively framed versions of those AI posts (A1). A separate 6,000-text corpus from 42 events supports model selection and threshold calibration. We evaluate OSM-Det, Fast-DetectGPT, Binoculars, and direct large language model (LLM) judges across five model families, then test disaster-domain calibration, a frozen-encoder linear readout, paired transformation sensitivity, and dataset artifact controls. Across fourteen frozen cross-family configurations, AUROC is 0.402-0.517 and the best prospective recall at a calibration-derived low-false-positive operating point is 3.6%; OSM-Det reaches AUROC 0.521 and 10.4% recall at a realised 6.7% false-positive rate. A disaster-trained linear head reaches AUROC 0.817, but a seven-feature surface classifier reaches 0.784 on the H0-versus-A0 contrast, and neutralising identified surface asymmetries reduces the head from 0.733 to 0.594. The head also separates A0 from A1 even though provenance is unchanged. The results show that text-based detection is not reliable enough to serve as an operational trust gate; multimodal claims, accountable sources, and other contextual evidence should be rested on to safeguard trust in disaster social sensing.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @osm-det 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/can-we-still-trust-d…] indexed:0 read:1min 2026-09-30 · —