cd /news/artificial-intelligence/a-survey-on-fake-review-detection-fr… · home › topics › artificial-intelligence › article
[ARTICLE · art-140734] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=· neutral

A Survey on Fake Review Detection: From Pre-trained Language Models to Large Language Models

A survey published on arXiv (2609.30292v1) reviews 211 studies on fake review detection from 2018 to early 2026, tracing the field from traditional machine learning and deep learning to pre-trained language model (PLM) and large language model (LLM) methods. The survey organizes the work by evidence source and fusion level — review text, sentiment, rating behavior, temporal metadata, user-product graphs, multimodal content, external knowledge, and LLM-generated signals — and analyzes reported performance trends on the Amazon, Yelp, and OpSpam benchmark families while flagging limitations from differing label construction, data splits, and evaluation protocols. It identifies open problems in adversarial generation, cross-domain transfer, uncertainty-aware fusion, missing-source robustness, interpretability, and trustworthy evaluation for AI-generated deceptive content.

read1 min views1 publishedSep 28, 2026

arXiv:2609.30292v1 Announce Type: new Abstract: Online reviews shape consumer decisions, platform governance, and corporate reputation.Fake reviews compromise this information channel by injecting deceptive evidence into rating systems, recommendation pipelines, and public trust mechanisms.The rise of large language models, or LLMs, has changed the problem in two directions.LLMs can generate fluent and context-aware deceptive reviews, while pre-trained language models, or PLMs, and LLMs also provide stronger semantic representations for detection.This survey reviews fake review detection from an information fusion perspective, covering 211 studies published from 2018 to early 2026.We organize existing work by evidence source and fusion level, covering review text, sentiment, rating behavior, temporal metadata, user-product graphs, multimodal content, external knowledge, and LLM-generated signals.We trace the development from traditional machine learning and deep learning to PLM-based and LLM-based methods, and examine how different approaches combine textual, behavioral, structural, and multimodal evidence.We also analyze reported performance trends on widely used Amazon, Yelp, and OpSpam benchmark families, while noting the limitations caused by different label construction procedures, data splits, and evaluation protocols.Finally, we identify open problems in adversarial generation, cross-domain transfer, uncertainty-aware fusion, missing-source robustness, interpretability, and trustworthy evaluation for AI-generated deceptive content.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @arxiv 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
→ Live at https://your-agent.zahid.host ✓
Get free account → Pricing
from €0/mo · no card required
LIVE [news/a-survey-on-fake-rev…] indexed:0 read:1min 2026-09-28 · —