cd /news/artificial-intelligence/evl-mcot-enhanced-vision-language-mu… · home topics artificial-intelligence article
[ARTICLE · art-74924] src=arxiv.org ↗ pub= topic=artificial-intelligence verified=true sentiment=↑ positive

EVL-MCoT: Enhanced Vision-Language Multi-CoT for Harmful Meme Detection

Researchers propose EVL-MCoT, an enhanced vision-language multi-chain-of-thought approach for harmful meme detection, achieving promising results on the HatefulMemes and MultiOff datasets. The method addresses limitations of existing dual-stream models by promoting multi-perspective thinking and using a prototype-guided decoding framework for finer visual-text alignment. Source code is publicly available at https://github.com/BGWH123/EVL-MCoT.

read1 min views1 publishedJul 27, 2026

arXiv:2607.22016v1 Announce Type: new Abstract: MEMEs are widely used on the internet and often carry strong elements of sarcasm or irony. Understanding their hidden meanings typically requires a joint interpretation of text and vision. Existing methods focus on the dual-stream vision-language model to extract the visual and text simultaneously, which lacks background information and prior knowledge about the comprehensive explanation of MEME. One feasible option is to adopt chain-of-thought (CoT). However, the simple CoT approach lacks multi-perspective thinking, which may compromise the reliability of the resulting answers. Moreover, it often relies on shallow feature fusion, lacking the fusion of local details and fine-grained visual-prompt text alignment. This limitation prevents a deeper understanding of the intricate connections between the visual and the text. Herein, an enhanced vision-language multi-CoT (EVL-MCoT) approach is proposed to address these limitations. By promoting multi-CoT, EVL-MCoT enhances consistency and reduces bias in the decision-making process. Additionally, we design a prototype-guided and context-guided decoding framework, which incorporates visual prototypes to guide the fusion process and enables the model to align textual and visual information more precisely. We achieve promising results on the HatefulMemes and MultiOff datasets. The source code has been publicly released and is available at https://github.com/BGWH123/EVL-MCoT.

── more in #artificial-intelligence 4 stories · sorted by recency
── more on @evl-mcot 3 stories trending now
sponsored brought to you by zahid.host 4,200+ EU-deployed projects
reading about agents? ship yours in a single git push.

Run your AI side-project on zahid.host

EU-based hosting, git-push deploys, automatic HTTPS, no cold starts. Free tier with a custom domain — perfect for shipping the agent you just read about.

$git push zahid main
Live at https://your-agent.zahid.host
Get free account → Pricing
from €0/mo · no card required
LIVE [news/evl-mcot-enhanced-vi…] indexed:0 read:1min 2026-07-27 ·