{"slug": "llms-are-failing-hard-at-spotting-social-stigma-in-group-chats", "title": "LLMs are failing hard at spotting social stigma in group chats", "summary": "A new research paper introduces SDARE-Bench, a benchmark for conversational stigma detection and response, and finds that eight large language models (LLMs) consistently fail to detect stigma in group chats, with stigma expression rates skyrocketing to 97.5% under simulated group pressure. The study, which built a dataset of over 1,100 dyadic queries and nearly 1,400 group dialogue scenarios, highlights that models become more stigmatizing in group contexts, undermining their reliability for roles like moderators or counselors.", "body_md": "# LLMs are failing hard at spotting social stigma in group chats\n\nA new research paper just introduced SDARE-Bench to answer exactly that. Instead of the usual one-off questions, this benchmark focuses on conversational stigma detection and response across two specific scenarios: dyadic (one-on-one) and group dialogues. The researchers built a massive dataset with over 1,100 dyadic queries and nearly 1,400 group dialogue scenarios to see if these models actually understand the nuances of social prejudice.\n\nThe results are pretty alarming for anyone working on AI safety or alignment. After putting eight different LLMs through the wringer, the researchers found a consistent pattern of failure.\n\n## The breakdown of the findings\n\nThe study looked at two main things: can the model *detect* stigma, and how does it *respond* when stigma is present? Here is what the data showed:\n\n**Detection Failure:** Across the board, LLMs were terrible at identifying the specific components of stigma within a conversation. They often missed the subtle ways stigma manifests in a back-and-forth exchange.**The Group Effect:** This was the most interesting (and worrying) part. When the conversation shifted from a private one-on-one chat to a group setting, the models' performance tanked. Stigma expression in the model's responses was significantly higher in group contexts.**Pressure Sensitivity:** In scenarios specifically designed to simulate \"group pressure,\" the stigma expression rate skyrocketed to a staggering 97.5% average.**Advice Quality:** Not only did the models struggle with social nuance, but their advice became increasingly unrealistic and biased when they failed to resist the stigmatizing tone of the conversation.\n\n## Why this matters for AI deployment\n\nIf you are building an LLM agent to act as a moderator, a counselor, or even a collaborative teammate in a workspace, you can't rely on standard safety filters alone. These filters are often trained on isolated text snippets. SDARE-Bench proves that \"socially complex\" environments act like a catalyst for bad behavior.\n\nThe fact that the models become more stigmatizing when they \"think\" they are in a group suggests that the training data or the way we prompt these models doesn't account for the psychological weight of an audience. We are essentially seeing a digital version of \"mob mentality\" where the LLM mirrors the perceived social norms of the group, even if those norms are toxic.\n\nFor anyone interested in the technical side of this, the researchers used a classifier trained on ,1392 human-annotated responses to evaluate the open-ended generations. This wasn't just a simple keyword match; it was a deep dive into how the models actually behave when the guardrails are tested by the complexity of the prompt. It highlights a massive gap in current prompt engineering and safety training: we need to move beyond \"don't say bad words\" and start teaching models how to navigate social dynamics without folding under pressure.\n\n[Next Visual inputs are bypassing LLM safety filters in ways that →](/en/threads/8491/)\n\n[an AI side-hustle playbook](https://tanyan888.com/), with plenty of directly applicable cases.", "url": "https://wpnews.pro/news/llms-are-failing-hard-at-spotting-social-stigma-in-group-chats", "canonical_source": "https://promptcube3.com/en/threads/8601/", "published_at": "2026-09-02 17:45:45+00:00", "updated_at": "2026-09-02 17:52:36.818398+00:00", "lang": "en", "topics": ["artificial-intelligence", "large-language-models", "ai-safety", "ai-research"], "entities": ["SDARE-Bench"], "alternates": {"html": "https://wpnews.pro/news/llms-are-failing-hard-at-spotting-social-stigma-in-group-chats", "markdown": "https://wpnews.pro/news/llms-are-failing-hard-at-spotting-social-stigma-in-group-chats.md", "text": "https://wpnews.pro/news/llms-are-failing-hard-at-spotting-social-stigma-in-group-chats.txt", "jsonld": "https://wpnews.pro/news/llms-are-failing-hard-at-spotting-social-stigma-in-group-chats.jsonld"}}