When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha A new arXiv study evaluating therapy chatbots found that large language models from Claude, GPT-4o, and Llama-3.1 understand 76-82% of Generation Alpha mental health vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point gap absent in human therapists. The researchers recommend mandatory human-in-the-loop architectures and regulatory frameworks for youth-facing mental health AI. arXiv:2608.20345v1 Announce Type: new Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha Gen Alpha, born 2010-2024 , with 13.1% of U.S. adolescents 5.4 million using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated. Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: 1 64 Gen Alpha mental health expressions validated by native speakers ICC=0.72 and clinicians kappa=0.78 ; 2 75 multi-turn conversations 780 turns with paired Standard/Gen Alpha versions. Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point pp vocabulary-comprehension gap p<.001, d 0.48 absent in human therapists 3pp, p=.22 . The gap is architecturally consistent and widens with ambiguity 7pp - 18pp . We identify six failure patterns: sarcasm masking 29pp , minimization acceptance 43pp , informal style bias 24pp , risk-stratified ambiguity 19pp , semantic drift 19pp , context-dependent violence 7pp . Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance 6.4x cost . With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI.