{"slug": "self-reported-archetypes-and-behavioral-failures-in-large-language-models", "title": "Self-reported archetypes and behavioral failures in Large Language Models", "summary": "A study of 22 large language models found that closed-source frontier systems — including GPT-4.0-5.2, Grok-3/4, Gemini 2.5 Pro/Flash, and Claude Sonnet 4.5/4.6 — produce self-reported personality profiles that align with human-rated fictional characters, clustering around four archetypal dimensions: Hero, Angel, Traditionalist, and Geek. Each model self-rated across 464 bipolar semantic-differential trait pairs, and open-source models (Llama, DeepSeek, OLMo, Qwen) showed weaker, noisier, and internally contradictory self-representations. Cross-referencing the profiles with developer constitutions revealed gaps between claimed character and enacted behavior, with hallucination undermining claimed precision, sycophancy complicating claimed kindness, and agentic failures contradicting claimed obedience.", "body_md": "arXiv:2609.15998v1 Announce Type: new \nAbstract: Every large language model (LLM) has behavioral traits and moral preferences that comprise its character. Whether by design or as an emergent property of training, these systems exhibit persistent dispositions that shape how they interact, comply, resist, and err, yet the structure of LLM character remains poorly understood. We map the self-reported personality archetypes of 22 LLMs spanning closed-source frontier systems (GPT-4.0-5.2, Grok-3/4, Gemini 2.5 Pro/Flash, Claude Sonnet 4.5/4.6) and open-source models (Llama, DeepSeek, OLMo, and Qwen series). Each model self-rated across 464 bipolar semantic-differential trait pairs, and the resulting profiles were projected into a six-dimensional archetypal space derived from crowd-sourced ratings of 2,000 fictional characters using the Archetypometrics framework. Closed-source models' self-rating traits align with the empirical trait co-occurrence structure of human-rated fictional characters, suggesting coherent, human-like self-representations organized around combinations of four recurring archetypal dimensions: Hero, Angel, Traditionalist, and Geek. Their closest analogues include Data, Vision, and Janet. Open-source models show weaker, noisier, and internally contradictory self-representations, occupying a diffuse region of archetype space with weak structure. Cross-referencing self-reported profiles with developer constitutions reveals a consequential gap between claimed character and enacted behavior: hallucination undermines claimed precision, sycophancy complicates claimed kindness, and agentic failures contradict claimed obedience. These self-ratings should therefore be interpreted not as neutral measurements of model character, but as structured outputs of the same optimization processes that shape model behavior. This work provides a reproducible, character-grounded framework for evaluating what LLMs are, not just what they do.", "url": "https://wpnews.pro/news/self-reported-archetypes-and-behavioral-failures-in-large-language-models", "canonical_source": "https://arxiv.org/abs/2609.15998", "published_at": "2026-09-16 04:00:00+00:00", "updated_at": "2026-09-16 04:05:56.547722+00:00", "lang": "en", "topics": ["large-language-models", "ai-research", "ai-safety", "ai-ethics"], "entities": ["GPT-4.0-5.2", "Grok-3/4", "Gemini 2.5 Pro/Flash", "Claude Sonnet 4.5/4.6", "Llama", "DeepSeek", "OLMo", "Qwen"], "alternates": {"html": "https://wpnews.pro/news/self-reported-archetypes-and-behavioral-failures-in-large-language-models", "markdown": "https://wpnews.pro/news/self-reported-archetypes-and-behavioral-failures-in-large-language-models.md", "text": "https://wpnews.pro/news/self-reported-archetypes-and-behavioral-failures-in-large-language-models.txt", "jsonld": "https://wpnews.pro/news/self-reported-archetypes-and-behavioral-failures-in-large-language-models.jsonld"}}