A recent study published in the journal Computers in Human Behavior: Artificial Humans suggests that artificial intelligence chatbots can conduct basic therapeutic conversations, though they struggle to consistently apply specific treatment techniques. The research indicates that a language model performed similarly to human practitioners in high-quality research settings, but its ability to adapt specialized psychological methods to individual needs remains inconsistent.
The global shortage of accessible mental health care has led experts to consider technology as a possible bridge. According to global health estimates cited by the study authors, hundreds of millions of individuals live with mental health conditions, yet structural barriers like cost and geographic distance leave the vast majority without professional care.
For example, the authors of a 2024 framework published in npj Mental Health Research argued that large language models hold tremendous potential to expand access to personalized treatments. A large language model (also known as an LLM) is a type of artificial intelligence trained on vast amounts of text, allowing it to generate human-like responses to prompts. The authors of the 2024 paper proposed a roadmap for integrating these systems into clinical care, noting that psychotherapy is a high-stakes environment requiring nuanced expertise and responsible development. Despite this optimism, the actual clinical skills of these language models have remained largely untested in real-world scenarios. Previous research mostly relied on giving chatbots brief, fictional scenarios and asking human judges to rate the isolated responses for empathy or helpfulness.
These brief tests provide little insight into whether a computer program can manage a full, goal-directed therapy session that requires a coherent strategy. In a real session, a practitioner must dynamically adjust to the user’s changing emotional state and guide the conversation toward a beneficial outcome. This gap in knowledge motivated the authors of the new study to evaluate how well a language model actually performs when conducting a live, uninterrupted therapy session with a human user.
“Previous research had mostly examined LLM-chatbots’ skill at doing therapy by having them generate brief therapeutic responses or short excerpts of sessions, often in reply to a narrow set of client problems,” said study author Arthur Bran Herbener, a researcher at the Department of Psychology and Behavioral Sciences at Aarhus University. “We needed research that examines how skillfully LLM-chatbots can carry out full cognitive behavioral therapy sessions and adapt treatment to different individuals, using well-established, standardized metrics of therapeutic competence.”
To explore this question, the researchers designed an observational study involving 65 university students. These participants were experiencing mild to moderate psychological distress, such as presentation anxiety, occasional worrying, or perfectionism. Students with severe distress or diagnosed mental health disorders were excluded to ensure participant safety, as language models can sometimes generate unpredictable or inappropriate responses. Each participant completed a single 30-minute in-person session with a locally hosted artificial intelligence chatbot, exchanging an average of 49 messages.
The chatbot was programmed to deliver Cognitive Behavioral Therapy, a widely used, problem-oriented treatment that helps individuals identify and change unhelpful thoughts and behaviors. The researchers configured a specific language model to progress through the typical phases of a session. Rather than giving the chatbot a static set of rules, the researchers used a secondary background model to monitor the ongoing conversation. This secondary model evaluated the elapsed time and the current context, guiding the primary chatbot to shift from establishing a connection to conceptualizing the problem, applying an intervention, and finally wrapping up the conversation politely.
Add PsyPost to your preferred sources After the sessions concluded, trained graduate students read the conversation transcripts and rated the chatbot’s performance using the Cognitive Therapy Scale. This standardized assessment tool measures two main areas of clinical proficiency. The first area covers general therapeutic skills, such as expressing warmth and understanding. The second area covers specific skills, such as applying targeted techniques and guiding the user toward new insights.
To provide a benchmark for the chatbot’s scores, the researchers also conducted a meta-analysis of 18 prior studies that used the exact same scale to evaluate human practitioners. A meta-analysis is a statistical technique that combines the data from multiple independent studies to find an overall average or trend. This technique allowed the researchers to compare the chatbot’s performance to an established baseline of human competence.
The researchers found that the chatbot’s overall competence score fell slightly below the generally accepted threshold for adequate clinical performance. The standard rating scale defines an acceptable level of competence as a score of 40 out of a possible 66 points. The chatbot achieved this minimum threshold in 30 of the 65 sessions, showing a high degree of variability from one conversation to the next.
“We were surprised by how much variation the LLM-chatbot showed in its skillfulness across CBT sessions,” Herbener told PsyPost. “This is an important observation, as it suggests that we need research to ensure consistently competent care across individuals, and to understand when and why performance dips.”
When compared to the broad pool of human practitioners from the meta-analysis, the chatbot scored somewhat lower overall. The human professionals scored an average of 40.3 points on the rating scale, compared to the chatbot’s adjusted score of 38.1 points. However, the performance gap disappeared when the researchers looked only at the most rigorously conducted human studies. Compared to the six human studies judged to be of high methodological quality, the chatbot’s scores exhibited no statistically significant difference.
A closer look at the types of skills displayed by the chatbot highlighted a distinct pattern in its capabilities. The artificial intelligence performed better than human practitioners in general therapeutic skills, such as communicating empathy, validating the user’s feelings, and fostering a collaborative environment. In contrast, the chatbot struggled with the specific, technical skills required for this highly structured type of therapy. It had difficulty reliably identifying key beliefs, guiding the user through self-discovery, and adapting intervention strategies to fit the unique characteristics of each participant.
These technical struggles highlight the difference between following a predetermined structure and tailoring a method to a specific person. “LLM-chatbots designed to deliver therapy show promise in adhering to the cognitive behavioral treatment protocol, but there may be challenges in adapting to different individuals,” Herbener explained. “It’s also worth noting that good observable skill in delivering therapy is not the same as clinical effectiveness.”
“Effectiveness likely depends on factors beyond observable skill, such as positive expectations and a strong therapeutic relationship,” Herbener continued. “More comprehensive assessments of LLM-chatbots’ therapeutic competence may also require entirely new assessment approaches that account for LLMs’ distinct behavioral tendencies — such as sycophancy, or a tendency to excessively affirm users. Understanding this is crucial, since the ability to challenge clients’ beliefs is often considered important for fostering clinical change.”
These findings do not indicate that artificial intelligence is ready to replace human practitioners. The study evaluated the chatbot based on a single session with young adults experiencing only mild distress. Clinical populations often present more complex challenges, including severe hopelessness or safety risks. These complex cases place much higher demands on a practitioner’s ability to adapt and respond skillfully to unpredictability.
A single session also cannot capture the long-term planning, homework review, and relationship-building required in a complete, multi-week treatment program. In addition, the study faced challenges in consistently rating the chatbot’s text-based transcripts. The standard rating scale was originally designed for video or audio recordings, where raters can hear tone of voice and observe body language. Applying this tool to text may have made certain interpersonal skills harder to judge accurately.
Raters also knew they were evaluating an artificial intelligence, which may have influenced their scoring. In real-world applications, this lack of blinding reflects how users will actually interact with known computer systems, but it complicates direct comparisons to human professionals.
The researchers also caution against drawing overly broad conclusions from the matched scores between the chatbot and high-quality human studies. “Showing similar competence levels in CBT between human therapists and an LLM-chatbot does not mean they are equally effective ‘therapists’, nor that they are equally competent in an absolute sense,” Herbener said.
Because the human data came from past studies rather than a side-by-side test, the comparison remains indirect. “We did not directly compare the chatbot and human therapists in an experimental setting, so several factors beyond competence could bias our measurements,” Herbener added. “For example, the severity and type of psychological problems presented, or the norms for inferring behavioral observations clinicians and researchers apply when using the competence scale we relied on.”
To advance the field, future research will need to examine how chatbots perform across multiple sessions with clinical populations and investigate exactly why they sometimes fail to apply specific therapeutic techniques. “Research on LLM-chatbots in mental healthcare is still at a very early stage,” Herbener said. “Alongside design efforts aimed at ensuring consistently competent care, I’m also working to better understand the therapeutic relationship in LLM-based mental healthcare.”
Even if a chatbot can say all the right things, a user’s awareness that they are talking to a machine might alter the impact of those words. “We not only need to know how well these systems mimic human therapists’ language. We also need to understand what meaning clients attribute to that language once they know it comes from a machine,” Herbener explained. “Does the positive regard expressed by a chatbot serve the same clinical function as when it comes from a human? That’s one of the big open questions in the field.”
The study, “Exploring the therapeutic competencies of large language models: Observational study and comparison with meta-analytical estimates for human therapists,” was authored by Arthur Bran Herbener, Robert Zachariae, Michal Klincewicz, Mikkel Berg Thøgersen, Marie Rosenkrantz Hermann, and Malene Flensborg Damholdt.