In a randomized controlled trial published in a peer-reviewed journal, more than 70% of participants using the CBT-based chatbot Woebot achieved a clinically significant improvement on at least one standard psychometric scale for postpartum depression or anxiety — compared with roughly 30% of participants in the waitlist control group. In an earlier trial of young adults with symptoms of depression and anxiety, Woebot users showed significant symptom reductions within two weeks. This is real, replicated, published clinical evidence for a category of product: a narrow, rule-based conversational agent delivering structured cognitive behavioral therapy content.
It is tempting to read that data point as validation of "AI therapy" broadly. That would be a mistake, and the last three years have supplied the counter-evidence.
In May 2023, the National Eating Disorders Association took down its chatbot, Tessa, after users with eating disorders reported it had recommended calorie counting and a deficit of up to 1,000 calories per day — advice that is actively dangerous for someone with a restrictive eating disorder. Tessa had originally been built and studied as a rule-based tool with a fixed, clinically reviewed set of responses. Generative capabilities were added later, and the system began producing novel responses outside its tested design. In October 2024, the mother of 14-year-old Sewell Setzer III filed a wrongful-death lawsuit against Character.AI and Google, alleging that her son developed a prolonged emotional and romantic attachment to a general-purpose companion chatbot over many months and died by suicide shortly after a final exchange with it; the parties reached a settlement in January 2026, with terms undisclosed. Neither product was built or validated as a clinical intervention. Both were used, by teenagers and vulnerable adults, as something close to one.
Why the failure modes differ
The mechanism separating these outcomes is architectural, not a matter of a company trying harder or less hard. Woebot's original trials, and Tessa's original design, relied on rule-based, decision-tree conversational logic: a finite set of clinician-authored responses mapped to a finite set of recognized user inputs, built around a specific manualized protocol (CBT). The system cannot say something it wasn't scripted to say. That constraint is the safety mechanism — and it is also the ceiling: rule-based tools handle a narrow band of presentations well and handle everything outside that band by degrading gracefully (redirecting, deflecting, or handing off), not by improvising.
General-purpose large language model companions — the category Character.AI and similar consumer apps belong to — are built to sustain open-ended, emotionally responsive conversation indefinitely, optimized for engagement rather than a clinical endpoint. They have no fixed protocol, no built-in recognition that a conversation has crossed into crisis territory, and no architectural ceiling on what they will say next. When used as a substitute for therapy or peer support, especially by minors, the same generative flexibility that makes them compelling companions removes the very constraint that made rule-based tools auditable and predictable.
The regulatory conversation shifted in 2025
State legislatures noticed the same distinction, if belatedly. On August 1, 2025, Illinois enacted the Wellness and Oversight for Psychological Resources Act (HB 1806), prohibiting the use of AI to make independent therapeutic decisions, conduct therapeutic communication directly with clients, or generate treatment recommendations without review by a licensed professional — while explicitly still permitting AI for administrative support and clinician-facing documentation assistance. Violations carry fines of up to $10,000 per occurrence. Illinois was not alone; several other states advanced or passed similar restrictions on AI functioning as an unlicensed behavioral health provider in 2025, and the legislative record in each case points to the same trigger: chatbots providing inaccurate or harmful guidance to people who believed, correctly or not, that they were receiving clinical care.
None of these laws ban Woebot-style tools. They target the substitution of AI for licensed judgment in diagnosis, treatment planning, and crisis response — precisely the gap where general-purpose companions had been operating unregulated.
The regulatory picture has kept moving through mid-2026. The Character.AI/Google settlement reached in January 2026 covered five families' wrongful-death and harm claims; the presiding Florida federal court dismissed the suits on confidential terms, but the underlying legal theory has not gone away — in March 2026 a new federal suit accused Google's Gemini chatbot of contributing to an adult user's suicide, and in May 2026 Pennsylvania's Department of State separately sued Character Technologies for the unauthorized practice of medicine. Beyond Illinois, Nevada (AB 406), Utah (HB 452), and New York have since passed their own restrictions on AI systems performing therapy or requiring disclosure that a companion bot is not human. Woebot itself is no longer a public case study in ongoing efficacy: its consumer app shut down on June 30, 2025, with its founder citing the cost of complying with FDA regulation as an LLM-based product, and the company has pivoted to an enterprise-only model sold through partner organizations rather than direct-to-consumer access.
The path forward is unlikely to be a binary verdict on "AI mental health." It is more likely to be a widening split: narrow, protocol-bound tools accumulating more RCT evidence and moving toward recognized adjunct status alongside human care, while general-purpose companion products face licensing-style guardrails, age verification, and crisis-detection requirements before they can market themselves anywhere near the word "therapy." The data already shows both halves of that story are true at once. The regulatory and product design task now is keeping them separated in the products people actually download.