A new generation of mental-health chatbots — Woebot, Wysa, Earkick, Therabot, Replika-in-therapy-mode, and direct ChatGPT conversations — is filling a genuine access gap. There's now even randomized clinical trial evidence that they work. They are not therapists, and they are not appropriate for everything that walks into a real therapist's office. The state of the evidence, and a practical division of labor.
Where the Evidence Stands
The most rigorous trial so far is the Therabot RCT (Heinz et al., NEJM AI, 2025): 210 participants with clinically significant depression or anxiety symptoms, four weeks of chatbot-delivered CBT-style therapy vs a control group. The treatment arm showed large symptom reductions — comparable in magnitude to in-person CBT trials at short follow-up. Subgroup signals suggested chatbots work well for people who would not otherwise access therapy (cost, time, stigma), and modestly for those already in treatment as adjunct practice.
Earlier-generation tools like Woebot and Wysa had smaller RCTs and meta-analyses showing modest benefit for depression and anxiety symptoms, with the strongest signals in self-guided CBT protocols. Pure-CBT chatbots typically beat pure-empathy chatbots on measurable outcomes. Symptom tracking, daily check-ins, and structured thought records add real value.
What's Genuinely Useful Right Now
- CBT-style thought records on demand — "Is this a cognitive distortion?" exercises are where these tools shine, especially late at night when a human therapist is unavailable.
- Mood and anxiety tracking — daily check-ins trend meaningfully; seeing patterns is half the intervention.
- Behavioral activation — get-out-of-bed prompts, scheduled small wins, "name one thing you'd enjoy today" exercises have measurable effects on depression.
- Sleep hygiene drills — guided wind-downs and CBT-I elements stack on basic sleep advice with real outcomes.
- Skills practice between sessions — for people already in therapy, chatbots as "homework partners" perform well.
- Anonymity-allowing first steps — for those not yet ready to call a clinician, the chatbots are a low-cost, low-stakes bridge.
What They Genuinely Can't Do
- Assess for suicide risk reliably across the board. Symptom-trigger responses vary by app and have failed in published audits. If the tool doesn't have a clear crisis protocol and a real human behind it, that's a hard limit.
- Diagnose — bipolar disorder, personality disorders, complex trauma, psychotic disorders — anything where the wrong frame causes real harm.
- Manage medication — initiation, taper, dose adjustments for SSRIs, stimulants, mood stabilizers require clinician judgment and lab monitoring.
- Substitute for safety planning in active crisis. If you're in crisis, contact a human (988 in the US, local crisis lines elsewhere).
- Real relationship — the therapeutic alliance is one of the strongest predictors of outcome in therapy, and it's exactly what an LLM can't replicate. For moderate-to-severe depression, complex grief, OCD, trauma — a real human still has the better evidence base.
How to Choose a Good One
- Look for published evidence in the specific population it's marketed for — depression, GAD, insomnia, alcohol use. Marketing without evidence is a red flag.
- Read the crisis protocol — if the tool doesn't know what to do when you say "I'm thinking about hurting myself," it shouldn't be marketed as a therapy product.
- Read the privacy disclosure — is your data used to train the underlying model? Is it sold? Can you delete it? These are the questions clinicians hear the most.
- Check whether a human is reachable — better apps route you to a real clinician when your inputs suggest risk.
- Be skeptical of "AI therapist" branding vs "CBT-I companion" or "anxiety skills trainer" — accurate framing is a feature, not a limitation.
A Practical Division of Labor
Use an AI companion for: building daily skills practice, mood tracking, sleep drills, low-stakes thought records, and between-session homework. Bring a human therapist (or prescriber) for: anything diagnosable, anything involving medication, anything involving active risk, anything involving trauma or complex grief, and anything that's gotten measurably worse over weeks despite your best efforts.
The two aren't competing tools — they're complementary. The most useful framing is: the chatbot is the daily exercise app; the human is the coach you see monthly. Neither alone is enough for serious presentations; both together are better than either alone.
The Bottom Line
AI therapy chatbots are now genuinely useful for low-to-moderate depression and anxiety symptoms, especially for people who wouldn't otherwise access care. The 2025 Therabot RCT gave them real clinical legitimacy — but they are still not therapists. If you have a question that starts with "Should I…," if you're pregnant or on psychiatric medication, or if you've had thoughts of self-harm in the past month, talk to a human. For everything short of that, well-built chatbots are a legitimate component of a layered routine.