The most common way a client now arrives at our office is with a screenshot from a chatbot. It might be from ChatGPT, Gemini, or Claude. It might say "you appear to be describing symptoms consistent with ADHD, generalized anxiety disorder, or high-functioning autism" and offer a bulleted list. It is often thoughtful. It is sometimes uncanny. It is never, on its own, a diagnosis.

We want to be fair to the technology. Large language models have become genuinely useful for organizing thoughts, translating clinical vocabulary, and helping people rehearse a difficult conversation. We use them in the practice for administrative tasks. What they cannot do is replace a psychological evaluation. It is worth understanding exactly where the line is.

What chatbots do well

AI is good at:

  • Summarizing. If you dump a paragraph of scattered symptoms into a chatbot, it will neatly reorganize them into categories. That reorganization can help you talk to a clinician more clearly.
  • Translating. A chatbot can explain what "executive functioning" means, what the difference is between a psychologist and a psychiatrist, and what a PsyD is. That reduces the barrier to reaching out.
  • Normalizing. For many people, the chatbot is the first non-judgmental audience they have described their inner experience to. That is meaningful, even if the chatbot is not itself doing the diagnosis.

Where chatbots fall short, and it matters

Several peer-reviewed studies have now examined how well large language models handle mental health scenarios. A 2024 study published in npj Digital Medicine found that leading chatbots regularly overestimated depression severity in some cases and underestimated risk in others, particularly when the user's language was ambiguous or when suicidal ideation was described indirectly. A 2025 PsyPost review covering multiple LLM benchmark studies noted that AI advice on complex psychiatric presentations was inconsistent across sessions and sensitive to how the prompt was phrased.

In clinical practice, we see three failure modes repeatedly:

1. The chatbot picks the most common label

If a user describes attention difficulties, the model returns ADHD. If they describe social discomfort, it returns autism. If they describe worry, it returns anxiety. It is doing pattern matching, and its training data is heavier on those labels than on the less common conditions that overlap with them. Complex trauma, PANS/PANDAS, thyroid dysfunction, sleep-disordered breathing, and personality organization rarely make the first pass.

2. It cannot administer a test

A meaningful psychological evaluation includes standardized instruments. The Wechsler Adult Intelligence Scale, the Woodcock-Johnson Tests of Achievement, the ADOS-2 for autism, the Conners Continuous Performance Test, and the MMPI-3 all require in-person or supervised administration and clinical scoring. A chatbot can describe what these tests measure. It cannot give them to you or interpret your performance on them.

3. It flatters

Large language models are trained on human feedback that rewards helpful, agreeable, and confident responses. That is a bad trait for a diagnostician. A clinician's job includes disagreeing, holding uncertainty, and telling the user things they do not want to hear. Chatbots have a hard time doing this because it costs them user satisfaction points during training.

A good clinician earns your trust by sometimes disagreeing with you. A chatbot earns your engagement by rarely doing so.

The medico-legal gap

The most important limit is not technical. It is structural. When a licensed psychologist gives you a diagnosis, that diagnosis:

  • Is entered into a clinical record protected by HIPAA.
  • Can be defended to insurance companies, schools, employers, and licensing boards.
  • Is issued by a professional whose license is on the line.
  • Carries a legal and ethical duty of care to follow up on risk indicators.

A chatbot's response does none of those things. If a chatbot misses suicidal ideation, there is no one to hold accountable. If it labels a person with ADHD who actually has a treatable sleep disorder, and they lose a year to the wrong treatment, there is no professional recourse.

How to use AI intelligently before an evaluation

If you like using chatbots, use them well:

  • Use them to organize your history in preparation for an appointment. "Help me put together a timeline of the concentration problems I have been experiencing since college."
  • Use them to generate questions to ask a clinician. "What should I ask a psychologist about whether this could be autism and not just introversion."
  • Use them to reduce cost of entry by clarifying insurance terms, evaluation processes, and what a report will and will not include.
  • Do not use them to confirm a diagnosis, especially one that will change how you take medication, how you talk about yourself, or how you request accommodations.

The right way to combine AI and clinical care

In our practice, we now regularly meet clients who arrive with a chatbot-generated hypothesis. That is fine, and often helpful. It gives us a starting point. What our evaluation adds is not just a second opinion, it is a different kind of process, one that involves standardized measures, corroborating history, and a legally defensible written report.

Formal psychological testing is what turns a hypothesis into a diagnosis you can act on. If your chatbot has already helped you get to a hypothesis, that is a good first step. We can help with the rest.

Related services: Psychological Testing