Ongoing
Testing large language models for mental health assessment
Deciding whether something a person describes counts as a traumatic event turns out to be strikingly unreliable, even between trained clinicians. That is a problem for research and for treatment, and it is the kind of judgement a language model might be able to make consistently. We test whether they can, including models small enough to run on a local machine so that no clinical data ever leaves the building, and we report where they fail as carefully as where they succeed.