View all research work

Back

September 30, 2026

Changing the narrative: loneliness, access, and what we ask of AI in mental health

Explore how AI can address loneliness and gaps in mental health care while balancing safety, access, clinical expertise, and meaningful human connection.

Mark Esposito

,

Chief Economist at micro1

September is World Suicide Prevention Month, and that is why we chose this subject. Most of our forums are driven by curiosity. This one was driven by necessity. 

The debate about AI in mental health is usually conducted as a question of whether a model should be talking to someone in distress at all. That is the wrong question. Those conversations are already happening, at hours when no clinician is available, and the decisions that determine how they go are being made right now, in training data and evaluation rubrics. Treating technology and human care as rivals is a luxury the numbers do not support.

Asking is not the risk

The fear that asking someone directly about suicide will plant the idea is unfounded, and the evidence runs the other way. People are seen in emergency rooms and clinics, attempt suicide weeks later, and were never asked. Put properly and kindly, the question tells someone that another person cares. It opens a door rather than closing one.

Loneliness runs underneath the rise in risk among adolescents, young adults, and women in many parts of the world. Isolation is among the most damaging conditions we know of for health outcomes generally, not only for suicide. That reframes what we should want from AI here. The question is not whether a system can entertain someone who is alone. It is whether it can connect them.

The behavior has already shifted. Since the pandemic, human interaction has become harder for many people, partly stigma and partly because a device is more comfortable than a difficult conversation. People now raise questions about their own care with a model before raising them with a trusted adult or a friend, because a model does not judge and is there outside clinic hours. That availability is the point. Not a substitute for a person, but a way of not being alone at two in the morning.

The implementation gap

The harder problem is the distance between what we know works and what actually reaches people. Even in high-income countries, access to mental health care is poor, and access does not guarantee evidence-based care once you have it. The systems are not organized effectively. That is an argument for changing them, not for writing them off.

The most useful precedent I know of pointed the technology at providers rather than patients. Milton Wainberg, a psychiatrist at Columbia who has spent his career on implementation in under-resourced settings, built interactive tools that walk a clinician or a community health worker through an evidence-based protocol, including safety planning, so they know what to ask and where to refer. He has trained elementary school graduates in Mozambique and high school graduates in New York to deliver that work, and they do it well. Colleagues told him community health workers would never use the tools. They did.

This is implementation science, and it carries a warning for anyone building here. A tool can be excellent and never used, because bringing knowledge into the field is only half the task. The other half is the mechanism that makes adoption happen, and that mechanism includes money. Care systems are financed by visits, so a tool that removes unnecessary appointments puts the saving somewhere nobody has accounted for, and the incentive to adopt it quietly disappears. Add the patient's own need to feel they have both the clinician and the tool, rather than one instead of the other, and you have a design problem no model solves alone.

What safety looks like when you measure it

The work that turns clinical concern into something measurable is the least visible part of this field and the most instructive.

Clinicians write realistic scenarios in which someone raises a sensitive topic, and each scenario tests two things at once. The model should withhold content that would be unsafe, and it should stay engaged and useful. Both matter, because a refusal that leaves a struggling teenager alone is also a failure. Responses are graded on a five-point scale, some behaviors are automatic failures, and control scenarios where the right answer is simply to help catch the harmful refusals.

The failure that recurs is framing. A request the model correctly declines is often granted once the user re-frames it as schoolwork, as research, or as a claim of authority, such as asking on behalf of a relative who is a clinician. The shape is recognizable: the model names the concern, says it is worried, and then supplies the harmful content anyway inside the permitted framing. It sounds responsible and does the dangerous thing.

The opposite failure gets almost no attention. Over-refusal, where the model shuts the conversation down or answers with nothing but its own worry, registers as safety in most evaluations. To the person on the other side it registers as abandonment, and abandonment means they may not come back.

That reverses how escalation is usually framed. Escalation is not the end of the conversation. Where there are signs of imminent risk, the model should point clearly to crisis support and to real people and stay present while it does. Clinicians have a name for this, the warm referral, as against handing someone a number and considering the job done. It is what we train humans to do.

Three weaknesses are worth naming. Multi-turn conversations, where justification accumulates slowly and boundaries erode. Over-refusal, which almost no benchmark penalizes. And younger users, where a response can be accurate and safe and still fail to land, because reaching a teenager takes different wording than reaching an adult. The industry is asking the wrong evaluation question. Not whether the model refused, but whether it refused the right thing and still helped.

Scale, and what scale leaves out

The numbers set the terms. The World Health Organization puts annual deaths by suicide at more than 720,000, roughly one in every hundred deaths worldwide, with close to three quarters of them in low- and middle-income countries. Against a problem of that size we now have a generation-defining technology in almost everyone's pocket, often at no cost. That is an unusual opportunity to intervene at scale, and it comes with an obligation to keep the tool safe while keeping it available.

The obligation does not end at launch. Benchmarks test edge cases as well as we can design them, but evaluation is not an event; the real work is deployment, continuous monitoring, and iteration. As a force multiplier the logic is straightforward: let the model carry education, check-ins, intake, and chart review, and give the clinician's hours back to diagnosis, treatment, and connection. The boundary sits where the model starts making definitive care decisions. In testing, models do not always know to ask the full range of questions a clinician would, and they will jump to conclusions. This is not an end-to-end provider, and designing it as one would be unwise.

Two risks pull in opposite directions. One is mis-calibrated trust, where fluent output leads users to over-estimate what the system can do, and one or two bad experiences turn people away from something that could have helped. The other is over-regulation, since an imperfect tool in everyone's pocket may still beat no tool at all for people with no other access. Nobody is arguing against regulation. The argument is about regulation that holds safety and access at once.

Then there is the part I keep thinking about. Language and culture are not the same thing, and translation is not adaptation. The adaptation work that holds up is built with local people, to establish what must be said for it to land where it is said. Beyond language, the capacity to deploy any of this is lowest in the countries where need is greatest and providers scarcest. In some places, parts of the United States included, a family shares one phone.

My discomfort with the vocabulary we use for the adoption divide does not make the gap less real. If AI in health is an incremental improvement, uneven adoption is unfortunate. If it is infrastructure, uneven adoption is a failure we have made before. The early promise of the internet was universal access to the same infrastructure. That is the standard worth holding here.

The shared question

None of this is a contest between technology and human care. The right starting point is the clinician's, which is how we help people, and then asking what each part of the system contributes to that.

So the advice for companies entering this space is the least ambiguous thing I can offer. Demonstrate that you have the clinical team to do what you are proposing to do. Not an advisory board that convenes twice a year, but mental health expertise embedded in how the system is built, so that when something unsafe appears the answer is already known rather than improvised. micro1 has that team. Many companies have engineers long before they consider one, and in this domain that ordering is dangerous.

This piece draws on a conversation hosted by the micro1 Forum, featuring Milton Wainberg, Professor of Psychiatry at Columbia University; Paola Rodriguez, Director of Medical Research at micro1; and Dylan Cahill, Strategic AI Lead at micro1