It is 1:40 in the morning and someone is typing a symptom into a box. This scene is old enough to have grandchildren. What changed is what happens next. Fifteen years ago the box returned a page of links — a forum thread from 2009, a Mayo Clinic page, something alarming from a site with a lot of ads — and the person at the keyboard had to sit there, half-literate in medicine, and weigh them against each other. It was miserable, but it was honest about its own limits. The ambiguity was visible. You could see the internet shrugging.
Now the box talks back. It produces a single, fluent, well-organized answer in the calm register of someone who has considered your case carefully. It has not considered your case carefully. It has produced text that resembles the output of careful consideration, which is a different thing, and the gap between those two things is where this whole subject lives.
The interesting shift in self-diagnosis is not technological, it is tonal. A results page forced the reader to judge; a chat interface absorbs the judgment and hands back a conclusion. Fluency is doing work that accuracy has not earned.
What the numbers actually say
It would be convenient to report that the new tools are simply worse than the old ones, or simply better. The evidence refuses to cooperate. A 2025 systematic review in npj Digital Medicine looked at self-triage tools for laypeople — apps and chatbots that advise where to seek care — and found accuracy ranging from 11.5% to 90% across individual systems, with study-level averages spread almost as widely. Large language models landed between 57.8% and 76% on triage accuracy. People deciding entirely unaided managed 47.3% to 62.4%.
Read those ranges again. The best tools are genuinely useful; the worst are worse than a worried person guessing; and the review’s own conclusion is that the tools should be neither universally recommended nor universally discouraged. Note also what was being measured: triage — should I go to the ER, the GP, or bed — not diagnosis. Only one included study even attempted to measure how accurately laypeople arrive at a diagnosis. The honest summary of the literature is that performance varies enormously and we know less than the interfaces imply.
Which is precisely the problem. A system whose real-world accuracy might be anywhere on an eighteen-point spread — and whose category as a whole ranges from 11.5% to 90% — is a system that should say “I don’t know” quite often. The New England Journal of Medicine ran a Perspective in 2025 asking, in its title, “Can AI Say ‘I Don’t Know’?” — and the fact that the question needs asking in a medical journal tells you the answer in consumer products is mostly no. Uncertainty is bad interface. It tests poorly. Nobody returns to an app that hedges.
The conversation runs both ways
The chatbot register doesn’t just change what comes out — it changes what goes in. A preregistered experiment with 500 participants, published in Nature Health, found that people who believed they were talking to an AI provided lower-quality symptom reports than people who believed they were talking to a physician. Same interface, same task; the mere belief about what was on the other end degraded the input. Users apparently calibrate their effort to their sense of the audience, and something about “AI” licenses vagueness — fewer details, less precision, the conversational equivalent of shrugging at a search box.
So the loop is worse than it looks. The tool speaks with more confidence than a results page ever did, while the user feeds it less information than they would give a doctor. The fluency flows in one direction and the carelessness in the other, and they meet in the middle to produce an answer that sounds authoritative and is built on a thinner foundation than either party realizes.
Meanwhile, the safety mechanism we’ve settled on is the disclaimer — the small type explaining that this is not medical advice. A study in Frontiers in Artificial Intelligence tested disclaimers across fifteen chatbot health scenarios and found their effects mixed and scenario-dependent; the authors state plainly that disclaimers “are not reliable safeguards against over-reliance.” Accuracy was the strongest predictor of trust, but tone barely mattered, because the chatbots were perceived as confident regardless of how they were phrased. Confidence, it turns out, is a property of the medium. You cannot disclaimer your way out of it.
The missing duty
Here is the material point, and it is not really about accuracy at all. A clinician who is wrong is accountable — to a licensing board, to a malpractice insurer, to a professional norm that treats error as a breach of duty. A consumer product that is wrong is covered by a terms-of-service document. A 2024 analysis in JMIR frames this as an unresolved question of professional responsibility: symptom-checker apps occupy the functional role of a triage nurse while bearing none of the obligations. They can even entrench existing biases in healthcare — performing worse for the populations already worst served — without anyone being on the hook for it.
This is the quiet trick of the category. The tools borrow the authority of medicine through tone while declining its liability through paperwork. The disclaimer is not a warning; it is a waiver, signed by someone at 1:40 in the morning who did not read it and would not have been offered an alternative if they had.
Fairness requires saying the other half. For the person who is uninsured, or rural, or rationing a doctor’s visit against rent, a decent triage tool is not a toy — it is the only consultation available, and “better than nothing” undersells it when the unaided alternative is a coin flip. The review’s finding that these tools shouldn’t be universally discouraged is not a hedge; it is the correct read. The problem was never that the tool exists. The problem is that it cannot say “I don’t know,” and that nobody has made it.
The old search results page, for all its misery, forced a moment of judgment — this looks serious, that looks like nothing, I should probably call someone. The chat interface removes that moment and calls the removal a feature. A tool that occasionally said “this is beyond me, and here is what a real answer would cost you” would be more honest, more useful, and completely unshippable. Which tells you who the product is actually for.