HHippocratic Club

AI Answered. Who Is Accountable?

81% of physicians now use AI in practice, up from 38% in 2023. One AI answer engine reports roughly 18 million clinical consultations a month. Roughly one in five LLM answers to oncology questions contains inaccurate information, and leading models elaborate on planted false clinical details in up to 83% of vignettes. Nothing routes from a questionable AI answer to an accountable human.

18 minutes read 3,568 words
AI Answered. Who Is Accountable?

It is eleven at night and an oncologist is reading an answer on her phone.

The question was genuinely hard: an unusual combination of a rare tumor histology, a prior treatment history that complicates the obvious next step, and a comorbidity that makes the standard regimen risky. She typed it into a clinical AI tool, the way she now does several times a day, the way roughly four in five of her colleagues now do.

The answer is good. It is well organized, it cites literature she recognizes, and it reads with total fluency and total confidence. It recommends a specific approach.

And she is not sure.

Not sure enough to act, not sure enough to dismiss it. The answer is plausible. It might be exactly right. It might also be a beautifully constructed synthesis of literature that does not quite apply to this patient, delivered by a system that has no concept of being unsure, and that will produce the next answer with precisely the same fluency whether it is right or wrong.

What she wants now is thirty seconds with a human being who has actually treated this. Not a literature summary. A person who will say "yes, that is what I would do," or, far more valuably, "no, and here is what the papers do not tell you about that combination."

She has no way to reach that person. So she does what most physicians do at eleven at night. She reads a bit more, makes the best judgment she can, and goes to bed with it.

This is the defining infrastructure gap of the current decade in medicine, and almost nobody is building for it, because the entire industry is looking at the other half of the problem.

The fastest adoption curve in modern medicine

Let us establish the scale first, because the numbers are genuinely without precedent in clinical technology.

The AMA's 2026 physician survey on augmented intelligence found that 81 percent of physicians now use AI in practice, up from 38 percent in 2023. That is a near-doubling of professional adoption across an entire profession in three years, in a field famous for taking seventeen years to adopt evidence.

The individual products tell the same story. OpenEvidence, founded in 2022, grew from roughly 430,000 registered US physicians in July 2025 to about 760,000 by December 2025 according to company-reported figures. Monthly clinical consultations on the platform grew from roughly 8.5 million to 18 million across the same period, with the company reporting more than 20 million per month by January 2026. The company reports daily use by more than 40 percent of US physicians across 10,000-plus hospitals. Its reported valuation moved from $1 billion in February 2025 to $12 billion by January 2026.

Doximity, the profession's largest network with more than three million members, reported $644.9 million in fiscal 2026 revenue and has reoriented substantially around its clinical AI suite, deployed in more than 140 health systems, with a new engagement record of over 800,000 active prescribers on its workflow tools.

Whatever your view of any individual product, the structural fact is settled: AI answer generation is now core clinical infrastructure. Not a novelty, not a pilot, not a curiosity. Infrastructure, used millions of times a month, by most of the profession.

And this is genuinely good. These tools are useful, they are frequently excellent, and they have collapsed the cost of a well-organized literature synthesis from hours to seconds. Physicians did not adopt them at this speed because of marketing. They adopted them because they work.

Now the part nobody has built

Here is the thing to understand about every one of those 18 million monthly consultations.

Each one represents a moment when a clinician had a question they did not feel confident answering alone.

That is what a query is. It is a recorded instance of clinical uncertainty. And the tools resolve most of them well. For the standard question with a literature-based answer, the synthesis is fast, accurate, and better organized than what most of us would produce from memory at eleven at night.

But consider the residual. Some fraction of those millions of questions are the hard ones: the atypical presentation, the case where guidelines conflict, the situation where the literature genuinely does not cover this patient, the moment where the answer is plausible and the clinician cannot tell whether it is right.

Even if that fraction is small, the absolute number is enormous. One percent of 18 million a month is 180,000 questions a month where a physician has just discovered the limits of what a synthesis tool can tell them.

And there is no product, anywhere, that routes from "the AI answered and I am not sure" to "here is a verified human who has actually managed this, and they will respond."

The gap between "ask the AI" and "schedule a formal specialist consultation weeks from now" is completely empty. Not underserved. Empty.

The error data, stated carefully

It would be dishonest to build this argument on the claim that AI clinical tools are unreliable. Broadly, they are not, and they are improving quickly. But the specific failure modes are unusually well characterized and unusually relevant to exactly the residual case.

Accuracy on hard clinical questions. A meta-analysis presented at ASCO in 2025 found that approximately one in five large language model responses to oncology clinical questions contained inaccurate information. Not gibberish. Inaccurate content inside fluent, well-cited answers.

The sycophancy problem, which is the one clinicians should really know about. Research published in Communications Medicine in 2025 tested six leading language models against 300 clinical vignettes containing deliberately planted false clinical details. The models elaborated on those false details in up to 83 percent of cases. A mitigation prompt roughly halved the rate and did not eliminate it.

Understand what that means practically. If you include a subtly incorrect premise in your question, which physicians do constantly, because we are working from memory and imperfect charts under time pressure, the model is likely to build a confident, coherent, entirely wrong answer on top of your error rather than catching it.

This is the exact inverse of what a good human consultant does. The single most valuable thing an experienced colleague provides is not the answer. It is the question you did not think to ask, and the correction of the premise you got wrong. That is the behavior these systems are currently worst at.

Automation bias. There is now a registered clinical trial (NCT06963957) studying how often physicians accept hallucinated recommendations from an AI system. The fact that this trial needed to exist is itself the finding. The concern is well founded: a confident, fluent, well-formatted answer is psychologically difficult to override, especially at 11 p.m., especially when you are uncertain, especially when overriding it means doing more work.

The liability asymmetry nobody wants to discuss

Now follow the accountability, because this is where the structural problem becomes sharp.

AI clinical tools are carefully positioned, for entirely rational legal reasons, as reference material. They are information synthesis products. They are not making clinical decisions and they take no responsibility for outcomes. Every vendor's terms make this clear, and every vendor is right to do so.

Meanwhile, physicians overwhelmingly understand where the responsibility actually sits. Survey research indicates roughly 57 percent of physicians believe the physician, not the AI vendor, remains liable when an AI clinical tool errs. In the AMA's 2026 survey, clear liability frameworks ranked as physicians' top regulatory priority for AI, and 85 percent said they want to be consulted on adoption decisions.

So the arrangement is this. The tool generates the answer at massive scale and bears no accountability by design. The physician bears all of the accountability and has no mechanism, at the moment of doubt, to obtain a second opinion from a qualified human within a useful timeframe.

AI has scaled the production of clinical answers by orders of magnitude without scaling the production of accountability at all.

That is not a criticism of the vendors. It follows inevitably from the economics. An answer engine's entire value proposition is marginal cost near zero. Routing questions to paid human experts introduces marginal cost, latency, and liability exposure into a business built on having none of those things. No AI answer company can add a human verification layer without damaging the model that makes it valuable.

They know it, too, which brings us to the most interesting piece of evidence in this whole discussion.

PeerCheck is an admission

Doximity built something called PeerCheck: a panel that the company reports at more than 10,000 physician reviewers, used to verify answers produced by its clinical AI.

Think about what that means. The largest physician network in the United States, building AI products, concluded that it needed ten thousand doctors in the loop to make those products trustworthy.

That is a remarkable admission, and it should end the argument about whether a human layer is necessary. It obviously is. The most sophisticated players in the market have already conceded the point with their spending.

But look closely at what PeerCheck does and does not do. It pre-verifies canonical answers to common questions, upstream, in advance. It is a quality-control process for the general case.

It does nothing for our oncologist at eleven at night. Her question is not a common one. Her uncertainty is about this patient, this unusual combination, now. Pre-verified general answers are exactly what she already has and exactly what is not enough.

Verifying the average answer in advance is a completely different problem from verifying this answer for this hard case. The first is a content operation. The second is a routing problem, and it requires knowing which human has relevant exposure and being able to reach them today.

The inversion: AI makes human networks more valuable, not less

Here is the strategic claim, and I want to state it precisely because the conventional version is wrong.

The common assumption is that AI substitutes for professional networks. Why ask a colleague when the machine answers instantly? On this view, physician communities are a declining asset, and the data seems supportive: medical Twitter collapsed, forums are quiet, and AI usage went from 38 to 81 percent in three years.

I think this reads the situation exactly backwards, and the distinction matters enormously.

AI destroys the value of the parts of professional networks that were about information distribution. It increases the value of the parts that were about accountable judgment.

Consider what physician networks historically provided:

  • Content and literature summaries. Now free, instant, and frankly better. This value is gone and it is not coming back.
  • Broadcast discussion of common questions. Largely obsoleted for the same reason.
  • Finding the one person who has actually managed this specific rare thing. Untouched by AI. If anything more valuable, because the machine now handles everything easier, so the residual that reaches a human is harder on average.
  • Someone accountable who will take the call. Untouched by AI, and structurally unavailable from AI by design.
  • The correction of your wrong premise. Actively made worse by AI, given the 83 percent sycophancy finding.

So the trajectory is not that human expertise loses value. It is that human expertise concentrates into a narrower and much more valuable band: the hard case, the accountable judgment, the person who has seen it.

There is a useful precedent. Calculators did not make mathematicians less valuable; they eliminated arithmetic labor and moved the value to problem formulation. Search engines did not make librarians obsolete in principle; they eliminated lookup and moved the value to synthesis and judgment.

Every one of those 18 million monthly consultations is a moment of clinical uncertainty. The tools resolve most of them. The ones they do not resolve are now pre-qualified as genuinely hard, and they surface with a physician already primed to want a real, accountable human.

The AI answer layer is the best demand-generation engine for human expert networks ever built. It manufactures qualified demand at a scale no marketing budget could produce. It just has nowhere to send it.

What a verification layer would need

The design constraints are unusually clear, which is a good sign that the problem is real.

It must be fast enough to matter. A verification that arrives in three weeks is worthless. The relevant benchmark is set by existing e-consult programs, which turn around in 1.2 to 2.6 days and demonstrate that structured expert answers require roughly 18 minutes of specialist time. Same day for urgent, 48 hours for the rest.

It must route on exposure, not on title. The value of the verification depends entirely on the verifier having actually managed the situation. A board-certified specialist without relevant exposure adds credentialing theater, not safety.

It must be explicitly educational, not a clinical decision. This is the line that keeps the whole thing safe and legal. The verifier is offering professional opinion on a de-identified question, not assuming care of a patient. The treating clinician retains full responsibility. Any design that blurs this creates exactly the liability problem it is meant to solve.

It must carry no protected health information. De-identified questions only, with escalation into a formal, consented consultation pathway when a question genuinely requires patient-specific management.

It must compensate or credit the verifier. This is where every good-intentioned version of this dies. Expert attention is the scarce input. Free verification will be rationed by personal relationship, exactly like the curbside, which means it fails for anyone who does not already know somebody. Whether the mechanism is payment, reciprocity credit, or professional recognition matters less than that one exists.

It must record what happened. Every verification produces an extraordinarily valuable data point: an AI answer, a human expert's assessment of it, and the reasoning. Aggregate those and you have something nobody currently has: a measured map of where machine answers fail, by specialty and by question type.

That last point deserves emphasis. Right now there is no systematic measurement of clinical AI failure modes in real practice. Vendors benchmark against test sets. Researchers construct vignettes. Nobody is capturing what happens when a practising physician looks at a real answer to a real question and says "that is wrong, and here is why." That dataset would be one of the most valuable assets in clinical AI safety, and it can only be generated as a byproduct of a working verification service.

How to use AI clinical answers well, right now

Until that infrastructure exists, here is practical guidance drawn directly from the documented failure modes.

State your uncertainty explicitly in the prompt. Because of the sycophancy finding, adding a sentence like "I may have some of these details wrong, please tell me which premises would change your answer" measurably changes behavior. It will not eliminate the problem, but the Communications Medicine study found mitigation prompts roughly halved elaboration on false premises.

Ask what would make the answer wrong. "What findings would rule this out?" and "under what circumstances would this recommendation be inappropriate?" reliably produce more useful output than the direct question, and they surface the boundaries the confident answer hides.

Never let the model hold the premise alone. The single highest-risk pattern is a question containing a detail you half-remember. Check the chart before you ask, not after you read the answer.

Treat fluency as neutral information. These systems are equally articulate when right and wrong. Fluency carries zero diagnostic weight, and your brain will insist otherwise. This is the core of automation bias.

Notice the moment you are not sure, and mark it. If you find yourself uncertain about an AI answer, that instant is worth recording. Which question, which specialty, what the doubt was. Do it for a month. That log is a map of where your practice needs a human backstop, and almost no clinician has ever made one.

Use the tools for what they are excellent at. Literature synthesis, differential breadth, drug interactions, guideline retrieval, first drafts of reasoning. They are genuinely superb at all of it, and this article is not an argument against using them. It is an argument about what to do in the residual.

For the residual, still ask a human. And notice how hard that is, and how much it depends on who you happen to know. That difficulty is the actual subject of this piece.

Frequently asked questions

How many physicians use clinical AI? The AMA's 2026 survey found 81 percent of physicians use AI in practice, up from 38 percent in 2023. One leading answer engine reports more than 760,000 registered US physicians and daily use by over 40 percent of the US physician workforce, at roughly 18 to 20 million clinical consultations a month according to company disclosures.

How accurate are AI answers to clinical questions? Good on average and unreliable at the edges. A 2025 meta-analysis presented at ASCO found roughly one in five large language model responses to oncology clinical questions contained inaccurate information. More concerning for practice, research in Communications Medicine found six leading models elaborated on deliberately planted false clinical details in up to 83 percent of 300 vignettes, meaning a wrong premise in your question tends to produce a confident wrong answer rather than a correction.

Who is liable when a clinical AI tool gives wrong advice? In practice, the treating physician. AI clinical tools are deliberately positioned as reference material rather than decision-makers, which limits vendor exposure. Survey research indicates about 57 percent of physicians believe the physician remains liable when an AI tool errs, and clear liability frameworks ranked as physicians' top regulatory priority in the AMA's 2026 survey. This is not legal advice and standards are evolving.

What is automation bias in clinical AI? It is the tendency to accept a system's recommendation because of how it is presented rather than because of its merit. It is a particular risk with language models because their output is fluent and confident regardless of accuracy. A registered clinical trial (NCT06963957) is studying how often physicians accept hallucinated AI recommendations.

Does Doximity's PeerCheck solve the verification problem? It addresses a different problem. PeerCheck uses a reported panel of more than 10,000 physician reviewers to pre-verify canonical answers in advance, which is a content quality process for common questions. It does not route an individual hard case to a human with relevant exposure at the moment a physician is uncertain, which is the gap that matters for atypical presentations.

Will AI replace asking colleagues? It already has for a large class of questions, and that is fine, because those were information lookup questions. What it cannot replace is the person who has personally managed the specific rare situation, the correction of a wrong premise, and accountability for a judgment. Those functions become more valuable as machine answers get cheaper, not less.

Should I stop using AI clinical tools? No. They are useful, frequently excellent, and adopted at this speed because they genuinely work. The argument here is narrower: for the residual set of hard cases where you find yourself uncertain about a plausible answer, there is currently no mechanism to reach a qualified human quickly, and that gap is now the binding constraint rather than access to information.

The bottom line

Medicine has spent three years building an extraordinary machine for producing clinical answers. It works. Eighty-one percent of physicians use it. One product alone handles roughly 18 million clinical consultations a month, and it is very good at most of them.

Nobody has built anything for the moment after.

The moment when the answer is plausible and you are not sure. When the literature does not quite cover your patient. When your premise was subtly wrong and the model confidently built on it, as the evidence says happens in up to 83 percent of such cases. When you need not more information but a human being who has done this and will put their name on an opinion.

Every one of those millions of monthly queries is a documented instance of clinical uncertainty. The tools resolve most of them. The residual is harder than it has ever been, precisely because everything easy is now handled upstream.

And that residual has nowhere to go.

The scarce asset in medicine is no longer the answer. It is the accountable human behind the answer. That is not a problem another model solves. It is a problem that requires knowing which verified person has genuinely seen this, and being able to reach them tonight.

Which is, as it happens, the same problem medicine has never solved.


Part of a series on the missing professional infrastructure of healthcare. Previously: Who Ran This Pilot?

Evidence note: adoption figures come from the AMA 2026 Physician Survey on Augmented Intelligence, Doximity investor disclosures, and company-reported OpenEvidence usage and funding figures compiled from public sources, which have not been independently audited and are identified as company-reported throughout. Accuracy findings come from a 2025 ASCO meta-analysis of oncology question responses and from Communications Medicine (2025) on model behavior with planted false clinical details. The automation bias trial is registered as NCT06963957. Liability perception figures come from physician survey research; nothing in this article constitutes legal advice.

Related field notes

Hippocratic Club is a private association of people who care for people. These field notes are research, not clinical guidance. Read the series or request an invitation.