A chief medical information officer is two weeks from signing a seven-figure contract for a clinical system that several thousand clinicians will use every day for the next decade.
She has done the demos. She has read the analyst report. She has the security review and the integration assessment. And now she is on the last and most important step of the evaluation: the reference calls.
The vendor has supplied three customers. All three are enthusiastic. All three describe an implementation that went well. All three are, in the industry's own phrasing, "coached," and in some cases compensated for their time.
She knows the list is curated. Everybody knows the list is curated. The vendor knows that she knows.
And she has no way to reach the fourth customer. The one who is eighteen months in, past the honeymoon, who could tell her what broke in month seven, which promised integration never arrived, and what her nursing staff will actually say in the first week.
The single most decision-relevant conversation available in health technology procurement is systematically unavailable to the person making the decision.
What curated-reference procurement has produced
Before discussing the mechanism, look at the output, because the output is the argument.
Physicians grade their electronic health records at 45.9 on the System Usability Scale.
That study, published in Mayo Clinic Proceedings and covering 870 physicians, put EHR usability in the bottom 9 percent of more than 1,300 usability studies conducted across all industries. On the letter-grade convention associated with the scale, 45.9 is an F.
The same research found each one-point improvement in usability score associated with 3 percent lower odds of burnout.
These are not obscure niche products. These are the systems chosen through the most elaborate procurement processes in healthcare, with committees, consultants, site visits, and reference calls, at enormous cost.
Now add the evidence picture for the newer wave. A JMIR analysis of 224 venture-backed digital health companies that had collectively raised $8.2 billion found that 44 percent had a clinical robustness score of zero.
No clinical evidence at all. For nearly half of a well-funded sector.
Which means that for a large share of what a health system evaluates, the reference call is not one input among many. It is the only evidence that exists. And it is selected by the seller.
The market built a company to solve this, which tells you something
The most interesting fact in this space is that the failure is so complete that an entire industry grew up to work around it.
KLAS Research operates a business whose core activity is telephoning healthcare providers one at a time to collect the peer experience data that does not circulate on its own. By its own published figures: more than 20,000 participating provider memberships, over 6,900 technology decisions shared, and more than 1,100 products and services measured, with roughly 95 percent of data collected by phone interview.
Providers get access in exchange for candid feedback. Vendors, in the company's own framing, "work with KLAS."
Think about what the existence and success of that business proves. Demand for candid peer reference data is so intense, and its natural circulation so blocked, that a substantial company exists to manually reconstruct it, one phone call at a time.
It also has an unavoidable structural feature: it sells to both sides. A research firm whose customers include the vendors being rated is in a genuinely difficult position when it comes to publishing the harshest available truth, and it cannot make peer references free without cannibalizing the business.
There is a second finding from the same organization that is arguably more important than any product rating it has ever published. Its Arch Collaborative, covering 300-plus organizations and more than 700,000 surveys, found that organizations running the same EHR produce widely different clinician experience, with an average physician Net EHR Experience Score of 23.4 against an award threshold of 60.
Same software. Wildly different outcomes.
Which means the product is not the main variable. Implementation is. And implementation knowledge is held by named individuals, not by products, which makes it exactly the thing a product rating cannot capture.
Why the honest answer never circulates
Six structural reasons, each sufficient on its own.
The vendor owns the channel and profits from curating it. Reference selection is a sales function. No vendor will voluntarily hand a prospect a list of dissatisfied customers.
Analyst firms sell to both sides. Making peer references freely circulate would destroy the business model that funds the research.
Health systems treat implementation experience as proprietary. Your organization's hard-won knowledge about what broke is, in the institutional view, a competitive asset and a legal exposure, not a public good.
Societies will not endorse or condemn products. Understandable, and it removes the natural neutral convener.
The clinician who has the answer is paid nothing. This is the quietly decisive one. A CMIO who has run the implementation you are about to run possesses precisely the information you need. Giving it to you costs them forty-five minutes and earns them nothing. The default outcome of any system where the valuable input is unpaid is silence.
Nobody accumulates the answers. Every buyer starts from zero. The same questions are asked of the same people repeatedly, and none of it is written down anywhere.
So the market's response is exactly what economics would predict: a paid intermediary that reconstructs by brute force what would otherwise flow freely.
Why this is getting worse right now
Three forces are converging, and they point the same direction.
Simultaneous evaluations have multiplied. A JAMIA study found 43 health systems evaluating across 37 AI use cases, with 77 percent citing immature tools as a barrier. Dozens of ambient scribe, in-basket, and revenue-cycle products are being assessed in parallel across hundreds of organizations.
The evidence base per product has thinned. Newer products are evaluated faster, with less published evidence, against a shorter track record. The 44 percent with no clinical evidence at all is a sector characteristic, not an outlier.
Tolerance for failure has collapsed. Post-2024 margin pressure means buyers now show, in one industry report's phrasing, "zero appetite for failed technology rollouts." The general enterprise picture is not reassuring: an MIT analysis of 300 generative AI deployments reported that 95 percent of enterprise pilots produced no measurable profit-and-loss impact.
More decisions, less evidence, less tolerance for error, and the primary evidence source still selected by the seller.
The buyer's real question is situational, which is why ratings cannot answer it
Here is the conceptual error underneath most attempts to fix this.
The instinct is to build a better ratings database. More reviews, better methodology, verified purchasers, a five-star average per product.
But look at what the buyer actually needs to know. It is never "is this product good." It is always:
"For us. At 250 beds. On this EHR. With our staffing model. In a rural market. With a nursing workforce that has been through two go-lives in three years. What actually happens?"
That question is situational, comparative, and unanswerable in aggregate. A five-star average across 400 institutions tells you almost nothing about your institution, particularly given the Arch Collaborative finding that identical software produces wildly different results depending on implementation.
Ratings are the wrong artifact. The right artifact is a reachable person in a comparable setting.
That reframe changes everything about what would need to be built. Not a review site. A routing layer that finds the person who ran your implementation somewhere else, and gets them on the phone.
What a working peer reference system needs
Verified identity and verified setting. Anonymity destroys the entire value. "A CMIO at a community hospital" is worthless; you need to know it is a 250-bed community hospital on the same EHR with a comparable service line, because comparability is the whole point.
Conflict disclosure attached to the person. Whether the answering clinician has any relationship with the vendor, has spoken at their user conference, or holds equity. Full disclosure makes a positive review credible; its absence makes every review suspect.
Compensation for the answerer. This is where good intentions have repeatedly failed. The clinician's time is the scarce input. Unpaid peer reference calls will be rationed by personal relationship, which reproduces exactly the existing inequity where well-connected buyers get honest information and everyone else gets the vendor list.
Structured capture, so answers accumulate. The tragedy of the current arrangement is that the same forty-five minute conversation happens hundreds of times and is never recorded. Capturing it, with the participant's consent and in a form that respects employer confidentiality, converts a repeated cost into an asset.
Hard boundaries on what is shared. Personal professional experience, yes. Employer-confidential contract terms, no. Pricing coordination, absolutely not, both because it is an antitrust concern and because it is unnecessary; the lesson is never the contract. No patient information, ever.
And no vendor participation whatsoever. Not sponsorship, not access, not visibility. The moment a vendor can influence the channel it becomes marketing, and its value goes to zero in about a quarter.
What you can do now
If you are buying
Find one reference the vendor did not give you. This is the highest-yield hour in your entire evaluation. Use your own network, a former colleague, a peer at a conference. One unfiltered conversation routinely outperforms the entire structured process.
Ask the vendor for a customer who stopped using the product. The response is informative regardless of content. A vendor confident in their product can usually name a churned customer with a legitimate, non-damning reason. Evasion is data.
Ask questions that curated references cannot survive. "What did your staff start doing that you did not design?" "What was in month seven that was not in month two?" "What would you tell someone to negotiate differently?" "Who on your team would tell me a different story than you are telling me?"
Ask specifically about the implementation, not the product. Given that identical EHRs produce Net EHR Experience Scores ranging from 23.4 to above 60, the implementation approach is the variable that matters most and the one no rating captures.
Write down what you learn and share it. Your peers are about to spend the same forty hours you just spent. A two-page internal memo, shared with two peer organizations, is worth more than most analyst subscriptions and costs nothing.
If you have implemented something
Answer the calls. You hold information that would save a peer institution months and possibly a failed rollout. It is one of the highest-leverage things you can do with forty-five minutes, and it is also how you build the relationships you will need when you are the one buying.
Be specific about your setting. Your experience is only interpretable if the person hearing it knows your bed count, EHR, staffing model, and market.
Say what you would do differently. This is the single most valuable sentence in any reference call and the one most often omitted for fear of looking critical of a past decision.
If you lead a system
Fund the outbound calls. Your CMIO answering peer reference requests is not a distraction from their job; it is how they build the network that makes their next evaluation better. Most organizations implicitly discourage it.
Ask for the unfiltered reference by policy. Make "one reference we sourced ourselves" a standing requirement of any major purchase. It is free and it changes what vendors bring you.
Frequently asked questions
How do health systems evaluate technology vendors? Typically through vendor demonstrations, analyst reports, security and integration reviews, and reference calls with customers the vendor selects. Because roughly 44 percent of venture-backed digital health companies have no clinical evidence at all, the vendor-selected reference call is frequently the primary evidence available for newer products.
How usable are electronic health records, really? Poorly, by physicians' own assessment. Research in Mayo Clinic Proceedings covering 870 physicians found a mean System Usability Scale score of 45.9, placing EHRs in the bottom 9 percent of more than 1,300 usability studies across industries, with each one-point improvement associated with 3 percent lower odds of burnout.
Does the product or the implementation matter more? Implementation, substantially. KLAS Arch Collaborative data covering more than 300 organizations and 700,000 surveys found organizations running the same EHR produce widely different clinician experience, with an average physician Net EHR Experience Score of 23.4 against an award threshold of 60.
Are analyst firms like KLAS reliable? They provide genuine value and collect data no one else collects, primarily through phone interviews with providers, covering more than 1,100 products. The structural caveat is that their business model involves both providers and vendors, which constrains how freely candid peer reference data can circulate and limits the model to aggregate ratings rather than routed conversations with named implementers.
What should I ask on a vendor reference call? Ask about month seven rather than month two, about what staff started doing that was not designed, about what they would negotiate differently, and about who else at their organization would tell you a different story. Also ask the vendor directly for a customer who stopped using the product.
Why don't clinicians share implementation experience more freely? Because it costs them time and earns them nothing, their employers frequently treat implementation experience as proprietary, and there is no channel that reaches the buyer who needs it. The clinician with the answer and the buyer with the question have no way to find each other, which is why an entire research industry exists to connect them by telephone.
The bottom line
The largest technology decisions in healthcare, affecting thousands of clinicians for a decade, are made using evidence supplied by the party selling the product.
Everyone involved knows this. The buyers know the references are curated. The vendors know the buyers know. And the process continues anyway, because there is no alternative channel, and the person who actually has the answer has no reason to pick up the phone.
The results are visible in the artifacts. An F grade for usability, in the bottom 9 percent of studies across every industry. Nearly half of a funded sector with no clinical evidence at all. Identical software producing wildly different experiences depending on implementation choices that no rating captures.
Somewhere is a clinical informaticist who ran exactly the implementation you are about to run, who knows precisely what will go wrong in month seven, and who would happily tell you if anyone asked.
Nobody has ever asked, because nobody knows who they are, and if they did, no one would pay them for the forty-five minutes.
Part of a series on the missing professional infrastructure of healthcare. Previously: Who Accepts the Transfer?
Evidence note: sources include Mayo Clinic Proceedings (2020) on EHR usability among 870 physicians; JMIR (2022) on clinical robustness across 224 venture-backed digital health companies; published KLAS Research figures on its provider network, product coverage, methodology, and Arch Collaborative findings; JAMIA (2025) on parallel AI evaluations across 43 health systems; and MIT NANDA analysis of enterprise generative AI deployments as reported in 2025. Analyst firm figures are company-published. This article describes structural features of business models rather than alleging any impropriety by any named organization.