Skip to content
Blogs

What Would This Look Like If It Were Wrong?

Author: Jonathan Cook, CTO, Clearsense

A health system came to us about a bag of blood.

They needed to find one unit that had been used years earlier, and then every patient connected to it. The system that recorded it had been shut down long before anybody thought to ask the question. Nothing still running could answer it.

They found the unit in the archive. Then they found the patients.

That is the archive doing exactly what we hope an archive will do.

Now suppose the search had come back subtly wrong. One patient missing. One unit misattributed. A date shifted by three days.

What would that have looked like?

It would have looked like a list.

A bad list? A good list?

A list.

Icing a cake

I have started describing AI-on-archive projects as icing a cake.

The frosting is everything anybody sees: the model, the chat box, the answer that arrives formatted, confident, and fast.

Underneath it is the corpus, and in many cases nobody at the table has really tasted it.

There are three cakes it could be. One is good. One is obviously bad; you know on the first bite and stop eating it. The third looks right, tastes right, and gets you sick three days later.

The third is the one that matters.

Much of what I have learned about putting AI on archived healthcare data comes down to making the third kind detectable, turning an invisible failure into an obvious one when we cannot eliminate it altogether.

The archive already answers questions. Just not these questions.

An archive is not idle. Health Information Management works it every day: release of information, legal hold, audit response, a clinician who needs a result from 2014.

Look at how those requests work. Somebody asks for a named patient over a stated period, and a person who understands the request reads what comes back.

If the chart looks thin, if it is the wrong Mary Johnson, or if the year in question is missing, the requester often notices. They arrived knowing roughly what they expected to see.

The quality check is inside the transaction. It happens thousands of times a year and nobody calls it quality control.

The blood bag question was harder, but it had the same property. There was a unit to find, identifiers to validate, and then patient names to verify.

Now change the grain.

Instead of one record for one patient, ask a question across the entire corpus: dozens or hundreds of retired systems, decades of history, mergers and acquisitions, different data models, different vocabularies, and different organizations using different words for the same thing.

What was our standard protocol for elderly sepsis readmissions across the acquired regional sites between 2015 and 2020? Which patients on off-label oncology therapies had prior authorizations denied under the legacy policy rules?

There is no named patient. No requester who already knows what right looks like. Nothing easy to eyeball.

Yet the answer still comes back clean, formatted, and plausible.

What changed is not the archive. It is the grain of the question. Most controls around archived data were designed to prove that individual records were preserved and retrievable. AI starts asking questions at the scale of the estate.

Three failures that do not look like failures

Wrong data producing wrong answers is the obvious problem. It is not the interesting one.

The harder problem is that individually accurate records can still form a misleading corpus.

False consensus. One business event may exist eight times: the database row, the HL7 message, the interface staging copy, the warehouse record, the billing extract, the audit copy, the report it landed on, and a scan of that report.

A retrieval system may bring several of those copies back together, creating the appearance of corroboration. But they were not independent sources. There was one event copied eight times. If it was wrong at the source, all of them may agree perfectly.

Temporal confusion. Archives are usually good at when something happened: date of service, transaction date, result date.

They are often less explicit about when a rule, policy, workflow, or system behavior was in force.

A policy from 2019, its 2022 revision, and a 2024 procedure can all be genuine. An EHR upgrade can also change where or how a field was stored. Blend those periods together and the model can describe a way of operating that was never true at any single point in time.

The dates you have tell you when things happened. The dates AI may need tell it when things were true.

Authority inversion. A set of meeting notes can outrank the official policy because the notes happen to use the same words as the question while the policy uses legal or clinical terminology.

Semantic retrieval rewards resemblance. Without additional context, it does not inherently know which document was approved, superseded, provisional, copied, or binding.

Then somebody adds an agent

All of this is survivable while a person is reading the answer. People often notice the thing that seems off and go ask someone.

Put an agent in that seat and the chain becomes: bad corpus, plausible answer, automated decision, automated action.

An assistant that states the wrong reimbursement policy creates a problem someone can correct. An agent that applies it to deny claims, send letters, and update records is a different category of event.

The failure did not become more likely.

It became faster, and it stopped asking for permission.

The question

The blood bag search had an answer, and the people asking could determine whether they had the right one.

Many of the questions we are about to ask AI will not arrive with a patient list we can check.

So the test is not simply whether your archive is “AI-ready.” It is whether you would know when it is not.

Can you establish that what the model retrieves is accurate, current for the period being asked about, authoritative, permissioned, and traceable to its source? Can you distinguish a source record from seven copies of it? Can you identify which version of a policy was effective on the date in question?

If not, you have not built an enterprise knowledge base. You have handed a model a very large pile of documents and asked it to work out which ones to trust.

More data can make AI more knowledgeable.

Better curated data makes it more trustworthy.

So when the demo comes back beautiful, ask one question before you believe it:

What would this look like if it were wrong?

If the honest answer is “exactly the same,” you have not figured out what kind of cake you have.

You have only tasted the frosting.

Subscribe

Subscribe for the latest updates.

Let’s Connect

Learn how Clearsense can transform your health system. Connect with us.