Most engineering problems have a right answer that you are either smart enough or well resourced enough to reach. Patient matching is not one of them.
There is no threshold that makes both kinds of error go away. There is no data source that covers everybody. There is no identity scheme that costs the patient nothing. Every decision in a patient matching system is a choice about which failure you would rather have, and the organisations that run this well are not the ones with better algorithms. They are the ones that made those choices on purpose and wrote down what they chose.
This piece sets out five trade-offs. Each has two options. Both options cost something, in both directions, and in each case the common industry answer is to pick one by accident and discover the bill later.
Key takeaways
Patient matching has no correct configuration. Tuning it means choosing which error you prefer, and the two errors do very different kinds of damage.
Match rates between organisations have been reported as low as 50 percent, including between organisations on the same electronic health record vendor.
The headline match rates that vendors and health systems quote are usually bought with human review, which means the number describes a staffing level as much as an algorithm.
Under TEFCA, a cryptographically verified patient identity still has to clear a demographic match at the responding organisation, and the threshold for that match is set by the responder's own policy.
Nobody is coming to solve this centrally. The federal duplicate-rate targets set in 2015 were missed, and the appropriations restriction on funding a unique patient identifier has been renewed year after year.
Why the numbers are worse than they look
The most cited figures in this field come from the Pew Charitable Trusts' 2018 work on patient matching, and they are worth stating plainly because they set the floor for any honest conversation.
Research by a contractor to the Office of the National Coordinator, reported by Pew, found match rates between organisations as low as 50 percent, including between organisations using the same electronic health record vendor. Within a single organisation, CHIME estimated matching can run as low as 80 percent, rising to 90 percent or better where data quality is high and uncertain matches get manual review. Intermountain Healthcare's exchange with the University of Utah improved from 10 percent to 95 percent through sustained work, with data validation and cleansing alone taking it past 60 percent.
Two things follow from that set. First, the same EHR product on both ends buys you nothing in particular, because matching is a function of data quality and policy rather than of software vendor. Second, the gap between 80 percent and 90 percent is not an algorithm improvement. Read the condition attached to it: manual review of uncertain matches. The better number is a staffing decision wearing a technical costume.
These figures are now several years old and should be treated as the shape of the problem rather than as current measurements. The shape has not changed. ASTP/ONC, the agency formerly known as ONC, set targets in its 2015 roadmap of under 2 percent duplicate records within facilities by the end of 2017 and 0.01 percent by 2024. Both were missed.
Trade-off one: which error you would rather have
A matching algorithm has one dial, and turning it does not reduce error. It moves error from one type to the other.
Turn it loose and you get false positives. Two different people are merged into one record. The resulting chart carries somebody else's allergies, somebody else's medications, somebody else's results. This is the error that hurts people, and it is extremely difficult to unpick afterwards, because downstream systems have already consumed the merged record and clinical decisions have already been made from it.
Tighten it and you get false negatives. One person exists as two records. The clinician sees half a history. Tests get repeated, interactions get missed, and the duplicate propagates into every system downstream of the index. This error is quieter, cheaper to fix, and much more common.
The genuine decision is which of these your system should prefer, and it depends entirely on what the matched data is used for. A record being assembled for direct clinical care at the bedside has a very low tolerance for false positives. A population analytics pipeline counting distinct patients has a much higher tolerance for a wrong merge and a much lower tolerance for duplicates inflating a denominator. The same organisation often needs different thresholds for different purposes, which is an argument for matching as a service with purpose-specific confidence bands rather than one global setting applied to everything.
Most systems do not make this choice. They inherit a vendor default and discover their position on the curve from incident reports.
Trade-off two: automatic resolution or a review queue
Every matching system produces three buckets. Confident match, confident non-match, and the uncertain middle. The middle is where the decision lives.
Send the middle to automatic resolution and you get a high automation rate, no queue, and no labour cost, in exchange for accepting whatever the algorithm concluded about the cases it was least sure of. The errors you would most want a human to catch are precisely the ones you have decided not to look at.
Send the middle to human review and you get the better accuracy figure, in exchange for a permanent staffing line. That queue does not shrink as the system matures, because new registrations keep feeding it. And a review queue that gets created but never resourced is worse than no queue, because records sit unresolved and behave as duplicates while the organisation believes they are being handled.
What makes this decidable is sizing the middle before you build. If the uncertain band is a handful of cases a day, staff it. If it is hundreds, you are designing a department, and that should be a conscious budget decision made with the people who will fund it. The same reasoning applies to the identity and merge paths in custom EMR and EHR software development, where an unresourced queue is one of the more common ways a technically sound build fails in production.
Trade-off three: your own data or somebody else's
Referential matching compares your demographics against large external identity datasets, typically postal address-change records and credit bureau data. Pew found it promising, and it does lift match rates, particularly for people who have moved.
The cost is who it covers. External identity datasets have thin coverage of children, and thin coverage of people experiencing homelessness or housing instability. Those are not edge cases. They are populations where a missed record does disproportionate harm, and they are exactly the people a safety-net provider sees most. A technique that raises your average match rate while lowering it for your most vulnerable patients has made your system less equitable and your dashboard more flattering at the same time.
There is a second cost. Patients object to credit bureau data being used to identify them in a clinical setting, and that objection is reasonable rather than uninformed. If you adopt referential matching, the honest position is to be able to say what data sources are used and why.
The alternative, matching only on what you collected yourself, keeps you inside data you can explain and audit, and accepts a lower ceiling. Neither is wrong. Picking referential matching without measuring its coverage across your own patient population is wrong.
Trade-off four: a verified identity or a demographic guess
TEFCA shows both approaches operating inside one framework, which makes it the clearest place to see the trade-off.
For Individual Access Services, where a person requests their own records, the Sequoia Project's Exchange Purpose Implementation SOP for IAS, version 2.1 dated 11 April 2025, requires the provider to verify the individual to at least NIST Identity Assurance Level 2, through a credential service provider approved by an approval organisation selected by the Recognised Coordinating Entity. Authentication has to meet at least Authenticator Assurance Level 2. The credential service provider issues a signed OpenID Connect token, and the query carries evidence of that proofing as an IAL2 claims token. Verification is required before first use, again when verified demographics change, and again after credentials expire. Queries may only carry demographics the credential service provider actually verified, from a minimum set of first name, last name, date of birth, address, city, state and zip code.
That is a strong identity. It is also a real cost to the patient: an onboarding process, a government ID, and in practice a device. Anyone who cannot complete it is excluded from the service entirely rather than merely matched poorly.
Now the part that makes this a trade-off rather than a solution. Under the same SOP, a responder must answer an IAS query when it has a valid IAL2 claims token and achieves an acceptable demographics-based match according to the responder's own policy. The cryptographic proof gets the request taken seriously. It does not get the record found. The record is still found by comparing names, dates of birth and addresses, and the bar for that comparison is set by whoever is answering.
So the industry's strongest available identity assurance still terminates in a demographic match governed by local policy. Any product that treats verified identity as having solved matching has misread the framework. This is the point where patient-facing design and identity architecture stop being separable, and it is why identity decisions belong in the same conversation as patient portal software development rather than in a later integration phase.
Trade-off five: standardise on write or normalise on read
Project US@, the technical specification for patient addresses published to improve matching, exists because address formatting variation is one of the largest single sources of match failure. Pew and Regenstrief found that standardising certain elements, address and last name in particular, improves match rates.
You can apply that at capture. Validate and standardise the address at the point of registration, and your data is clean from then on. The cost is friction in the registration flow, including for patients whose addresses do not fit the standard neatly, and a longer build on every intake surface. We treat this as a core design question in intake forms software, because an address field is where most of a matching system's future accuracy is quietly decided.
Or you can apply it at match time. Normalise both sides of the comparison when you run it, change nothing about capture, and keep the registration flow fast. The cost is that every single match pays the normalisation cost, your stored history stays inconsistent, and anything that reads the address without going through your normalisation layer, including reports, letters and outbound exchange, sees the raw version.
Doing both is defensible and expensive. Doing neither is the default, and it is the reason the single highest-yield intervention in most matching programmes is not an algorithm change at all. It is address data quality.
What you actually have to decide
None of this is a recommendation, because none of these trade-offs has a correct side. What is not optional is making them explicitly.
Four things are worth writing down before a line of matching code is committed. Which error your system prefers, and for which use cases. What your uncertain band is expected to contain in volume terms, and who is staffing it. What external data sources you use, and what your coverage looks like across your own patient population rather than nationally. And what your match threshold actually is, expressed as a number, owned by a named person, and reviewed on a schedule.
That last one matters more than it sounds. A match threshold is a clinical safety policy that happens to be stored in a configuration file. Treating it as a technical setting means it gets changed by whoever is tuning performance that week, with no record of why.
Measuring any of this requires that match outcomes are themselves data you can query, which is a reporting requirement most matching implementations add late and painfully. It belongs in the design from the start, alongside the rest of the healthcare data analytics platform work, rather than being reconstructed from logs after the first serious incident.
Nobody is going to centralise this away. The appropriations restriction that prevents the Department of Health and Human Services from spending funds to promulgate a unique patient identifier has been carried forward year after year despite sustained advocacy to remove it, and the draft of version 7 of the United States Core Data for Interoperability, published for comment in February 2026, proposes no changes to the core demographic elements that matching actually depends on. The problem stays local, which means it stays yours.
At Woltrio we treat patient identity as an architecture decision with a named owner rather than a feature, which is where our healthcare software development services start on any product that will hold records for the same person twice.


