Mimir
Under the hood

How we resolve twelve spellings of one owner

The county never promised that “SMITH REALTY L.L.C.” and “SMITH RLTY” were the same company. Or different ones. Somebody has to decide, and be able to prove why. That somebody is us, and here is how we do it.

By Sean Poon, Founder7 min read

Ask a county index who owns a building and it gives you an honest answer: whatever string somebody typed into a form years ago. Ask twice across two documents and you can get two different strings for the same party. Deeds get recorded. Nobody ever promised they would agree with each other.

That is the actual problem under every “who owns what” question. Not access, the records are public. Not volume, 160M+ records fit in a database. The problem is identity: the same building lives under a block-and-lot in one system, a parcel id in a second, a document reference in a third. The same person appears as a dozen spellings across a dozen entities. Until something collapses those, every query you run is quietly wrong in both directions: it misses holdings and it double-counts them.

Why we refuse to solve it with a match score

The industry-standard answer is probabilistic record linkage: score every candidate pair, keep the pairs above a threshold. It works, in the sense that a demo works. It fails in the two ways that matter for anything a lender or a court will rely on.

  • It is unstable. Retrain the model or nudge the threshold, and identities you cited last quarter quietly dissolve. An answer that changes when nobody filed anything is not a fact. It is a mood.
  • It is unexplainable at the unit level. A 0.91 aggregate score cannot tell a reviewer which piece of evidence did the work, so it cannot be checked without redoing the whole job by hand.

Mimir’s identity spine is deterministic instead. Records are normalized, then reduced to fingerprints, content-addressed tokens built from the fields that actually carry identity. Identical fingerprints collapse. The same inputs produce the same canonical ids on every run, forever, which is what lets an answer cited in a memo be re-derived and re-checked a year later.

The ladder: strong evidence first, weak evidence never promoted

Resolution runs as a ladder of tiers. The top rungs are the boring, unarguable ones: exact parcel identifiers, exact document references. Lower rungs use progressively softer evidence: normalized addresses, then constrained name-plus-address combinations. The discipline is that evidence never gets promoted above its rung. A bare name match, alone, resolves nothing, because names are how the county spells a party, not who the party is.

The instructive failures are the ones that force the ladder to stay honest. Upstate New York tax-map “print keys” look unique, but they are only unique within a town. Treat them as county-scoped and you silently merge buildings across town lines. A sentinel like block=N/A is not a key. A shared mailing address is strong evidence, right up until the address is a bank lockbox or a registered-agent office serving thousands of entities, which is why address evidence is capped before it can fan out into a fake empire. Every one of those caps exists because the alternative was a specific, observed, wrong merge.

People are harder than buildings

Buildings at least have parcel numbers. People have spellings, initials, abbreviations and the occasional typo, distributed across Secretaries of State and county recorders with no shared key at all. So the person spine is built the same way but more carefully: separate fingerprints per source, resolution constrained by co-occurring evidence (an address, an entity relationship, a recorded instrument) and a hard rule that ambiguity fails closed. Two people who might be the same person stay two people until a filing says otherwise.

The test we hold ourselves to: every merge decision must be explainable from filings a reviewer could pull independently. Not “the model was confident.” This address, this agent, this instrument, this date.

That bar is why the resolved corpus can carry per-field citations at all. A probabilistic identity cannot honestly cite anything, because its parts have no stable existence to point at. A deterministic one can, and 21.4M owner records currently do.

See your market.

Twenty minutes on the live terminal.

See your market