AI lease extraction is our flagship feature, so we owe you the honest version of how well it works — with real numbers, which nobody in this industry seems willing to publish. So we measured it, and here is everything: the method, the score, and the three answers we argued with.
How we tested (and why these leases)
Measuring accuracy requires an answer key, and nobody's real leases belong in a blog post. So we generated 50 test leases with known ground truth: four different document layouts (formal numbered sections, letter-style, dense single-block, terse cover-sheet), realistic messy variety — one to three tenants with names like "Ana-Lucia Rodriguez-Vega" and "Robert Johnson Jr.", rents with odd cents, dates written three different ways, fields that are simply absent — plus ten deliberately adversarial documents: addenda that contradict the lease body on the deposit, addenda that change the rent, deposits that don't exist, key terms buried in attachments, and one page saved sideways.
Every PDF went through our actual production pipeline — the same upload, the same extraction, the same model configuration a customer hits — and the results were scored automatically against the answer key. Nine fields per lease, 450 field checks total. One honest limitation: these are digital-text PDFs, not phone-photographed scans; the review screen exists precisely because the physical world is messier than any test set.
The score: 447 of 450, and zero real errors
Field by field, extraction matched the answer key 447 times out of 450 — 99.3%. Perfect scores on every property address (50/50), every lease start and end date (100/100), every rent amount (50/50, including the $2,400.50-style cents and the addendum overrides), every tenant name set (50/50, hyphens and suffixes included), every unit label — including correctly answering "no unit" on the 19 single-family homes — and every phone and email, including correctly reporting absent ones instead of inventing them.
The three mismatches? Documents that said the deposit was waived. Our answer key wanted "blank"; the model wrote "$0." Look at that disagreement long enough and you'll conclude the model read the document at least as well as our answer key did. Zero wrong amounts, zero wrong dates, zero wrong names, zero invented values.
The adversarial ten
The traps were the interesting part. On the three leases where an addendum contradicted the body on the deposit, the extraction followed the addendum — correct, since the later document controls — and volunteered its reasoning: "confirmed by addendum." On the two where an addendum changed the rent, it took the addendum figure both times. The sideways-scanned page read fine, at a self-aware medium confidence. Nothing in the adversarial set produced a silently wrong number.
Confidence, calibrated
The system rated 42 documents high confidence, 8 medium, 0 low — and every medium came with a stated reason (a waived deposit, a missing phone number, an absent unit designation). That's the behavior we tuned for: the extractor is deliberately conservative, because a flag costs you ten seconds of review while a silent wrong number costs you a dispute.
Where humans stay essential
Every extraction still passes a human review screen before anything is saved — that's architecture, not apology. AI reads; you approve. The same philosophy runs through the whole product: AI drafts renewals, you send them; AI suggests payment matches, you apply them.
Want to test it on your own lease? The free lease audit tool runs the same pipeline on one document, no signup. If it earns your trust, plans are on the pricing page.