Case Study

The model wasn't the problem: rebuilding AI extraction in a lease deal pipeline

A commercial real estate firm's deal pipeline used AI to read letters of intent and fill in the financial terms. The numbers kept coming back wrong. The fix wasn't a better prompt.


The system

A commercial real estate investment firm runs its leasing deals through a pipeline built on Airtable and Zapier. When a letter of intent (LOI) arrives by email, automation validates the sender, matches the deal to a property and suite in the firm's property management system, creates the deal record and its CRM twin, provisions a document folder in SharePoint, and runs an AI abstraction that reads the LOI and fills in about 35 financial and legal terms: base rent, term, improvement allowances, abatements, brokerage.

Those terms drive everything downstream: internal approvals, the weekly leasing pipeline report, the financial revision history. A wrong number there doesn't stay put.

My part in this pipeline: I diagnose what breaks, design the fix, and build it.

Data & UI
Airtable
Deals, intake archive, approvals, run logs
Extraction
OpenAI API
Reads the original LOI file and the email
Documents
Zapier + Microsoft Graph
Conversion, SharePoint folders, e-signature

The symptom

Account managers reported that the abstraction was returning an empty base rent and wrong figures on negotiated LOIs. The natural read, and the one everyone started from, was extraction quality. Better prompts. Maybe a better model.

What the measurement showed

Before touching the prompt, I compared the file the model received against the file that actually arrived. A Zap converted incoming Word documents to PDF before extraction, and the conversion rejected every tracked change.

212 insertions and 131 deletions went in. Zero came out. Everything the firm had written in redline (the negotiated rate, the multi-suite breakdown, the corrected term) was deleted before the model ever saw the file, and everything they had struck out survived. The model had been faithfully reading the counterparty's original draft.

Fed the original file, the same pipeline, with the same prompt and the same model, recovered every missing number.

That one measurement reframed the project. The problem wasn't intelligence, it was input. Every design decision below follows from it.

What I rebuilt

1
Read the file exactly as it arrived. Conversion can only subtract, and here it subtracted the negotiation itself. I verified against the API documentation that the model reads .docx, .doc and .pdf directly, and that the largest file in the base sat well under the size limit. Nothing forces a conversion anymore.
2
Let the model resolve the redlines, and pick the model for reliability. On the test document the model chose the inserted value over the struck one every time, and when struck and inserted text sat side by side it left the field out rather than invent one. A cheaper model on the same input glued a struck "3" and an inserted "5" into "35". Reliability won, at roughly $0.35 per deal for both calls.
3
Treat the email as a first-class source. Sometimes the operative terms live only in the email body, and a renewal negotiated entirely over email has no document at all. Document and email are extracted separately, with the same 35-field schema and different prompts, so a failure names its side.
4
Merge with asymmetric rules, on purpose. Numbers are compared: if the document and the email disagree, nothing is written and the conflict is reported, because a filled field looks settled and a silently picked winner is worse than a blank. Text is never compared: the document wins and the email only fills gaps. Measured, the two sources state the same term in different words constantly, and the email was the one that misread people out of signature blocks.
5
Separate deciding from writing. The decide step is pure: no table access, no network, the same inputs always produce the same plan. The write step is the only one that touches records. That split made it possible to grade extraction on real production deals with zero risk, and it surfaced pending corrections before anything was overwritten.
6
Make it re-runnable, and log every run. The abstraction used to live inside deal creation, so every test created a real deal in two systems, and deals created before a fix could never benefit from it. Now a checkbox runs it on any deal, overwriting existing values is an explicit per-run choice, and each run writes one row to a log: what it wrote, what disagreed, what's pending. The counts in that row are derived from the list of decisions, never passed in separately, because a count and a list that can disagree make a report that can lie.

What I didn't build, and why

The alternatives matter as much as the design. Each of these was considered, and several were built far enough to measure.

AlternativeWhy it lost
Run each extraction twice and compareCatches randomness, not systematic error. The observed failures were systematic.
A second AI to arbitrate conflictsThe precedence rules fit in five rows with one ambiguous case. Code decides that auditably, and for free.
One AI agent per topicThe errors were about grounding, not breadth. Splitting multiplies the waiting inside a runtime that bills for it.
Strip tracked changes in codeBuilt and proven in Zapier, but the Airtable script sandbox can't decompress a .docx, so keeping it meant an extra round trip on every run. Shelved, not deleted.
A rich per-field schema (value, quote, status, unit)Designed and discarded before build. The errors it targeted were symptoms of the sanitized input, and forcing a status on all 35 fields demands answers the model has no basis for.
A pattern check on the premises fieldMeasured across all 144 intakes, the obvious pattern matched 7%. People compare that field; machines don't.

Known limits, written down at delivery

A unit ambiguity the merge can't see. One commission term is stored in a currency field while the value is usually a percentage. When both sources agree on "4", there's no disagreement to report. The fix belongs in the schema, not the extraction.

Multi-suite deals stop at the next step. The extraction now states every suite in the document, but the suite linking downstream is still single-suite.

Writing the limits down is part of the delivery. The next person to touch this shouldn't discover them in production.

The same instinct, one layer over: folders that survive a rename

Every deal also gets a SharePoint folder, and the firm reported two symptoms: renaming a folder broke its link, and renewals got a duplicate folder instead of reusing the tenant's. Underneath the two symptoms were three problems, and the one carrying the weight wasn't in the report: everything in the chain referred to folders by path, and a path doesn't survive a rename.

The design addresses folders by their Microsoft Graph drive and item ids, which survive renames and moves. The existing path-based Zaps (at least 14 of them) keep running untouched, with the path demoted to a cache that a scheduled sweep keeps fresh. The backfill of the existing folders checks its own work: the folder identifier already stored on each deal matches the eTag Graph returns, so a mismatch flags a folder that was moved instead of quietly writing a wrong id.

Once references survive a rename, the naming complaint shrinks from a migration to a one-line prompt change. This part is designed and verified against the live SharePoint tenant, and wasn't shipped yet at the time of writing.

Also in this pipeline

A financial revision history that merges a burst of edits by the same person into one negotiation counter, so a deal's history reads as four or five counters instead of a hundred field changes. And a stage transition log that records every move from one deal stage to the next, which is what makes a weekly pipeline report possible: you can't report on movement you never recorded.

What I delivered

Every missing figure recovered without a prompt or model change
Extraction re-runnable on any deal, old or new
Numeric conflicts reported, never silently resolved
Zero-risk evaluation against production deals
One auditable log row per run
Rejected alternatives and known limits on record

What this demonstrates

When an AI system gives wrong answers, the model is the most visible suspect and often not the guilty one. The first job is to look at what the model actually received.

The second job is to design the pipeline so that when it is wrong, it's wrong loudly: disagreements reported instead of guessed, decisions separated from writes, every run on the record. That's what makes AI output safe to put into a system people make money decisions from.

Measure the input before you tune the model.