The system
A commercial real estate investment firm runs its leasing deals through a pipeline built on Airtable and Zapier. When a letter of intent (LOI) arrives by email, automation validates the sender, matches the deal to a property and suite in the firm's property management system, creates the deal record and its CRM twin, provisions a document folder in SharePoint, and runs an AI abstraction that reads the LOI and fills in about 35 financial and legal terms: base rent, term, improvement allowances, abatements, brokerage.
Those terms drive everything downstream: internal approvals, the weekly leasing pipeline report, the financial revision history. A wrong number there doesn't stay put.
My part in this pipeline: I diagnose what breaks, design the fix, and build it.
The symptom
Account managers reported that the abstraction was returning an empty base rent and wrong figures on negotiated LOIs. The natural read, and the one everyone started from, was extraction quality. Better prompts. Maybe a better model.
What the measurement showed
Before touching the prompt, I compared the file the model received against the file that actually arrived. A Zap converted incoming Word documents to PDF before extraction, and the conversion rejected every tracked change.
Fed the original file, the same pipeline, with the same prompt and the same model, recovered every missing number.
That one measurement reframed the project. The problem wasn't intelligence, it was input. Every design decision below follows from it.
What I rebuilt
What I didn't build, and why
The alternatives matter as much as the design. Each of these was considered, and several were built far enough to measure.
| Alternative | Why it lost |
|---|---|
| Run each extraction twice and compare | Catches randomness, not systematic error. The observed failures were systematic. |
| A second AI to arbitrate conflicts | The precedence rules fit in five rows with one ambiguous case. Code decides that auditably, and for free. |
| One AI agent per topic | The errors were about grounding, not breadth. Splitting multiplies the waiting inside a runtime that bills for it. |
| Strip tracked changes in code | Built and proven in Zapier, but the Airtable script sandbox can't decompress a .docx, so keeping it meant an extra round trip on every run. Shelved, not deleted. |
| A rich per-field schema (value, quote, status, unit) | Designed and discarded before build. The errors it targeted were symptoms of the sanitized input, and forcing a status on all 35 fields demands answers the model has no basis for. |
| A pattern check on the premises field | Measured across all 144 intakes, the obvious pattern matched 7%. People compare that field; machines don't. |
Known limits, written down at delivery
A unit ambiguity the merge can't see. One commission term is stored in a currency field while the value is usually a percentage. When both sources agree on "4", there's no disagreement to report. The fix belongs in the schema, not the extraction.
Multi-suite deals stop at the next step. The extraction now states every suite in the document, but the suite linking downstream is still single-suite.
Writing the limits down is part of the delivery. The next person to touch this shouldn't discover them in production.
The same instinct, one layer over: folders that survive a rename
Every deal also gets a SharePoint folder, and the firm reported two symptoms: renaming a folder broke its link, and renewals got a duplicate folder instead of reusing the tenant's. Underneath the two symptoms were three problems, and the one carrying the weight wasn't in the report: everything in the chain referred to folders by path, and a path doesn't survive a rename.
The design addresses folders by their Microsoft Graph drive and item ids, which survive renames and moves. The existing path-based Zaps (at least 14 of them) keep running untouched, with the path demoted to a cache that a scheduled sweep keeps fresh. The backfill of the existing folders checks its own work: the folder identifier already stored on each deal matches the eTag Graph returns, so a mismatch flags a folder that was moved instead of quietly writing a wrong id.
Once references survive a rename, the naming complaint shrinks from a migration to a one-line prompt change. This part is designed and verified against the live SharePoint tenant, and wasn't shipped yet at the time of writing.
A financial revision history that merges a burst of edits by the same person into one negotiation counter, so a deal's history reads as four or five counters instead of a hundred field changes. And a stage transition log that records every move from one deal stage to the next, which is what makes a weekly pipeline report possible: you can't report on movement you never recorded.
What I delivered
What this demonstrates
When an AI system gives wrong answers, the model is the most visible suspect and often not the guilty one. The first job is to look at what the model actually received.
The second job is to design the pipeline so that when it is wrong, it's wrong loudly: disagreements reported instead of guessed, decisions separated from writes, every run on the record. That's what makes AI output safe to put into a system people make money decisions from.
Measure the input before you tune the model.