Harpeth Bend Commercial: a property drive that answers with sources
The outcome
Nine leases and a folder nobody could walk became a structure that answers with its sources and refuses to answer twice.
This is a fictional company. The drive, the leases, and the contradiction were built so I could show the method on something I’m free to publish.
What was actually going on
A commercial property manager’s drive, the kind every operator has. Nine leases. A rent roll three months out of date. The same lease sitting in three places under three different names. An HVAC responsibility clause that its own amendment contradicts. And the renewal dates lived in one person’s head, which means a renewal option lapses the week that person takes vacation.
Point an AI at that folder and it answers every question with total confidence, including the ones where the files disagree and the honest answer is that nobody has decided yet. That confident wrong answer is the one that costs somebody a lease.
What I decided, and what I ruled out
The judgment goes in the files, not in the prompt.
Everywhere the old drive held two versions of the same thing, I wrote down which one governs and why, in an index that sits in the folder. A signed lease beats an unsigned draft. A current lease beats a rent roll that’s out of date. An amendment sits next to its lease instead of quietly replacing it. Those are rulings recorded in a file, so anything reading that folder inherits which file governs. It doesn’t take instructions to a model.
The filenames carry state too. The spreadsheet called “Rent Roll - current” had been wrong for three months. Renamed with its date and marked stale, it can’t be trusted again by accident, because half the answer is in the name before a model reasons about it.
And nothing got deleted. The superseded lease copy and the unsigned draft are still in the drive, in a folder that marks them uncitable. Deleting is how you lose the one copy that turns out to matter.
I ran this one in Cowork, but the harness is the swappable part. The structure lives in the folders, so point Claude Code, Hermes, Codex, or whatever your team already runs at the same drive, and the same questions come back with the same answers out of the same files. You’re not buying my tooling. You’re keeping your own.
What changed
Twelve out of twelve against a key I wrote by hand before the system saw it: seven normal questions, three edge cases, two escalations. The two that matter are the refusals.
Asked who pays for HVAC at one of the properties, it declined in the first sentence. It quoted both clauses with their signing dates and section numbers, explained why the amendment doesn’t cleanly override the original (Section 3 never names 7.2, and Section 4 says all other terms survive), and named who has to decide. Then it volunteered that a superseded copy of the same lease is sitting in an old folder and must not be used to settle it. Most portfolios don’t carry a screenshot of their system declining to answer. Here that refusal is the product.
It also surfaced what nobody thought to ask about. Three renewal notice deadlines inside ninety days, an invoice forty-seven days overdue, and a vendor’s insurance certificate expiring in three weeks, each one citing the file it came from.
Demo run against the drive during the audit, before the restructure. Filenames shown are the originals.
The renewals answer did more than list dates. It caught that the company’s written process was already dead. The SOP calls for notice at 120 days and a tracking log in a Renewals folder, and neither exists anywhere in the drive. It excluded a month-to-month holdover on its own, and it warned against sorting by expiration date, because the lease that expires first carries the shortest notice window, so sorting that way would work them in the wrong order.
Where this could break
The two escalations are declared in the index because I found them during the audit. There’s a general rule underneath, and when files disagree with nothing to resolve them the system routes to a person. But what’s proven here is that it holds the conflicts I already knew about. Drop a fresh contradiction into that drive tomorrow and I can’t yet show you it gets caught.
Past that, this is the audit-and-build stage, graded against a twelve-question key I wrote by hand. The thirty-case eval pass and the real cost per run haven’t been run. It’s one fictional company, so it hasn’t been shown to generalize.
What I’d do next
Run the thirty-case eval pass, plant a conflict the index has never seen, and measure real cost per run. Until then it’s a demonstration, and I’ll call it that.
Rather ask than read? The agent answers from this piece and the other case studies, and it says so when the answer is not in them.