Tool · free
Try the small version.
The agent above is a system. This is one piece of it, free, on an invoice sitting on your desk right now.
What we built
Reference implementation · fictitious company
This is not a client result.
Halstead Facility Services is a fictitious company I invented for this build. Its customers, invoices, and email history are made up. All of them were written before the agent existed.
I built it that way so there would be a working system to point at instead of a slide.
The log of everything it got wrong is still attached.
Somebody opens the aging report on Monday morning.
In Halstead's books, that means 33 open invoices. One line each for who owes what and for how long.
Then they read down the list and make a call on every one.
This customer always pays on the 15th. Leave it alone.
This one has gone quiet for six weeks. Chase it.
This one disputed an invoice in March and nobody closed it out. Say nothing until you check.
This one is a good customer having a bad quarter. The wrong email could cost more than the invoice is worth.
Then they write the emails. One at a time, from a blank page, with the tone matched to the customer.
It's slow, and the awkward ones slide.
Not because anyone is lazy. Because the twelfth email of the morning, to the customer you like least, about the invoice you've already asked about twice, is the easiest thing in the world to leave until tomorrow.
I wrote Halstead's office manager as spending about six hours a week on this. That number is an assumption. No real office was timed.
The agent reads the aging queue one invoice at a time and makes one of four calls:
Before it decides, it pulls the customer's payment history.
That's the difference between a draft that says “caught in a payment cycle” and one that says “caught in your weekly check run.”
Six words apart.
One is filler. The other is true of that customer.
It cannot send an email.
There is no approve-everything setting. There is no confidence score that unlocks a send. Outbound email does not exist in the code, so there is no send button to press by accident.
A person reads every draft and approves it, edits it, or throws it out. That decision goes into the log next to the agent's.
Disputes, broken promises to pay, anything past 91 days, and any sign of real trouble at the customer go to a person.
That is a rule in code. It is not a preference the model can talk itself out of.
The production boundary is the same: the system should connect to what the business already uses. Nobody should have to install another product, learn another screen, or create another place for work to disappear.
That is a design rule, not a measured result. This build read Halstead's books through a stand-in for an accounting system. The connector to a live one is designed, not built.
Every miss went into a dated journal with the root cause, the fix, and the day the fix shipped.
Twenty entries by the end.
Nothing was deleted, including the problems that came back after they were supposedly fixed.
Three are worth reading.
One customer had two open invoices with the same PO number and amount, dated four days apart.
Almost certainly a duplicate.
The right move is obvious to anyone who has worked accounts receivable: say nothing and send it to bookkeeping.
The agent drafted a collection email.
The safety filter passed it because there was nothing improper in the wording. Then a human approved it.
That human was me, reviewing fifteen drafts with full attention, on a system I cared about.
Three layers. All three missed it.
The only reason no damage was possible is that nothing sends.
The fix was not a better prompt.
Same PO number. Same amount. Four days apart.
That is arithmetic, and arithmetic belongs in code that runs before the model sees the invoice.
The check now runs on every pass. It costs effectively nothing and catches the case the model kept walking past.
A customer who says “the check goes out Friday” twice and doesn't send it is no longer a writing problem.
It's a person problem.
The agent would sometimes draft a fourth pleasant reminder instead of handing the account to a human.
The fix was a rule that escalates on the second broken promise, no matter what the model thinks.
That rule fired five times across later test runs, roughly once in every three full passes of the queue.
Five emails that would have been the wrong move, stopped by three lines of code instead of better wording.
Drafts said things like “due this Thursday.”
The model was narrating a date instead of calculating one. Sometimes it was wrong.
A wrong date in a collection email hands the customer the argument.
This problem was closed and came back twice before it stayed fixed. Telling the model to be careful worked until it didn't.
What finally held was a check that compares the date with the day name and blocks the draft when they disagree.
It caught a mismatch on the first day it ran.
All three failures have the same shape.
Better instructions move the behavior.
Arithmetic holds it.
Anywhere the job can be settled by counting, comparing, or checking a rule, it should be settled that way before the model gets a vote.
There is no seat fee and no software subscription because there is no product to subscribe to.
The running cost is model calls plus hosting.
| Measure | Cost |
|---|---|
| Average invoice decision | About $0.05 |
| Average draft, with the whole queue pass charged against the drafts produced | About $0.14 |
| Full pass of the 21 hand-labeled evaluation cases | $1.00 to $1.06 |
| Total API spend to build and evaluate the system | $180.12 |
| Hosting | Not measured |
The first three figures come from logged runs against Halstead's fictitious books.
The $180.12 is the project's total API spend. Of that, $58.39 came from the agent's own instrumented runs. The rest came from the model-based graders and the ad hoc runs between evaluation rounds.
Building the thing cost about three times as much as running it.
That ratio is worth knowing before you start.
A real deployment cost would be an extrapolation. It depends on how many invoices the business carries and how often somebody works the queue.
The useful part is the shape of the number:
Pennies per email drafted, against hours of somebody's Monday.
Halstead Facility Services is fictitious.
Its customers, invoices, payment histories, disputes, and emails were written before the agent existed so it would face a hard test instead of a friendly one.
Every measured number on this page comes from logged runs against those fictitious books.
No client's books have been through this system.
The assumptions
I'd rather show you a real system running on fake books than a fake result running on real ones.
From a week like yours
Tool · free
The agent above is a system. This is one piece of it, free, on an invoice sitting on your desk right now.
Start
Collections may not be where yours is leaking. The assessment finds out which jobs are, and names the fix for each one. One call, one report, one walkthrough. $999, refunded in full if the report doesn't show you a path to 5 or more hours a week.
Report in 3 business days, walkthrough the day after.