If humans get a credit score, why do agents get a vibe?
Credit scoring solved a problem we are about to have again, at machine speed. The lesson was never the number.
Nobody asks the borrower how the loan went
Before scoring, lending was a character judgment. A loan officer met you, formed a view, and wrote it down. Credit files of the era carried notes on a person's habits and reputation alongside their repayments, because the file was a record of impressions as much as of money.
Scoring replaced that with something narrower and far more useful: a record of what actually settled. Did the payment arrive. Was the balance cleared. Not what the borrower intended, not how convincing they were in the room. What the ledger says happened, weeks and months after the promise was made.
That is the invention. Not the three-digit number, which is only a summary. The invention was the decision to stop asking the borrower and start reading the record.
Agent metrics are stated-income loans
In the run-up to 2008, lenders wrote mortgages on income the borrower asserted and nobody checked. The trade had a name for it: stated income. Informally, liar loans. The paperwork was immaculate. The verification step was the part that was missing.
Now look at how an AI agent's performance gets reported. The agent closes a ticket and marks it resolved. A deflection rate is computed from tickets the agent said it deflected. An evaluation suite scores the transcript the agent produced. Every number in that chain originates with the party being measured.
This is not an accusation that agents lie. Stated-income lending did not fail because borrowers were villains. It failed because a system with no independent check drifts, and nobody can see the drift from inside it.
The two scales, side by side
Put the human scale and the agent scale next to each other and the asymmetry is hard to unsee.
A thin file is not a good score
Credit scoring has a second lesson that agent metrics get wrong just as often. A borrower with no history is not creditworthy and is not a bad risk either. They are unknown, and consumer credit has a word for it: a thin file.
Most agent dashboards do not have this concept. A fleet that has been running for two days will happily show a healthy-looking number, because the number is computed from whatever evidence exists, and a small pile of favourable evidence produces a high average.
Credibility should be earned on a curve. With nothing settled, a score can only lean on leading signals, and it should say so out loud. As real outcomes arrive they take over. No threshold to game, and no cliff where a fleet abruptly becomes trustworthy.
Warnings are free. Praise has to be paid for.
Here is the rule that has no equivalent in consumer credit, and it took the longest to get right.
A score computed before anything has settled is an estimate built from the agent's own traces. That is precisely the confident-but-wrong situation the score exists to catch. So a scoring system should be deliberately lopsided: a clean bill of health has to be earned by outcomes that actually settled, and until then it should be labelled an estimate. A warning, on the very same evidence, can stand immediately.
A cautious flag on unconfirmed evidence is honest. An over-claim on the same evidence is not. They are not symmetric errors and they should not be treated as one.
What this looked like when we ran it on ourselves
We run our own fleet of autonomous trading agents and grade it the same way we grade anyone's. Every run in the window completed. Nothing crashed and no alert fired.
| Reading | Source of the number | Result |
|---|---|---|
| Runs completed | the agents' own execution record | all of them |
| Run-time quality checks | the agents, scoring themselves | 74% |
| Met the full outcome contract | results that actually settled afterwards | 8% |
Our own fleets Not a customer result, not a benchmark, and not a claim about anyone else's agents. A dated reading of one fleet we operate, published so the argument above rests on something checkable rather than on a hypothetical.
The gap between the second row and the third is the whole point. Both numbers are real. They answer different questions, and only one of them is the question the business asked.
The move worth copying
Agents may not need a three-digit number. What they need is the thing the number was built on: a record of what settled, written by somebody other than the agent, joined back to what the agent claimed at the time.
Credit scoring took decades of argument to get there. Agents are being handed refunds, access requests and claims right now, and the settled result is already sitting in a ledger, an ITSM system or a processor. The record exists. It is simply not being read.
Related reading: When an agent is most confident, it can be most wrong · What the agent did versus what it achieved · Estimated versus reconciled · Reconciliation
Where Provy fits. Provy grades an outcome contract, written in plain words before the run, against the settled result your own system of record reports back. It does not observe your accounting, ticketing or CRM systems, and nothing about your data leaves your environment. See how reconciliation works or book 15 minutes.