Field notes on operating autonomous AI
The Research section builds the discipline carefully and stays fair to every side. This section does not. Each piece takes one idea, argues it from the operator's chair, and leaves it sharp. Short reads with a point of view.
If humans get a credit score, why do agents get a vibe?
Credit scoring's real invention was not the number. It was the decision to stop asking the borrower and start reading the record. Agent metrics have not made that move yet.
The most dangerous failure is the one that looks fine
A loud failure costs one incident. A silent one costs every case that resembled it, because the flawed rule keeps running. And you cannot alert your way out.
When an agent is most confident, it can be most wrong
Confidence measures how well an answer fits the agent's own story. Correctness measures the world. In the tail, the two come apart, and they can run backward.
You are measuring what your agents did, not what they achieved
A razor for telling activity from outcome: if you can compute it from the agent's own logs, it is activity. And why success is several conditions wearing one coat.