I measure whether systems do what they report doing — my own included. Mostly that means evaluation integrity: leakage audits, walk-forward validation, multiple-comparisons correction, and pre-registered tests allowed to come out negative.
Most of what is here is a negative result, because most of what I measure turns out not to work.
Current work is on the AI power buildout — what fraction of the interconnection queue everyone quotes actually gets built. That lives at tirramind.com/queue.
Posts
-
An agent harness that lost to a single API call
-
Thirteen Ways a Pipeline Lies
-
No detectable forward-return edge in CFTC Commitments-of-Traders positioning anomalies
-
Welcome to “How Things Work”!
subscribe via RSS