I measure whether systems do what they report doing — my own included. Mostly that means evaluation integrity: leakage audits, walk-forward validation, multiple-comparisons correction, and pre-registered tests allowed to come out negative.

Most of what is here is a negative result, because most of what I measure turns out not to work.

Current work is on the AI power buildout — what fraction of the interconnection queue everyone quotes actually gets built. That lives at tirramind.com/queue.

Posts

subscribe via RSS