Discussion about this post

User's avatar
Michael Lopez Chiesa's avatar

Really sharp piece, and the keeper line is "a test that could have proven you wrong and didn't." Patching is the inverse: each rule is tuned to its own case, so the only number that moves is the one measured on those cases, which is Goodhart's law in slow motion, the metric stops measuring the model and starts measuring your patches. The one thing I'd add is that clean train/val/test fixes contamination but not the ticket-shaped sampling bias, so even an honest number can be honest about the wrong region of the input space, the failures like the ones you've seen rather than the ones you haven't.

No posts

Ready for more?