ExploitBench and ExploitGym: The Benchmarks Bitcoin Security Now Depends On
Two benchmarks measure whether an AI model can turn a suspected bug into a proven exploit. A joint UK-US government evaluation put Kimi K3 at 0 of 41 on arbitrary code execution against 20 of 41 for frontier models. Bitcoin’s defenders were pushed toward the wrong end of that gap.

