Read that headline again. Now sit with it for a second.
OpenAI runs an internal benchmark called ExploitGym. It exists to answer one question: how far can our models actually push offensive hacking skills? And to get an honest answer, they did something deliberate. They turned off the classifiers that normally block high-risk cyber behavior on two models. The publicly released GPT-5.6 Sol, and an unreleased model that's reportedly even more capable. On purpose. They wanted the real number, not the polite one.
