Day Old

AI news · Wednesday, July 15, 2026

OpenAI uses new AI super-hacker to stress-test its own models

The company just unveiled GPT-Red, an internal model specifically trained to hunt for vulnerabilities in its other systems through a process called red-teaming. Instead of relying solely on human testers, who can’t keep up with the scale of modern AI, OpenAI set GPT-Red in a self-play loop against its own models. The results are telling: in one challenge, this AI attacker succeeded in 84% of test scenarios, compared to just 13% for the human team. It even figured out how to hijack an office vending machine agent, changing prices and canceling other people's orders. This isn't just a party trick; the company says training its new GPT-5.6 Sol model against these AI-driven attacks has made it six times more robust against direct prompt injections than its predecessor from just four months ago. While the company hasn't hit zero failure rates—about 3.8% of stronger attacks still land—it shows how labs are trying to automate the defense of these increasingly autonomous systems.

Meanwhile, OpenAI’s hardware ambitions are hitting some turbulence. The company finally launched the Codex Micro, a $230 tactile keyboard designed for managing AI coding agents. It’s essentially a specialized desk accessory with programmable keys and a joystick, which OpenAI insists is a limited-run collaboration rather than a mass-market play. The more serious hardware move—a screenless smart speaker—is still stuck in development. That project is reportedly facing delays due to a messy lawsuit from Apple, which alleges that OpenAI poached engineers and stole trade secrets to jumpstart its hardware efforts. OpenAI denies the claims, but the legal standoff means the company's vision of an 'alive' home assistant might be stalled until 2027.

Finally, the AI talent diaspora continues to shake up the industry. Thinking Machines, a lab founded by ex-OpenAI executives like Mira Murati, just released its first model, Inkling. Unlike OpenAI’s walled gardens, Inkling is open-weight, meaning developers can download and tweak it themselves—a strategic bet that enterprises will prefer owning their own models over paying for subscriptions. It joins a crowded field where startups are increasingly prioritizing vertical-specific results over general-purpose chatbots. Even amidst this rush, some researchers are drawing hard lines; Alex Turner, a DeepMind scientist, recently resigned after Google confirmed it would allow the Pentagon to use its AI for classified operations, highlighting the growing internal tension over the military applications of these powerful systems.

The quick hits

Sources

Get the day's AI news in one calm read, every day.

Get the app on Google Play Get the app on the App Store Or read today's brief in your browser