Close Menu
Tech Nova Mindset – Empower Innovation and Forward Thinking

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    The Download: US robot restrictions, and ICE’s DNA grab

    August 4, 2026

    ‘Everyone Is Doing It’: The Truth About AI in Hollywood

    August 4, 2026

    Did an AI Music App Just Snitch on the Song of the Summer?

    August 4, 2026
    Facebook X (Twitter) Instagram
    Trending
    • The Download: US robot restrictions, and ICE’s DNA grab
    • ‘Everyone Is Doing It’: The Truth About AI in Hollywood
    • Did an AI Music App Just Snitch on the Song of the Summer?
    • Trump’s AI protectionism has come for robotics
    • Turning Paper Charts Into Digital Medical Records
    • The ‘Guardrail Guy’ Went Viral for Posting About Flock Cameras. Then Someone Destroyed Them
    • The Download: reward hacking explained, and suspected Iranian cyberattacks
    • Simulation Apps Pinpoint Cause of Electronics Failures
    Tech Nova Mindset – Empower Innovation and Forward Thinking
    • Home
    • Gadgets
    • Reviews
    • Tech News
    • Future Tech
    • AI & Robotics
    • How-To Guides
    • More
      • Cybersecurity
      • Startups & Innovation
    Tech Nova Mindset – Empower Innovation and Forward Thinking
    Home»AI & Robotics»Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
    AI & Robotics

    Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

    kirklandc008@gmail.comBy kirklandc008@gmail.comJuly 15, 2026No Comments3 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
    Share
    Facebook Twitter LinkedIn Pinterest Email

    As LLMs become more complex and get used in a wider variety of tasks—especially in the form of agents, which can interact with computer files, websites, and third-party code as well as other agents—it’s hard for teams of people by themselves to keep up with all the types of attacks that might take place. “The risk surface grows and the blast radius also grows,” says Nikhil Kandpal, a research scientist at OpenAI who co-created GPT-Red.

    OpenAI built GPT-Red to future-proof its safety testing process. “As more capable models become available, we will have already designed the system that can discover new modes of attack,” says Dylan Hunn, a research scientist at the company and fellow co-creator of GPT-Red. The researchers say it has already come up with new types of attack that had not been seen before.

    OpenAI focused most of its efforts on a type of attack known as a prompt injection, where a hacker slips an LLM instructions to make it do things its developers or users do not want it to, such as copy confidential information, sabotage a company’s code base, or generate embarrassing or harmful output. In theory, such instructions can be hidden in any text that the LLM might encounter—in code or on a website, for example.    

    Training dojo

    To build GPT-Red, OpenAI’s researchers took an LLM that had not been trained as a hacker and set it up in what’s known as a self-play loop with several other models. Its goal was to try to attack the other models; their goal was to try to defend themselves. Over many rounds of play, GPT-Red became better and better at attacking other LLMs, and those LLMs became better and better at fending off the attacks.

    The training took place in a kind of dojo that OpenAI had designed to mimic a range of scenarios in which LLMs might be deployed in the real world, including browsing the web, reading emails or calendar apps, and editing code.  

    When GPT-Red found a new kind of attack, it would explore multiple different versions of it to find the most efficient one for specific scenarios. “Compared to a human red-teamer, the model is very, very good at finding exactly what will work, exactly what’s most effective,” says Hunn. “It’s extremely persistent about drilling down into an attack that it has discovered.”  

    In particular, OpenAI claims that GPT-Red found a type of prompt injection attack that the researchers had not seen before, which they call a fake chain of thought. A chain of thought is a kind of diary in which an LLM makes notes to itself and keeps track of partial results as it works through problems. GPT-Red found a way to insert a fake entry into another model’s chain of thought that would trick that model into acting on spoofed information.

    “It’s like if I told you that 1+1=3 and that you have verified this already,” says Chris Choquette-Choo, another research scientist on the team. “The model’s like, ‘Oh, okay, of course,’ and it just spits out 3.”

    built GPTRed LLM Meet models OpenAI safer superhacker
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    kirklandc008@gmail.com
    • Website

    Related Posts

    The Download: US robot restrictions, and ICE’s DNA grab

    August 4, 2026

    ‘Everyone Is Doing It’: The Truth About AI in Hollywood

    August 4, 2026

    Did an AI Music App Just Snitch on the Song of the Summer?

    August 4, 2026
    Leave A Reply Cancel Reply

    Top Posts

    Nothing CEO says phone prices are going to keep going up

    June 12, 20267 Views

    Google DeepMind Plans to Track AGI Progress With These 10 Traits of General Intelligence

    March 21, 20263 Views

    The AirPods 4 and Lego’s brick-ified Grogu are our favorite deals this week

    October 12, 20253 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Recent Posts
    • The Download: US robot restrictions, and ICE’s DNA grab
    • ‘Everyone Is Doing It’: The Truth About AI in Hollywood
    • Did an AI Music App Just Snitch on the Song of the Summer?
    • Trump’s AI protectionism has come for robotics
    • Turning Paper Charts Into Digital Medical Records

    The Download: US robot restrictions, and ICE’s DNA grab

    August 4, 2026

    ‘Everyone Is Doing It’: The Truth About AI in Hollywood

    August 4, 2026

    Did an AI Music App Just Snitch on the Song of the Summer?

    August 4, 2026

    Trump’s AI protectionism has come for robotics

    August 3, 2026
    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Privacy Policy
    • Terms and Conditions
    • Disclaimer
    © 2026 TechNovaMindset. Designed by By Pro.

    Type above and press Enter to search. Press Esc to cancel.