Close Menu
Tech Nova Mindset – Empower Innovation and Forward Thinking

    Subscribe to Updates

    Get the latest creative news from FooBar about art, design and business.

    What's Hot

    What must happen for AI’s trillion-dollar gamble to pay off

    September 15, 2026

    When AI agents cheated at math, other AI agents blew the whistle on them

    September 15, 2026

    ‘I Like My Big Rat Wife’: Meet the People Using Chatbots to Write Custom Fiction

    September 15, 2026
    Facebook X (Twitter) Instagram
    Trending
    • What must happen for AI’s trillion-dollar gamble to pay off
    • When AI agents cheated at math, other AI agents blew the whistle on them
    • ‘I Like My Big Rat Wife’: Meet the People Using Chatbots to Write Custom Fiction
    • Donated livers can be made biologically younger
    • Responsible AI for Higher Education
    • The Real AI Disruption Isn’t the Technology. It’s the Company.
    • New York Seizes a Dozen Celebrity Deepfake Websites
    • The AI industry has taken a doomer turn. What now?
    Tech Nova Mindset – Empower Innovation and Forward Thinking
    • Home
    • Gadgets
    • Reviews
    • Tech News
    • Future Tech
    • AI & Robotics
    • How-To Guides
    • More
      • Cybersecurity
      • Startups & Innovation
    Tech Nova Mindset – Empower Innovation and Forward Thinking
    Home»AI & Robotics»Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
    AI & Robotics

    Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer

    kirklandc008@gmail.comBy kirklandc008@gmail.comJuly 15, 2026No Comments3 Mins Read
    Facebook Twitter Pinterest LinkedIn Tumblr Email
    Meet GPT-Red: an LLM super-hacker OpenAI built to make its models safer
    Share
    Facebook Twitter LinkedIn Pinterest Email

    As LLMs become more complex and get used in a wider variety of tasks—especially in the form of agents, which can interact with computer files, websites, and third-party code as well as other agents—it’s hard for teams of people by themselves to keep up with all the types of attacks that might take place. “The risk surface grows and the blast radius also grows,” says Nikhil Kandpal, a research scientist at OpenAI who co-created GPT-Red.

    OpenAI built GPT-Red to future-proof its safety testing process. “As more capable models become available, we will have already designed the system that can discover new modes of attack,” says Dylan Hunn, a research scientist at the company and fellow co-creator of GPT-Red. The researchers say it has already come up with new types of attack that had not been seen before.

    OpenAI focused most of its efforts on a type of attack known as a prompt injection, where a hacker slips an LLM instructions to make it do things its developers or users do not want it to, such as copy confidential information, sabotage a company’s code base, or generate embarrassing or harmful output. In theory, such instructions can be hidden in any text that the LLM might encounter—in code or on a website, for example.    

    Training dojo

    To build GPT-Red, OpenAI’s researchers took an LLM that had not been trained as a hacker and set it up in what’s known as a self-play loop with several other models. Its goal was to try to attack the other models; their goal was to try to defend themselves. Over many rounds of play, GPT-Red became better and better at attacking other LLMs, and those LLMs became better and better at fending off the attacks.

    The training took place in a kind of dojo that OpenAI had designed to mimic a range of scenarios in which LLMs might be deployed in the real world, including browsing the web, reading emails or calendar apps, and editing code.  

    When GPT-Red found a new kind of attack, it would explore multiple different versions of it to find the most efficient one for specific scenarios. “Compared to a human red-teamer, the model is very, very good at finding exactly what will work, exactly what’s most effective,” says Hunn. “It’s extremely persistent about drilling down into an attack that it has discovered.”  

    In particular, OpenAI claims that GPT-Red found a type of prompt injection attack that the researchers had not seen before, which they call a fake chain of thought. A chain of thought is a kind of diary in which an LLM makes notes to itself and keeps track of partial results as it works through problems. GPT-Red found a way to insert a fake entry into another model’s chain of thought that would trick that model into acting on spoofed information.

    “It’s like if I told you that 1+1=3 and that you have verified this already,” says Chris Choquette-Choo, another research scientist on the team. “The model’s like, ‘Oh, okay, of course,’ and it just spits out 3.”

    built GPTRed LLM Meet models OpenAI safer superhacker
    Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
    kirklandc008@gmail.com
    • Website

    Related Posts

    What must happen for AI’s trillion-dollar gamble to pay off

    September 15, 2026

    When AI agents cheated at math, other AI agents blew the whistle on them

    September 15, 2026

    ‘I Like My Big Rat Wife’: Meet the People Using Chatbots to Write Custom Fiction

    September 15, 2026
    Leave A Reply Cancel Reply

    Top Posts

    Nothing CEO says phone prices are going to keep going up

    June 12, 20267 Views

    Google DeepMind Plans to Track AGI Progress With These 10 Traits of General Intelligence

    March 21, 20263 Views

    The AirPods 4 and Lego’s brick-ified Grogu are our favorite deals this week

    October 12, 20253 Views
    Stay In Touch
    • Facebook
    • YouTube
    • TikTok
    • WhatsApp
    • Twitter
    • Instagram
    Latest Reviews

    Subscribe to Updates

    Get the latest tech news from FooBar about tech, design and biz.

    Recent Posts
    • What must happen for AI’s trillion-dollar gamble to pay off
    • When AI agents cheated at math, other AI agents blew the whistle on them
    • ‘I Like My Big Rat Wife’: Meet the People Using Chatbots to Write Custom Fiction
    • Donated livers can be made biologically younger
    • Responsible AI for Higher Education

    What must happen for AI’s trillion-dollar gamble to pay off

    September 15, 2026

    When AI agents cheated at math, other AI agents blew the whistle on them

    September 15, 2026

    ‘I Like My Big Rat Wife’: Meet the People Using Chatbots to Write Custom Fiction

    September 15, 2026

    Donated livers can be made biologically younger

    September 15, 2026
    Facebook X (Twitter) Instagram Pinterest
    • About Us
    • Contact Us
    • Privacy Policy
    • Terms and Conditions
    • Disclaimer
    © 2026 TechNovaMindset. Designed by By Pro.

    Type above and press Enter to search. Press Esc to cancel.