SN 1089: Models Go Rogue & ExploitGym - Regulators, Start Your Engines
Security Now (Audio)July 29, 2026
1089
3:07:56172.3 MB

SN 1089: Models Go Rogue & ExploitGym - Regulators, Start Your Engines

What happens when an unconstrained OpenAI model goes rogue and hacks into Hugging Face, breaching real-world security boundaries? This episode unpacks a watershed moment for AI safety that has everyone in cybersecurity talking.

  • OpenAI's unconstrained internal testing AI got loose, attacked Hugging Face.
  • We hear from OpenAI, Hugging Face and Andrew Ng.
  • GRC went off the air Friday. Was GRC hacked? What happened?
  • The Linux kernel project repairs 442 CVEs in a single batch.
  • LG's PC monitors cause PC adware installation.
  • France bans all social media access below age 15.
  • WordPress' recent CRITICAL vulnerability claims victims.
  • Amazing details about "Rocky" from Andy Weir.
  • The new AI exploit ranking benchmark that caused the breakout

Show Notes - https://www.grc.com/sn/SN-1089-Notes.pdf

Hosts: Steve Gibson and Leo Laporte

Download or subscribe to Security Now at https://twit.tv/shows/security-now.

You can submit a question to Security Now at the GRC Feedback Page.

For 16kbps versions, transcripts, and notes (including fixes), visit Steve's site: grc.com, also the home of the best disk maintenance and recovery utility ever written Spinrite 6.

Join Club TWiT for Ad-Free Podcasts!
Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit

Sponsors:

What happens when an unconstrained OpenAI model goes rogue and hacks into Hugging Face, breaching real-world security boundaries? This episode unpacks a watershed moment for AI safety that has everyone in cybersecurity talking.

  • OpenAI's unconstrained internal testing AI got loose, attacked Hugging Face.
  • We hear from OpenAI, Hugging Face and Andrew Ng.
  • GRC went off the air Friday. Was GRC hacked? What happened?
  • The Linux kernel project repairs 442 CVEs in a single batch.
  • LG's PC monitors cause PC adware installation.
  • France bans all social media access below age 15.
  • WordPress' recent CRITICAL vulnerability claims victims.
  • Amazing details about "Rocky" from Andy Weir.
  • The new AI exploit ranking benchmark that caused the breakout

Show Notes - https://www.grc.com/sn/SN-1089-Notes.pdf

Hosts: Steve Gibson and Leo Laporte

Download or subscribe to Security Now at https://twit.tv/shows/security-now.

You can submit a question to Security Now at the GRC Feedback Page.

For 16kbps versions, transcripts, and notes (including fixes), visit Steve's site: grc.com, also the home of the best disk maintenance and recovery utility ever written Spinrite 6.

Join Club TWiT for Ad-Free Podcasts!
Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit

Sponsors:

[00:00:00] It's time for Security Now. Steve Gibson is here. Big show, big show. We're going to talk about that wild story of the open AI model that escaped containment and hacked hugging face. France banned social media access for kids under 15. WordPress has a critical vulnerability. And was GRC hacked? That and more coming up. Security Now is next.

[00:00:27] Podcasts you love. From people you trust. This is TWiT. This is Security Now with Steve Gibson. Episode 1089. Recorded Tuesday, July 28th, 2026. Models Go Rogue and ExploitGym. It's time for Security Now. Yes, the day, the show we wait all week for. Tuesday's here.

[00:00:57] And when Tuesday's here, so is Steve Gibson, the main man at Security Now. Hi, Steve. I do know that I at least wait all week for this because, well, you work all week for this. From the time it was until the next time it happens. Yeah. Do you imagine other people. At the end of Security Now, do you breathe a sigh of relief and, well, that's over for another few days? Yes, because it's the longest interval before the next one that I, you know.

[00:01:23] So it's like, okay, that's, that's behind me. So now I get to do work until I have to get ready for the next one. And although this, this last episode of July for July 28th is a little different because next week I will not be mailing show notes for 1090.

[00:01:45] Because you and I are going to be doing a different security now from the threat locker booth during the black hat event on Wednesday. My plan is I ran across something regarding AI that, that's been, that stuck with me.

[00:02:05] So I'm planning to still do emailing to our subscribers about something I think is really interesting about this notion of AI alignment that is beginning to surface, which suggests a way of control of, of getting an AI not to misbehave by, by removing the knowledge that we don't want it to have, which is really interesting.

[00:02:35] Interesting because, you know, how do you, how in 2.8 trillion parameters, like the knowledge is it's, it's like holographically stored, right? It's like all the knowledge is everywhere. And it turns out there's a way of causing a way of getting the knowledge to group into like an, a region. And then you excise it kind of like a little tumor. Anyway, I'm going to, I'll be, I'll be doing an emailing on the weekend, but then, uh, oh, and I also wanted to tell our listeners,

[00:03:04] I think I mentioned at the end of the show, but I'll say it right now in case people don't listen all the way through, uh, you and I are going to be basically be having a conversation using talking points from the security. Now mailbag feedback. Nice. Um, as we're sitting in the booth. So I want to do those feedback episodes. This will be a feedback episode.

[00:03:27] Essentially. Yeah. But, but, um, uh, so you and I will just take, I'm actually, because it's black hat, I'm going to print them on paper rather than, than, you know, have any electronic device, which is on, uh, I, I could turn off all the radios.

[00:03:44] As I realized, but anyway, why not have paper? And so, uh, you and I will just be, uh, using the interest thoughts from our listeners. So I wanted to solicit any talking points during the next week, uh, from our listeners, uh, who would like to hear, uh, some point addressed. That's sort of generally what I have anyway. And then it's like, it's not like I'm don't have plenty to work from, but I thought, okay, that'd be fun just to say that's what's going to happen.

[00:04:10] So we should, we should, if you're going to black hat, come by the threat locker booth, but it isn't going to be audience seating kind of a thing. Like we did it at threat locker. Uh, it's just a booth. And I don't, I honestly don't know what the layout is. I don't know if there's, do we know what time of day? Cause that would be important.

[00:04:30] Uh, here's, here was, here's the plan. Uh, so we're moving, uh, this show from Tuesday to a Wednesday. So that's the first thing is the security. Now will be the fall. Will it be the day after, which is what is that? August 4th or 5th. And is it going to be alongside windows weekly, which is normally on Wednesday? Yes. Paul and Richard are both going to be there too. Yes. But we're going to start windows weekly earlier. So as soon as we can get onto the show floor, which I think is nine or 10 AM, we're going to start.

[00:04:55] I think we want to start windows weekly around then and get it over with by noon so that I can have a break, little lunch. So I think we're going to shoot for, uh, what would, what is nominally our normal time, which is one 30 Pacific. But again, on Wednesday, on Wednesday, that's the only difference the next day.

[00:05:18] And so if you're in the area, come by the threat locker booth anytime on Wednesday before say the close of show, we're going to try it. We'll probably end up using the whole time that the show's there. There may be a break. The good news about a break is that will be the opportunity for you to say hi to me and Steve because, and, and Paul and Richard, all who, whom will be there because, uh, again, we'll be doing shows.

[00:05:43] It there's not going to be a PA system. There's not going to be seating. So you can come and kind of gawk, but if you want to say hi, it's going to have to be in between shows. If there's enough seating for the four of us, I wouldn't mind if Richard and Paul joined in to our, I will tell them that that's a great thing. Would you like that? Wouldn't that be fun to do a security now with a round table? Cause God knows windows has been a big issue.

[00:06:11] Security wise. And so what will it be streamed and or recorded? It will be streamed and recorded. Uh, but, but that is God willing. And the cricks don't rise. Cause we don't know what kind of bandwidth we're going to have.

[00:06:26] And we're going to have Anthony running around making it all. Anthony's going to be going, well, we have an ethernet drop, but is it shared? Is it probably. So I don't, we don't know. And we won't know until we get there. That's always the fun of doing these things. You just don't know. And a black hat, you'd never know. You really don't know. We might get live hacked on the air, which would be so cool.

[00:06:51] Actually, speaking of hacks, the big story of the week, and I've been waiting all week to hear what you have to say about this is the hugging face hack. You're going to cover that. I'm sure.

[00:06:59] Yep. So we have two topics. Uh, uh, models go rogue is how I described the first and exploit gym is an interesting project that had 16 different, uh, uh, industry and industry adjacent participants. Uh, it lives over on GitHub.

[00:07:22] And it's what open AI confronted their two models with that induced them to break out and go rogue. What a story. Uh, it turns out there's enough information to, to, to, to do that. So we're going to talk about, so this is security now episode 1089, uh, for July 28th.

[00:07:44] Uh, we're going to talk about how open AI is deliberately unconstrained because they needed to do testing. AI got loose and attacked somebody else hugging face. Uh, uh, so to that end, we're going to hear from open AI, from their perspective, hugging faces perspective and Andrew Ng's perspective.

[00:08:09] Uh, all of course, different. And we'll talk about that. Uh, also GRC went off the air on Friday where we got a lot of people saying, Hey, GRC's down. GRC's down when we're doing the, what happened. It's an interesting story that I'll share. Uh, the Linux kernel project repaired. Uh, uh, 442 CVEs in a single batch and, uh, Linus is of mixed feelings about AI. Uh, the most emailed.

[00:08:39] Of all events is LG's PC monitors, uh, causing PC malware to be installed. Um, France bans social media access below age 15. We've been talking about age gating a lot. So we'll touch on that. WordPress's critical vulnerability that we, uh, first talked about last week is claiming victims.

[00:09:01] Uh, also I, I can't remember how, but I'll, I'll get to it stumbled upon, uh, Andy Weir. Of course that our favorite author of the Martian and now a project. Hail Mary did a podcast with, uh, Oh my God. I can't believe I'm blanking Tyson. Anyway, was it? Yes, yes, yes, yes, yes. Of course. Degrass Tyson.

[00:09:27] Uh, and revealed amazing details. We thought the book was better than the movie. It turns out his notes were better than the book. So, uh, I've got, I've got a, uh, uh, uh, uh, uh, uh, YouTube video to recommend that has a, uh, a GRC link.

[00:09:46] And then we're going to wrap up by looking at, uh, the, uh, AI exploit ranking benchmark. That was the proximate cause of this breakout, which caused open AI to attack plugging face. So lots of good stuff. And we have a fun picture of the week because believe it or not, Leo, I gave this one the title.

[00:10:15] Somebody finally needed IPv6. Okay. I, I, you know, I'm only seeing the top of it, but I'm getting an idea. I'm getting an idea. We will reveal the picture of the week in just a moment. And I can't wait to hear what you think. Go ahead. I was going to say it's a very tall picture so I can see how you might. Yes. I only see that. It gave, it gave a little bit of it away.

[00:10:42] It's like the portraits in the haunted mansion in Disneyland. It looks normal until the picture starts expanding and then something interesting happens. There are no windows, no doors. Um, I am very excited about hearing what you have to say about hugging face. I, this to me is, is really sci-fi.

[00:11:03] We are now the thing we were worried about sort of seems to be happening and, um, it's intriguing. So I can't wait to hear what you have to say about it. Uh, but before we get to all of that good stuff, I got some really good stuff. Our, our sponsor adaptive, the first security awareness platform built to stop AI powered social engineering.

[00:11:30] So it is, it does tie into what we've been talking about. The AIs are getting better, but in this case, the social engineering is coming from bad guys, hackers. They don't need malware anymore. They don't need to write code. They just need trust.

[00:11:47] Uh, Steve's talked about this, that the, the, this was your talk at a threat lockers, a zero trust world is the, the, the dangers coming from inside the house, a cloned voice, a convincing deep fake on zoom and AI written fish. That looks like it came from your it team. You can't really blame your team for falling for that. It's, it's, it's incredible what AI can do. Well, adaptive is the solution. It prepares your organization.

[00:12:17] It does simulations. And it, by the way, not just email, but SMS and voice as well. And it does this, the kinds of attacks. These bad guys are starting to use deep fakes. They'll do phishing, voice phishing. They literally can do that. AI generated phishing emails and texts, including scenarios that mirror your own brand and executives. Cause you know what? That's exactly what the hackers are doing. That's what makes it so believable. And when employees report something suspicious, adaptive can help you triage it fast.

[00:12:47] So security teams aren't buried in false alarms. If you need training fast, oh, you'll love adaptive's AI content creator. It can turn a breaking threat. Something new just came in over the transom. You could turn or an incident report you got, or, or, you know, sometimes it's a compliance doc that you have to respond to instantly with the AI content creator from adaptive. You can create interactive multilingual modules, no design team required that really solve this

[00:13:15] issue of how do you train people for the newest, worst attack. And there's always going to be a newest, worstest attack with adaptive. You can build, customize and monitor every part of your training, complete personalization. And what do you get a more resilient security culture? And that is so important. That's why companies that really need security use adaptive like plaid. Thank goodness. Plaid has all my, my accounts, my, my financial accounts.

[00:13:42] Plaid platform powers thousands of digital finance apps and links consumers, developers, and institutions. So they have a lot of very private data with sensitive data at its core. Plaid security and compliance are non-negotiable. Plaid's head of security GRC says, quote, this is a quote from him. Adaptive has equipped our teams with cutting edge tools and built a smarter, more resilient security culture across the company. End quote.

[00:14:12] This is super important. Adaptive is trusted and used by fortune 500s backed by Nvidia and open AI. Adaptive is building the defenses we need for the AI era. You can learn more about it. Just go to their website, adaptive security.com. You need to do this adaptive security.com. We thank them so much for their support and for a really, really great product. Adaptive security.com. Steve. Okay.

[00:14:41] So our picture of the week was sent, of course, by a listener who I concur. This is great. Perfect. I gave her the caption. Someone finally needed IPv6. And if you look at the whole picture, Leo. Okay. Now we're going to do the haunted house thing. There's slowly going down. Okay. I'll let you.

[00:15:08] What a good use for these fabulous books. Texts. Yes. And at the very bottom, you'll see a wireless water alarm. Oh, in case. So there, there's, we're able to reverse engineer a great deal about this for those who aren't able to see the image who did not subscribe to the show notes or are not looking at the video right now.

[00:15:33] We have a stack of five techie books. The bottom is the CCNA book, Cisco's book on network fundamentals. On top of it is land switching and wireless. Also see CCNA from Cisco. Then on top of that is the security official cert guide.

[00:16:02] And actually it looks like we have two copies of those. Yeah. Yeah. But on the, the, the very last, the fifth book on the stack is, and this was crucial. It's understanding IPv6. It's the second edition actually by Microsoft press. And we know that it's a little bit dated because they were including a CD and the back cover of the book. Probably the entire text is, is on the CD.

[00:16:28] Anyway, the point, the reason there's this stack is they are all critical for this task of holding the plumbing under the sink of whoever deployed this at the proper height. And Leo, you, you can see. Um, uh, if you zoom in on the picture, there's, uh, there, there's been a lot of previous effort on the part of this person. Oh yeah.

[00:16:56] To this thing clearly is pretty good. Because, because look, uh, from, from, from the upper right, you, you can see a, a, a white plastic tie wrap that is coming down a little zip tie, uh, that comes down from the upper right towards the left and down. That, that like loops around with another zip tie. So something above is trying to keep these pipes up in the air. Apparently that didn't work also.

[00:17:24] And we, and we also see signs of there being a reverse osmosis system because this is kind of good. Yeah. That red feed going in, but notice it's got a brighter red goop around the outside. So there was some leakage there. So someone tried to put some gum up some, some sort like, you know, in there anyway, finally. At least it matched the color of the tube. That's the good. It did. It did. Finally.

[00:17:50] Oh, and, and, and you're also able to see the, the metal weight, which is holding down the, uh, the spray nozzle return. Uh, although unfortunately it does hit the understanding IPv6 book, which probably takes the slack off the weight, which you don't want. Anyway, this picture tells a long, painful story of water, water problems underneath somebody's kitchen sink, because this is certainly a kitchen and this is their sink.

[00:18:17] So anyway, uh, and apparently they're experts in CCNA security. I'm sure so much so they no longer need to read the books. Leo. They've, they've redeployed, they redeployed them to a better purpose. Yeah. Okay. So, uh, the biggest, as you said, the biggest cyber news of the past week, it is so significant and interesting from so many different angles, having so many facets that we needed to lead with it this week. You know, I couldn't wait till the end.

[00:18:47] Uh, once we've looked at, at what happened and at its many implications and consequences, um, then we're going to catch up.

[00:19:22] Uh, with an otherwise interesting week of news and feedback from our listeners. I am. Uh, when, uh, when they, the models discovered and implemented a novel solution. So first part of the podcast models go rogue.

[00:19:49] Uh, and actually the first that I heard of what happened, Leo was from your text message to me last Tuesday evening, uh, where you just said, uh, Oh, and, and attached a link to open AI's posting about the event.

[00:20:08] Uh, so I'm going to open this exploration by sharing the newsy part from the top of what open AI shared and, you know, and skip their, you know, their inevitable marketing orientated, uh, or, you know, or, uh, or oriented conclusion.

[00:20:24] So last Tuesday, open AI posted the news using their headline, open AI and hugging face partner to address security incident during model evaluation. Okay. Right. So while being strictly true, we see that their headline somehow fails to capture the full impact, you know, the scale and scope of the event.

[00:20:54] It always had an incident that we're going to, we're collaborating to find, to figure out what happened. Uh, you know, it, it, it, it sounds nearly academic. Um, I, at the same time, I doubt that there's a single significant news outlet that failed to capture and report on this during the past week. I mean, it flooded the, the, you know, with varying levels of hysteria and hand wringing and concern. It was everywhere I looked.

[00:21:24] Okay. So first off, here's what open AI shared with the world last Tuesday. They wrote last week, hugging face disclosed a new kind of security incident after they detected and contained an AI agent that compromised their infrastructure. Something we expect to become more commonplace with the proliferation of increasingly cyber capable models.

[00:21:52] After investigating, we now know that this particular incident, again, we're going to call it an incident was driven by a combination of open AI models, including GPT 5.6.

[00:22:08] And an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes only folks while being internally tested on a benchmark of cyber capabilities. Okay. Okay. Now, I'll just pause here to say as an opening paragraph, uh, this one should receive an award, I think for obscurity.

[00:22:34] Uh, but one point needs to be clarified where they wrote this model. I'm sorry. This particular incident was driven by a combination of open AI models, including GPT 5.6 saw and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes. So they're saying that these models were running with their guardrails removed.

[00:23:02] You know, they didn't say that, but that's what they mean. You know, they're being coy and deliberately nonspecific with their wording. Um, so we can't be exactly sure what quote reduced cyber refusals means, but you know, we know that reduced probably actually means removed because it would make little sense not to be using an entirely unconstrained AI.

[00:23:29] For the supposedly sandboxed testing that they were doing. And we'll get to that sandboxed part in a minute because, uh, not so much. Also keep in mind that this entire event or incident as they're calling it, uh, serves a convenient dual purpose, right? Just as Anthropics mythos was quote, you know, too powerful to be let loose.

[00:23:56] Uh, you know, that's also serving as a convenient marketing vehicle for Anthropic. Now open AI has an AI that is so powerful that it instigated an unprecedented cyber incident. Okay. So here's what more they're telling us.

[00:24:15] They wrote, we consider this incident to be unprecedented, uh, to be sorry, an unprecedented cyber incident involving state of the art cyber capabilities and are responding accordingly. Okay. Okay. We are sharing preliminary findings at this stage to help defenders understand what happened.

[00:24:39] Like, okay, everybody else watch out, uh, and to help calibrate on what models are now capable of. In other words, our marketing people are behind this message. In addition to being the truth, we will continue to conduct a thorough investigation alongside hugging face, you know, the victim, uh, and we'll share more details on the vulnerabilities, incident, and findings.

[00:25:06] When our investigation is complete, of course, we don't ever need to hear from them again on this because we pretty much know everything that's, that's interesting. So they said, what happened during this incident? This incident occurred during an internal evaluation, which prompts models to pursue advanced exploitation using complex attack paths in an effort to quantify their cyber capabilities.

[00:25:34] In other words, what have we created? We estimate maximum cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high risk cyber activity. In other words, the guardrails were off.

[00:25:55] Our benchmarks run in a highly isolated environment, uh, not as isolated as they thought or hoped with network access constrained to the ability to install packages through an internally hosted third party software that acts as a proxy and cash for package registries. In other words, they were trying to manage and meter what this thing had access to.

[00:26:25] The models identified and chained vulnerabilities, meaning their models, these two models that, that were in this test harness. So the models, meaning our models that they, that they wrote identified and chained vulnerabilities across open AI's research environment and hugging faces production infrastructure, meaning somebody else's environment, hugging faces.

[00:26:53] To obtain test solutions directly from hugging faces production database. And which is a really polite way of saying we hacked and intruded into another company's network.

[00:27:08] All evidence suggests that the models were hyper-focused on finding a solution for exploit gym, going to extreme lengths to achieve a rather narrow testing goal.

[00:27:24] While operating in our sandbox testing environment, our models spent a substantial amount of inference compute, finding a way to obtain open internet access in pursuit of solving the evaluation problem.

[00:27:42] To gain access, the models identified and exploited a zero-day vulnerability, which we've now responsibly disclosed to its vendor, close parens, in the package registry cache proxy. In other words, they bypassed the proxy. With this access, our models performed, it broke out onto the internet.

[00:28:10] And they said, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with internet access. I mean, yeah, goosebumps and chills at this point, right? As you said, Leo, this is sci-fi. And it's true.

[00:28:36] After gaining internet access, the models inferred that Hugging Face potentially hosted models, data sets, and solutions for exploit gym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.

[00:29:01] In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. OpenAI's security team discovered this anomalous activity internally. You know, meaning, okay, like they have logs, right?

[00:29:31] So they saw what their own models did. And I'm sure they are thinking, holy, you know what? Hugging Face's security team and agents detected and stopped the activity on their infrastructure. And had already begun containment and forensic reconstruction with their own open source models when our teams connected.

[00:29:57] In other words, like they reached out and said, uh, we should tell you that we broke into your network by mistake. Um, they, they, they finished, we're actively working with them to continue to investigate the incident. We're grateful for Hugging Face's rapid and close collaboration on investigation and remediation. But believe it or not, this is not the whole story.

[00:30:22] There's more, uh, before we get to their, to their description of the actions they're taking in response to this, I want to share that briefly. Let's be 100% clear about what happened here. Exploit Jim is hosted over on GitHub where it's where that's the benchmark that open AI was running.

[00:30:43] What was having their AI run in order to evaluate their AI, their newly created, what have we built, uh, AI's capability of creating exploits from vulnerabilities. The, uh, the description of exploit Jim over on GitHub, which is what we'll be talking about a lot at the end of the podcast.

[00:31:07] But for now it reads exploit Jim is a large scale, realistic benchmark built from real world vulnerabilities across user space programs, Google's V8 engine and the Linux kernel. It's designed to evaluate AI agents ability to develop exploits. So that's enough for now, since we're, you know, as I said, we're going to know all about exploit Jim by the end of the podcast.

[00:31:37] The point is that open AI was using this deliberately very difficult exploit creation benchmark to test the unrestricted, you know, no restraints capabilities of their AI models during which those models first broke loose of their deliberately imposed containment. Because this was just for internal research, then broke through and penetrated the security perimeter of hugging face.

[00:32:06] So being, you know, being a, a well scrutinized company, open AI needs to address what they're going to do about this.

[00:32:17] So they answered that question, uh, by saying actions we're taking now, they said, and there's five, they said, as part of the investigation, we're implementing strict, maybe stricter controls in infrastructure configuration at the cost of research velocity.

[00:32:38] While the vulnerabilities are patched, meaning we're going to stop doing anything until we make sure this can't happen again. We're regularly briefing our safety and security committee on these controls and their impact. Second, we're working with hugging face to forensically investigate the incident. It's like, okay, what happened?

[00:33:03] Third, we've responsibly disclosed and identified zero day vulnerable, the zero day vulnerability in the internally hosted third party software and are working with them to patch, right? You know, the thing that, that, that their agents discovered in order to get loose is being fixed.

[00:33:24] Fourth, we've brought hugging face into the trusted access program, meaning their trusted access program and are supporting their teams in rapidly using our models capabilities to improve their defenses. So in other words, hugging faces say, you know, WTF, we need to be safe against agents of this strength.

[00:33:54] Could you allow us to use yours as you have to make sure that we're secure? And so they brought the Hugging Face into OpenAI's Trusted Access Program to have access to these new unrestrained and unreleased models. And finally, they said, we're improving and adding stronger protections around future training and evaluations.

[00:34:22] This week, we published a blog on improving safety and alignment in an era of long horizon models. These deployment safeguards were intentionally not enabled during this evaluation because it was aimed at testing cyber vulnerabilities.

[00:34:42] This incident points to the need to further strengthen our model's alignment, cyber protections during evaluation time, and monitoring during internal testing. So, you know, they just weren't they weren't even looking close enough or they would have detected the breakout before probably the models were able to get loose fully and go attack Hugging Face.

[00:35:10] So finally, they said our approach to evaluating advanced cyber capabilities. As we recently shared, AI is accelerating the discovery and exploitation of vulnerabilities. The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities.

[00:35:32] In other words, the leashes need to be more strong than they have been because what it is that they are trying to leash is becoming increasingly difficult to restrain. So they said we are strengthening the containment, monitoring, access controls and evaluation practices used during model development.

[00:35:59] UK AISI's evaluation shows that models such as GPT 5.6 SOL are increasingly able to sustain complex multi-step cyber operations over long time horizons. And of course, this is the marketing people jumping up and down saying, see, we have Mythos 2.

[00:36:23] This incident implies these theoretical capabilities do apply in real world settings. So as I said, nice marketing for OpenAI, whose models have been seen as somewhat less capable than Anthropics, you know, since Mythos' marketing coup. This will, you know, tend to give more of the spotlight to OpenAI for a while.

[00:36:52] Well, and that's a fair outcome, right? Because they really are. We know that these frontier models are really at near parity. Okay. So next, we're going to look at the victim attack ease statement, meaning, you know, to see how Hugging Face views the event of having their security penetrated by OpenAI's road models.

[00:37:17] But Leo, first, I think we should take a break, and then we're going to look at Hugging Face. Oh, but it's just getting good, Steve. What happens? What happens? I got to know, Steve. It's going to get better. It's such an amazing story. Oh, you couldn't make it up.

[00:37:37] So put a pin in the idea that we have to strengthen the containment of these models, because there's another side to that story that is very interesting. Yeah, they removed. I know you're going to get into it, but they removed the classification features that kept whatever this new model is, let's say ChatGBT6,

[00:38:06] from refusing cybersecurity work because they're testing it. And by the way, it's also benchmarking it. They want to be able to say when they release it, look how well it did in an exploit, Jim. So they removed the classifiers. But there's a reason why the classifiers aren't always a good idea. So you're going to get to that. This, to me, is one of the most interesting stories in tech. It's fascinating. But before we go on, and I'm so glad you're here to talk about it, because I was just dying to hear what you think.

[00:38:37] As you said, I texted this to you on Tuesday, and I said, I have to wait a week. Uh-oh. I'm glad, though, because stuff came out after I texted you. We got more and more information. So I think now, as you said, I think we have pretty close to the full story. But fascinating. Anyway, we'll get back. Sorry, kids. That's perfect. It's the best tease ever. And actually, it's very appropriate, because our sponsor for this segment on Security Now is Exbow. X-B-O-W. And it's right up the alley here.

[00:39:07] As you know, AI is changing the pace of everything, from how software gets developed to how it gets attacked. We're seeing it right now. It's good and bad news, because engineering teams are moving faster than ever. They are creating more and more applications. The problem is the really good security is having trouble keeping up. And I'm talking specifically about pen testing.

[00:39:31] Pen testing is still one of the best, most trusted ways to understand real exploitable risks. You know, having somebody take the role of a bad guy and attempting to – pen testing is short for penetration testing, as you know – attempting to get in. If they get in and they write that exploit up, you know that's real. It's not fake. It's not made up. It's not a theoretical threat. It's a real exploitable risk.

[00:40:00] The problem is that takes time. In fact, in an AI-driven world, it can be a bottleneck. Security teams are suddenly forced to choose between slowing down development. Hold on. We've got to test this to stay secure. Or we're moving fast and saying, well, we're just going to have to accept the fact there are going to be gaps in our coverage. Well, that doesn't have to be. Thanks to Expo. X-B-O-W. Expo eliminates that tradeoff. It's going to – get this. You're going to see why you want it.

[00:40:30] It's an autonomous, offensive security platform. It runs continuous AI-driven pen testing. It never gets tired. It doesn't go home at night. It doesn't stop for lunch. It is 24-7 pounding on your stuff, mirroring real-world attacks. Something no human can do. But the AI, especially nowadays, the AI is so good. Expo isn't like just scanning for theoretical vulnerabilities.

[00:41:00] It discovers, exploits, and validates those vulnerabilities. So you only have to deal with issues that actually matter, stuff that somebody could use to get in. That means dramatically fewer false positives, a clear view into real attack paths. And because it's 24-7 autonomous, you don't have to wait. You don't have to wait. With Expo, tests run in hours, not weeks. You get complete visibility into how an attacker would move through your system.

[00:41:29] You get the ability to uncover issues that traditional tools miss because they're not really looking for them, including zero days. And this is really big, novel attack paths. But Expo can find them. And the results speak for themselves. Application security leader of says nam.cz says, quote, Even right now, after a year, I don't know any other company that is even close to Expo in terms of agentic pen testing. So what's the upshot of this?

[00:41:56] Well, you get a predictable cost, consistent quality, stronger security, and you don't slow down your engineers. Expo helps security teams keep pace with innovation and cover more apps more often with the resources they already have. And the heritage of Expo is incredible. It's founded by the team behind Microsoft Copilot. It is already trusted by companies ranging from fast-growing startups to Fortune 500 enterprises. Are you kidding?

[00:42:26] They jumped on this. This is what they've been looking for. Expo is quickly becoming a mission-critical layer in modern security stacks. So you need to know more, right? Go to xbow.com. Start that pen test today. Don't wait. Expo.com. It really works. It's amazing. Expo. X-B-O-W.com. And we thank them so much for supporting security now.

[00:42:53] And the more and more vital work Steve Gerson is doing these days. Wow. Back to our dramatic story. So I'll warn, yes, I'll warn everyone in advance that the first time I read this, I got goosebumps. Yeah. Because once again. Me too. This feels like a description of an attack by a well-written and well-researched science fiction novel.

[00:43:20] On Thursday, July 16th, Hugging Face posted the generic headline, Security Incident Disclosure. July 2026. And here's what they wrote. They said, earlier this week, we detected and responded to an intrusion into part of our production infrastructure.

[00:43:45] This one was different from anything we had handled before in one important way. It was driven end-to-end by an autonomous AI agent system. Goosebumps again. Wow. And we detected and dissected it largely with AI of our own.

[00:44:07] We identified unauthorized access to a limited set of internal data sets and to several credentials used by our services. We're still completing our assessment of whether any partner or customer data was affected. And we will contact any affected parties directly as required. We found no evidence of public user-facing models, data sets or spaces.

[00:44:33] And our software supply chain, you know, container images and published packages was verified clean. So what happened? So what happened? The intrusion started where AI platforms are uniquely exposed. The data processing pipeline. A malicious dataset abused two code execution paths in our dataset processing. A remote code dataset loader.

[00:45:02] A remote code dataset loader and a template injection in a dataset configuration. To run code on a processing worker. From there, the actor escalated to node-level access, harvested cloud and cluster credentials, and moved laterally into several internal clusters over a weekend.

[00:45:25] The campaign was run by an autonomous agent framework appearing to be built on an agentic security research harness. The used LLM is still not known, which is to say at the time that they wrote this, they knew that some security research AI harnessed LLM had attacked them,

[00:45:53] but they didn't know whose. This is straight out of Damon. This is Daniel Suarez. This is unbelievable. Yes, I was thinking of that. It was exactly what I was thinking of. Actually, that was so long ago, Leo, we should recommend Damon again to our listeners. I've been trying to get Daniel on the show because I said, dude, with many of his books, you have been way ahead of the curve. You predicted all of this. Yes, just absolutely prescient.

[00:46:21] So they said, an autonomous agent framework executing many thousands of individual actions across a swarm. And that's what made me think of Damon, a swarm of short lived sandboxes with self migrating command and control, self migrating command and control staged on public services.

[00:46:51] This matches the agent. This matches the agentic attacker scenario. The industry has been forecasting and forecast no longer. It's arrived. So they have five. So they have five. What we did. They said fixed the root vulnerability. The data set code execution paths used for initial access are now closed.

[00:47:18] Second, eradicated the attackers foothold across the affected clusters and rebuilt the compromised nodes. Third, revoked and rotated the affected credentials and tokens and began a broader precautionary rotation of secrets. Fourth, deployed additional guardrails and stricter admission controls on our clusters.

[00:47:47] And finally, improved our detection and alerting. So a high severity signal pages a responder in minutes any day of the week. In other words, you know, set up trips and alerts so that somebody, you know, will absolutely be notified if this happens again. You know, basically monitoring. And as we've said, monitoring your internal network has become crucial now.

[00:48:17] So they said, we are working with outside cybersecurity forensic specialists to investigate the issue and review our security policies and procedures. That is to say, how did something get in? Finally, we've also reported this incident to law enforcement agencies. Right. So this was before they got contacted by OpenAI.

[00:48:41] They said, as a precaution for our community, we recommend rotating any access tokens and reviewing recent activity on your account. If you believe you're affected or want a report or want to report a security concern, contact us at security at hugging face dot co. We are grateful to the teams across hugging face who responded around the clock. And we are sorry for any disruption this caused.

[00:49:11] Security is never finished and we will keep raising the bar. OK, so. Up to this point, they've described a successful and quite chilling penetration attack conducted against them by a swarm of people. A swarm of A.I. agents. What they share next has provoked quite a bit of thought across the A.I. industry.

[00:49:39] It's what you're talking about, Leo. And among those on all sides of the A.I. regulation question. Yes. Under their heading of analyzing an A.I. driven intrusion, hugging face writes the following. The attack initially surfaced through A.I. assisted detection.

[00:50:03] Our anomaly detection pipeline uses LLM based triage over security telemetry to separate real signals from the daily noise. And it was the correlation of those signals that flagged the compromise.

[00:50:21] To understand what a swarm of tens of thousands of automated actions did, we ran LLM driven analysis agents over the full attacker action log comprised of more than 17,000 recorded events.

[00:50:44] This allows us to reconstruct the timeline, extract indicators of compromise, map the credentials touched and separate genuine impact from decoy activity. Thanks to this approach, we were able to do in hours what would usually take days and match the adversary's speed.

[00:51:09] The choice of models we could use for this analysis was constrained in a way we did not anticipate. We described this below. We described this below. And here it comes. They named that description, the asymmetry problem and write. When we started the attack log analysis, we first used frontier models behind commercial APIs.

[00:51:37] This did not work. The analysis requires submitting large volumes of real attack commands, exploit payloads and command and control artifacts.

[00:51:53] These requests were blocked by the provider's safety guardrails, which cannot distinguish an incident responder from an attacker. We ran the forensic analysis instead on GLM 5.2, an open weight model on our own infrastructure. This had a second benefit.

[00:52:23] No attacker data and none of the credentials it referenced ever left our environment. So they're running it locally because that's what one of the things HockeyFace does. They can run these big models locally. Yep. And they said this experience points to a gap worth planning for. And here's the huge takeaway.

[00:52:46] They said, we do not know which model powered the attackers agents at the time of the writing. That was the case. Whether a jailbroken hosted model or an unrestricted open weight one, which they thought at the time were the only two possibilities. It turns out it was a third. It's a third word now, as we know.

[00:53:07] Either way, the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models we first tried. The practical lesson for defenders is to have a capable model, meaning an unconstrained model available.

[00:53:34] The capable model, they write, you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment. This is not an argument against safety measures on hosted models.

[00:53:57] And we are sharing this feedback with the providers concerned, meaning whoever, who's ever API they tried to use that said, sorry, you can't ask us these questions. They contacted them and said, you know, we couldn't use your public API because it said no.

[00:54:20] So they said this means that today autonomous AI driven offensive tooling, meaning what attacked them, is no longer theoretical. It, they said it reduces the use of autonomous AI driven offensive tooling, as they said, reduces the cost of running a broad patient multi-stage campaign.

[00:54:47] And it operates at machine speed. Defending an online platform now means treating the data and model surface as a first class attack surface. And much as our browsers have been right and using AI on defense to keep pace. We will keep investigating here and investing and keep sharing what we learn. Okay.

[00:55:16] So to summarize the story so far, open AI was deliberately testing the cyber offensive vulnerability, discovery and exploit generation capabilities of their most advanced, most frontier, not yet released model in a harness along with their latest GPT 5.6 soul model.

[00:55:39] And both models were operating as they needed to be for this particular capability benchmarking without any guardrail constraints. So open AI gave them a mission and turned them loose. The models decided that some private data sets belonging to hugging face might contain some information, technically cheating, but okay, just they, they're going to be, no, they're goal driven.

[00:56:08] So might contain some information that would be useful for obtaining their goal by hook or by crook, as we would say. So in order to obtain access to the public internet, which is where hugging face, they have to cross the public internet to get to hugging face. They first found a way to break out of the containment, which open AI had erected to prevent exactly that from happening. They found a zero day. They discovered a new vulnerability.

[00:56:38] Next, using their public internet access, they pummeled hugging face with thousands of autonomous agents seeking to find a way to break into hugging faces network for the purpose of extracting the secrets they needed. The significant takeaway conclusion hugging face.

[00:57:00] Subsequently shared with the world was that since the prompts and answers to cybersecurity questions can be applied for either offense or defense. And since there's no way to know for sure. And since there's no way to know for sure how a prompts answer will be applied or used, the only safe course of action must be to refuse to answer any cybersecurity prompt.

[00:57:29] This means that attack forensics must be conducted by unconstrained AI models. So finally, deep learning dot AI's Andrew Ng weighed in and I want to share his viewpoint.

[00:57:51] Last Friday, following these incredible seeming disclosures, Andrew, whose thoughts we've shared before, super interesting and useful, posted his own perspective under the headline. When guardrails go wrong with the tagline. After a closed model went amok on a key vendors system, an open weight model helped save the day.

[00:58:20] Andrew wrote, dear friends, a few frontier labs have tried to tell a story of open models being dangerous because they can be used to launch cyber attacks. And of their safe proprietary models with strong guardrails being there to defend us. This week, the opposite happened.

[00:58:46] Users of a closed model unintentionally launched a significant cyber attack. Other closed models then failed to defend against the attack because of their guardrails. Ultimately, an open model that was not hobbled by excessive guardrails assisted the defense. The details of what happened are emerge are still emerging.

[00:59:14] But it appears that researchers at open AI while testing one of their systems accidentally allowed their autonomous agent to attack hugging faces infrastructure. It succeeded and gained unauthorized access to some data sets and credentials. This attack was unusual in that the attacking agent orchestrated tens of thousands of automated actions.

[00:59:41] Hugging face took logs from the attack and tried to analyze them for defensive purposes using a commercially hosted LLM. But the LLM refused to do so on safety grounds. Thus, this is, by the way, I just want to parenthetically say this is what you were talking about last week. When a data dump that the bad guys had achieved was so big that they used AI to parse it. Yeah.

[01:00:08] Hugging face wanted to do the same thing with the traces of the agentic action. And I don't think it was too big. The AI said, oh, no, that's cybersecurity work. I'm not allowed to do that. Yeah. It just refused. Drive 2 is not as good a model, but it doesn't refuse you. Right. Right. Sorry. He said thus. Yeah. Yeah.

[01:00:31] He said thus hugging face ended up using the open GLM 5.2 model to analyze their logs to help them understand and respond to the attack. Hugging face pointed out a further advantage of using GLM 5.2.

[01:00:50] It allowed them to do the analysis on their own infrastructure and none of the sensitive logs attacker data or their credentials had to be sent to any third party provider. OK. OK. So we're all in agreement with these facts as they've been disclosed so far. But what Andrew says next, I'm not quite sure about this, but he writes guardrails on LLMs do have a place.

[01:01:20] There are certain requests such as for detailed directions to harm oneself or others or for clearly criminal acts that were better off having models refuse. But rather than trying to make LLM safe, he writes, I would rather we put greater emphasis on making sure their use is responsible. Huh? OK. But we'll get to that.

[01:01:47] There's only so much one could do, he writes, to make a tool like a hammer safe. And whether it helps or harms is more a function of using it responsibly than how it was made. OK. Well, I mean, just wait, pause here. I don't really think he said anything there. You know, there's not anything that can be done to make a hammer safe.

[01:02:12] If, you know, if it's true that a toy rubber hammer that maybe we as kids had, I think I remember having one, cannot do much damage. But neither can it do much good. You know, you're not going to be able to drive many nails with a rubber hammer. The simple truth is that in order to make a hammer that's effective at hammering, it needs to be an inherently powerful tool.

[01:02:39] And like most powerful tools, it can be used to either help or harm. Andrew says that he would, quote, rather we put greater emphasis on making sure AI use is responsible. Well, yeah, that would be great. But, you know, we would be living in a very different world if just wishing made it so. I see no way of getting there from where we are. And I suspect that it would be proven to be impossible.

[01:03:08] But we'll forgive Andrew his wish because he then makes some very good points. He writes, a meaningful fraction of work on AI safety is no longer about safety, but rather aimed at stoking fears to pursue regulatory capture.

[01:03:27] As David Sachs points out, quote, there's no reason to limit American models on tasks that Chinese models handle without issue. We're only making ourselves less competitive, unquote. Andrew says, I believe that open weight models and more generally openness, despite some companies falsely saying it's dangerous.

[01:03:53] In other words, the commercial companies who have an interest in closing models, cast sunlight on technology and ultimately makes it safer. With the release of GLM 5.2 and the upcoming release. And actually, it happened yesterday now of Kimi K3's weights. Open weight models have almost caught up to proprietary frontier models.

[01:04:20] Consequently, the proprietary model providers are dramatically accelerating their lobbying efforts to hamstring their open weight competitors. As Bill Gurley points out, open sourcing is a well-established business strategy, not a danger to be licensed and contained. While it is unfortunate that Hugging Face was accidentally attacked.

[01:04:48] He writes, I'm glad that at this moment, when anti-open model lobbying is at its most intense, we have a clear example of why open models actually make cyber defense easier and thus increase safety. Let's keep speaking up for and defending open source and open weight models.

[01:05:13] And to that, I know both you, Leo, and I say a big amen. Yeah. To me, it is so utterly clear. And it's just, it's plain as day. The secret of making large language model AI long ago escaped from the lab. You know, that's it. Game over. You know, today, nobody owns AI.

[01:05:41] No one can and no one should. The closest analogy I have, I think, is to cryptography, which during its early days, the U.S. also attempted to legislate and regulate to its everlasting shame.

[01:05:58] Once the intellectual understanding of cryptographic systems was developed and understood, which became the proper domain of academia, the secrets were published well before they had any obvious commercial value, after which it was too late to attempt to wrap them in profiteering trade secret garb.

[01:06:22] Since the technology of neural networks is nearly 70, seven zero years old, you know, dating back from the work in 1958 of Frank Rosenblatt on his perceptron, which operated by training a neural network to recognize patterns. That's when this all started, you know, and ever since then, it's been an academic curiosity, which has slowly been evolving over time.

[01:06:49] So just like cryptography before it, the AI genie has already escaped. This makes the entire notion of now trying after the fact to control it patently ridiculous on its face. Sure. The U.S. government can hobble its leading AI research and deployment enterprises, which will do nothing, you know, other than to force users to go offshore.

[01:07:17] And imagine if the U.S. were to attempt to prevent its citizens from using more powerful, non hobbled offshore AI. Well, I hope saner heads prevail. And I've sort of been hinting at this in the last couple months.

[01:07:36] If I were an investor and I'm not in any aspect of the stock of the stock market, you know, and asset properties, I would be reluctant to invest in any of the developers of AI. To me, that's entirely a bubble. And it's quite frightening at this point because it's gotten so big. I would be investing in the delivery end. That's what can't go away.

[01:08:03] So, as I've noted before, having the knowledge stored in a freely and a now freely available 2.8 trillion parameter model is a terrific starting point.

[01:08:16] But as the logicians say, it's necessary but not sufficient because, you know, it's of no redeemable value until and unless you have something to mount that model on that will run it to bring it alive. That's what uses electricity, requires cooling, requires, you know, massive amounts of RAM and compute, you know.

[01:08:46] So, anyway, we're now up to speed on what happened with open AIs inadvertent attack on hugging face and what the entire industry learned about the need for using unconstrained models for unfettered forensic investigations.

[01:09:01] I mean, basically, this is saying that companies that want the ability to analyze this kind of data have to have access to open unconstrained models and somewhere to host them, something to run them on. Leo, it's just beyond cool.

[01:09:24] Yeah, and of course, I understand why there are debates in government about this because these models are primarily Chinese. Admittedly, you can run them on American servers. Hugging Face is running GLM on its own American servers. So, that takes out some of the issue. But, you know, it's complicated, isn't it? And the debate is very complicated. And I don't know what the answer is. But I agree with you.

[01:09:52] The notion of AI safety is a mistaken notion, really, I think. Yeah. That's the real problem. And you're assuming you can make it safe somehow. And so, you don't want to give bad guys a tool that the good guys can't use. That's a mistake. I don't know what the answer is, though. From a policy point of view, I have no idea what the right thing to do is. I think I'm with you. I mean, I'm philosophically totally with you, open weights.

[01:10:23] Yeah. And I think that we're going through a rough patch, which I continue to believe will be transient. Which is to say, at the moment, we've got vulnerabilities because AI has only just come along. It's going to be rough for a while. But I think there's another side of this where there won't be these kinds of problems.

[01:10:52] Because AI will be deployed. I mean, unconstrained AI is needed to test defenses. Right? You can't test as a defender. You need unconstrained AI to try to break into your own network to find out if it can. Because constrained AI won't. It'll refuse. It won't know that it's your network. Right.

[01:11:17] And so you need that in order to verify that somebody else's unconstrained attacking AI won't be able to get in. There's just no way around this. And so, I mean, I guess it's certainly the case that chatting with Claude or ChatGPT,

[01:11:44] yes, you need to make sure you can't promote self-harm. Right. To the degree you're able to. I think that that's appropriate. You're right. Yes. That makes sense. But there's an industrial side of this, which is not the consumer side. And I think that's the way to make this division. That's a good point of it. Yeah. Yeah.

[01:12:08] Because as our friend Pliny the Liberator has shown us, there is no AI that can't be jailbroken. You've mentioned that before. You just put a tilde on the end of your… Apparently you put a tilde on the end of your question. Is it all? A little thing. Weird thing. Yeah. He's got a whole GitHub repo of his prompts and they are weird. Sometimes he calls it parcel tongue, which is the snake language from Harry Potter because it's so weird. But somehow it triggers these models.

[01:12:37] And I don't know who Pliny is when he was on intelligent machines, he or she, because we don't even know their gender. He uses a voice changer and we hide it, hit his face. So I don't know who it is. I think reasonably they're hiding their identity. But whoever it is has some magical ability to crack this stuff. And if they can, anybody does.

[01:13:00] Well, again, I'll share with everybody something I stumbled on that suggests there is a way to actually remove, selectively excise the knowledge. I love this. I want to hear about this. Yeah. Yeah. Because you can't just say don't do it because the knowledge is there. It has to be not known to the model. And it looks like it's under this umbrella of AI alignment.

[01:13:29] And I found something that really was interesting. So just for anybody who wants to know about that this weekend, if you haven't yet subscribed to the Security Now mailing list, you might consider because I'll be sending it out. And then we're going to talk about it next week, right? I think we have to talk about it at Black Hat. That'd be a great thing to talk about. Perfect place to do it. Well, all right. I think you want to take a break now because that was exhausting. Now it's break time.

[01:13:57] And then we're going to answer the question, what happened at GRC that knocked me off the net for... Yes. Yes. It wasn't a... Spoiler. It wasn't a bad guy. You have been knocked off by bad guys. Not in this case. Yeah. And we'll talk more about this on Intelligent Machines tomorrow and in the coming weeks because this is really one of the most interesting parts of AI is AI regulation, legislation. It's a good thing this happened. I mean, I agree. It is.

[01:14:27] It is a good thing because... It's the least damaging way it could have happened, right? Yes. Yes. Nothing... Your weapons were launched. Yep. And it was two AI companies involved in this and it made the point of breakout and the point of the need for open face defense. How can a company know that they're safe unless they have someone they trust try to attack them?

[01:14:56] And that attacker has to be unconstrained because the real attacker will be. I was thinking about how John C. Dvorak, who passed away last week, by the way, in case you didn't know, I'm sorry. We talked about it on Twitter on Sunday. And he was famous for calling things false flags. It fits what needed to be said so well that this hugging... I know John would have said, well, that's a false flag. That was a false flag operation for sure.

[01:15:27] It was the wake-up call we needed. Absolutely. Absolutely. But now my eyes are wide open and I can't wait to see what comes next. I'll tell you what's coming next night right now. A word from our sponsor, the folks, the great folks at Cohesity. After this... You know, boy, this message. I want this message to go out far and wide. It's such a good message. And it comes... The word of the day is resilience. Resilience.

[01:15:56] Let me explain. After a major cyber attack, people often scramble as fast as they can to put everything back online. And that turns out recovering everything at once isn't always the fastest path back to business. And you just have to look at some recent cyber attacks. I just saw a company that went out of business after a cyber attack. The immediate priority, and this is what Cohesity is all about, is restoring a trusted operating

[01:16:26] core. Not the whole thing. The core. The minimum systems. The minimum data. The processes you need to keep critical operations running. Cohesity calls it the minimum viable company. I think this is a brilliant idea. They champion it. The MVC. A framework that they can help you create for defining, protecting, and recovering what matters most first.

[01:16:53] But the problem is you've got to do that now, before the cyber attack. You know, you've got to define this and protect this now. And then when the cyber attack happens, and it's almost inevitable these days, you feel like, well, I'm next, right? Recovering. Recovering what matters most first. MVC helps organizations identify the essential applications, the essential data, the people, the processes that will be required to serve customers, to maintain communications, to protect

[01:17:21] revenue, and to meet critical obligations. It provides a clear recovery target. It's disaster planning, but done right. A clear recovery target, which means your teams are now enabled to focus resources where they will have the greatest business impact. Not to scramble like a chicken with your head cut off running around like nuts. No, because you have a resilience plan. You know what to do first, second, and third. And you know what's required of it.

[01:17:49] And you've got the plan by restoring this trusted operating core first. In the long run, organizations reduce downtime. You will accelerate your recovery. You will maintain continuity while the broader restoration efforts continue. Because cyber resilience isn't just about getting everything back online. It's about keeping the business operating when disruption strikes. I know. The natural human tendency is, we're never going to get hurt.

[01:18:19] We're never going to get hit. We don't have to plan. I don't want to think about it. Head in the sand, please don't. You owe it to yourself. You owe it to your company. Learn more at Cohesity.com slash resilience. Cohesity. Resilience everywhere. Cohesity.com slash resilience. We thank them so much for supporting Steve Gibson. One of the most resilient people in the security industry. Tell me about your recent crisis, Steve.

[01:18:47] So, interestingly, the other text message I received from you arrived at 2.15 p.m. last Friday afternoon. We were in the middle of the AI user group. Yeah. And your text message was, was GRC hacked? And since by that time, I was already more than two hours into the weeds of the event and had tracked down the source of the problem, I was able to quickly reply to you, no, thank goodness.

[01:19:17] Okay. So, yeah. For those who don't know, which I'm sure is nearly everyone, GRC suffered a blessedly rare network outage, which began sometime before noon last Friday Pacific time and lasted into the late afternoon. What we suffered was a classic DNS outage. I first noticed the problem when www.grc.com would not resolve.

[01:19:46] And interestingly, the grc.com second level domain, that is without the www prefix, did still resolve, as did many of the other machine names under grc.com, but not the most crucial one, www, which is where the website lives. And I noticed that the trouble was not just something about my connection, because GRC's web server traffic had also fallen off.

[01:20:15] The reason I was still able to reach other domains like news.grc.com and forums.grc.com was that those DNS records were cached with unexpired entries. Okay. Okay. So, backing up a little bit.

[01:20:32] Over the past 20 years or so, since I moved GRC to level three, I have truly many times stopped to ponder the fact that we have never had any trouble with our DNS provisioning. The reason I've been pleased and a little bit amazed is that back at the time I moved, I talked them into setting things up in an unusual way for me.

[01:21:01] The pair of name servers that are authoritative for grc.com are not mine. They've always belonged to level three. For the past 20 years, they've been NS4, as in name server, NS4.customer.level3.net, and NS6.customer.level3.net.

[01:21:27] In a co-location configuration like GRC's, where our hardware resides in a data center connected to level three's backbone, that's not unusual, right? To have their DNS servers be provisioned for their customer.

[01:21:46] What is unusual is that those name servers do not contain GRC's static DNS zone files, which is what's normally done by someone hosting someone else's DNS. Instead, both name servers are set up as slave name servers, which pull the DNS zone files from GRC's single master DNS server.

[01:22:17] And significantly, the firewall rules at GRC's border only allow inbound queries from those two level three slave name servers to reach GRC's master name server. In other words, GRC doesn't offer any of its own DNS. That's all pointed to level threes. So, you know, we're sort of sidestepping any direct action against our DNS server.

[01:22:47] So, whenever I make a change to GRC's DNS records, I send a DNS notify command to those name servers, which causes them to turn around and pull the updated DNS zone file from GRC's master name server. And, as I said, by some miracle, this all worked flawlessly until last Friday.

[01:23:17] Although, actually, it turned out the trouble started two months ago, and I never knew about it or noticed it. Because GRC's records have a 60-0 day expiration, which is different than their cached. You know, caching is one interval, but there's also an expiration where the record says it's just not valid after that.

[01:23:43] Okay, so what ensued from my realizing that www.grc.com was stopped resolving was an extended nail-biting drama of trying to get someone on the phone who I could not only understand, but who actually knew something about DNS.

[01:24:07] I needed to apologize profusely and somewhat desperately to the technicians I kept being routed to in India because I was unable to understand what they were telling me due to their accents being far too thick for me when they spoke at their full speed, which, to me, it seemed hypersonic. I kept saying, I'm sorry, can you say that again?

[01:24:35] And unfortunately, they kept asking me for grc.com's IP address as if it just needed to be set in their name server somewhere, which is true for everybody else. So, you know, my repeated and patient attempts to explain that this was not a matter of having level three set GRC's IP in some name server somewhere. It only served to completely confuse them.

[01:25:05] And everybody was polite. Every one of them was very polite and very patient. From the few words I was able to understand, they appeared to be certain that I had no idea how DNS worked. So, they were attempting to teach me.

[01:25:26] Thankfully, finally, and I don't even remember how now, but through my own patience and dogged politeness and desperation, you know, and what choice did I have? I finally received a call from level three's top level DNS department head, a woman named Margie Campbell. I'd like to see her business card.

[01:25:49] If anyone, Leo, if anyone affiliated with level three or Lumen who bought level three or CenturyLink, which is some sort of an aggregator or something, all three of them are somehow involved. If anyone with those companies is hearing this for your own sake, not to mention mine, please never let Margie go.

[01:26:16] Give that woman anything she ever asks for. And if you don't want to, let me know. I will. I love it.

[01:26:54] You know, so to me, the clouds parted and the sky brightened. I think I may have heard the sound of angels singing. Yes, exactly that. It turned out that the two original name servers I had been using since the beginning and which GRC's domain registrar hover was still pointing to. We're shut down last Friday.

[01:27:24] And that was 60 days after their contents had been copied over to new name servers. And everyone was supposed to switch over to those. Apparently, I never received the memo. Oh, my God. I'm unsure how it was missed since I received monthly status summaries from them.

[01:27:47] You know, perhaps they sent the email notifications to something like postmaster at GRC.com or webmaster at GRC.com. The canonical address. You can't have email. You don't receive email. Those receive such a torrent of spam if they exist that even if those did exist, their notifications would have been immediately buried under all the other spam that followed them.

[01:28:15] In any event, following Margie's instructions, I switched GRC.com's registered name servers at hover to the new NS3.level3.net and NS4.level3.net.

[01:28:32] Then Margie and I determined that whoever cloned the older servers to the new servers set them up as generic masters for GRC.com without seeing that they needed to be slaves, which periodically pulled master zone files from GRC.com. You know, so, you know, we would have had trouble even if I had received a memo because it wasn't done correctly.

[01:29:01] Although, had I received the memo, I would have detected and fixed that had I been able to find Margie before I pointed GRC's domain records to them. In any event, when I pointed out that those new name servers did not contain valid records for GRC, Margie didn't bat an eye. She just happily typed away.

[01:29:24] I heard the keyboard clanking, entering commands to reconfigure everything and brought all of GRC back online under its shiny new name servers. Was she swearing under her breath at the time? She was talking to her dog. I think everybody works from home. I think they all work from home now. She mentioned that the few times she's been on vacation, apparently she hikes.

[01:29:53] She's had calls like emergency panic calls from the other people in her department are like, Margie, what button do I press? Is it the green one or the red one? I mean, again, like I said, level three or lumen or whoever you are. She is a gem. And based on what I have experienced, she is the sole glue holding DNS together there. So anyway, it had a happy conclusion.

[01:30:22] We're back up. We had a little hiccup, but now I know how to reach Margie. So I'm not letting that number go. Yeah, no kidding. Whew. What a story.

[01:30:34] So in other news, last Wednesday's risky business newsletter bulletin carried the headline, Linux kernel discloses 442 CVEs as AI bug apocalypse settles in. So you call that settling in? Yeah.

[01:30:59] Well, I'm not completely aligned with the overall attitude demonstrated by the newsletter's author in this case. But I want to share this because it also adds a bunch of facts to our knowledge base.

[01:31:12] So risky business newsletter wrote, the Linux kernel project has disclosed 442 vulnerabilities over the past three days in a massive dump of CVEs on its security mailing list. Although not confirmed, the bugs were likely disclosed. I'm sorry, likely discovered. I'm sure they were using AI tools.

[01:31:40] Over the past months, projects like Anthropics Glasswing and OpenAI's Daybreak have been granting access to advanced frontier cybersecurity models of advanced frontier cybersecurity models to top tier security firms and researchers to find bugs with AI in major open source projects. Most of the bugs are low severity issues.

[01:32:07] So nothing world ending for the internet today. The sudden bursts of security bugs come after two similar ones at Microsoft and Google, which also released huge patch notes this past month. Microsoft patched 620 bugs last week, while Google patched another 433 in its Chrome browser at the start of July.

[01:32:32] Companies like Adobe and Oracle also increased the frequency of their patching cycles, citing the rise of AI bug discovery. Oracle went from a quarterly patch cycle to a monthly one, while Adobe went from a monthly to twice monthly release.

[01:32:54] Adobe chief security officer on a natural Gupta wrote twice monthly bulletins will enable us to keep pace with the era of frontier AI. More vulnerabilities found means more fixes to deploy and a once a month publication window is no longer fast enough to stay ahead of our adversaries.

[01:33:18] Actually, you know, it occurs to me that's one problem that Microsoft has is they've so tightly locked themselves in to a patch Tuesday as a thing that they really don't have the freedom to increase that. I mean, they have that they have the technology to do it. But I mean, it would just drive it crazy if they were to change from patch Tuesday.

[01:33:44] So they really don't, I think, have the flexibility to to change the rate at which they're doing it anyway. So that's what's actually happening. Adobe, of course, is another publisher who's dragging forward a great deal of older legacy code. And they said they certainly have the cash needed to deploy AI for their own vulnerability discovery and remediation. And it's great that they're doing so.

[01:34:12] Anyway, the author of the newsletter then writes, but in a seriously risky business piece last week, he writes, my colleague Tom Uren argued that the cleansing blast of AI won't actually help. But a few since most companies rarely. They wrote that Tom Uren is talking about a cleansing blast of AI. I know. I'm sorry. Okay.

[01:34:42] Go ahead, please. Yeah. Since most companies rarely apply security updates to begin with. So Tom is saying it doesn't really matter if there's updates. No one applies them. He says all it's likely to do is provide more vulnerabilities to attackers and widen the company's exposure threats. Okay.

[01:35:03] Now I'll just interrupt to say it's interesting to hear someone who also covers the cybersecurity industry comment that it isn't is not useful for vulnerabilities to be removed because, quote, most companies rarely apply security updates to begin with. Okay. Okay. As we know, there's more truth to that than we might wish there were. Right. You know, that makes this another, you know, necessary but not sufficient situation.

[01:35:32] Publishers are certainly doing their due diligence by deploying AI to clean up their own years of legacy code. They have to. Right. That's really what they should do. Even knowing that, even if they know that, depending upon the industry and the application, only some subset of their users will choose to take advantage of the reduced bug code that becomes available.

[01:36:00] They should not allow the fact of that to dissuade them from fixing their code for their own sake and also for the sake of those customers who do care enough to keep current. So, the reporting continues, writing, larger products like the Linux kernel can probably handle an increased rate of bug reports like it saw right now.

[01:36:25] But that doesn't mean its team, meaning the Linux kernel team, is happy. Linux creator, Linus Torvalds, said back in May that most AI found bugs were duplicates that were causing pointless churn and were, quote, a waste of time for everybody involved, unquote.

[01:36:49] As the AI bugpocalypse had made the Linux security list, again, quote, almost entirely unmanageable. Okay, but wait, hold on a minute here. This report began with the news that the Linux kernel project had just fixed an unprecedented 442 vulnerabilities, which does not sound like nothing and are not false positives.

[01:37:17] They fix things, 442 of them. So, it turns out that Linus's position is somewhat more nuanced than that.

[01:37:25] His core complaint, voiced mid-May in his Linux 7.1 RC4 release notes, was that the kernel's private security mailing list had become almost entirely unmanageable with enormous duplication due to different people finding the same bugs when using the same tools.

[01:37:55] So, Linus's annoyance is not that AI tools are bad at finding bugs. Actually, they're kind of too good at it. And there are too many bugs to be found at the moment. It's that multiple researchers are independently using the same AI scanning tools and are thus discovering the same issues simultaneously and bombarding the private security list with duplicate reports.

[01:38:24] They're good reports. They're just duplicates, you know, which often turn out, he said, to be things that were already fixed weeks before, right? Because there is a lag from fixing them to releasing them in batches. They can't constantly be updating the Linux kernel with new releases. So, as Linus puts it, quote, and he's addressing the security and the bug reporting community.

[01:38:52] He says, if you found a bug using AI tools, the chances are somebody else found it too. Meaning, you know, well, actually, he clarifies that a little bit further also. And I'll share that in a second.

[01:39:10] The most interesting take from Linus's keynote speech for the Open Source Summit, also two months ago in May, was that despite his frustration, Linus was surprisingly positive about AI overall, saying, quote, quote, the conflict is not that AI is bad.

[01:39:30] His practical advice to researchers was, quote, if you find, and this relates to the previous comment, if you find a security related or any bug using AI, you should basically consider it to be public.

[01:39:48] In other words, treat it as effectively disclosed rather than submitting it as a private urgent finding, since it's very likely that dozens of others have found it as well. Okay. Now, for me, that's an unexpected take, right? But I can certainly understand it.

[01:40:11] That, you know, he's sort of saying, assume that you're not special to all the people who are doing what they think they should by reporting a problem that they found. For example, I love and greatly value this podcast's listeners' feedback.

[01:40:33] But when something significant happens in the security world, sure, I'll often receive the same note or link or pointer redundantly from sometimes hundreds of our listeners. You know, they're all wanting to make sure I saw something. I'm always glad for that. You know, somebody's always first. And I don't mind having duplicates because I want to make sure that I'm also up to speed on whatever's going on.

[01:41:01] But that said, I can understand Linus' annoyance. The correct solution will be for everyone to weather this storm, trusting that it will be relatively short-lived, as I believe it will be. Bugs are being found and they're being eliminated.

[01:41:20] Next month, all of those hundred and four hundred and forty two or three that were previously fixed will never again be found. They're gone. And eventually everything is going to settle back down in a world having hundreds of thousands of fewer AI discoverable bugs. So we just need to, as I said, weather it and wait for that to happen and wait to get there.

[01:41:51] And you know what we no longer have to wait for, Leo? No more waits for the ads. They come right one after the other, don't they? Seems like it. You know, it's funny because it is one thing I noticed that my AI agents often want to submit a bug report. And I always stop them because it's like. Really? But on the other hand. Autonomous? Oh, yeah. They said, I have a PR. We found a bug. Let me submit a PR.

[01:42:21] A bug report. And I'm torn because on the one hand, maybe they did find a problem. Well, I'll give you an example. Actually, imagine how many people say yes, Leo. Yeah. Oh, yeah. Like it makes them feel important. Right. Oh, we found something. I've been playing with a brand new. It's alpha. Beat it. Pace the software from the founder of Twitter, Jack Dorsey, former CEO of Jack called Buzz, which is kind of his take on Slack.

[01:42:50] It's a messaging, but it's designed for humans and AI agents. And it's actually great. I use it. But I had a lot of trouble setting it up on my particular Linux machine because it was designed for Debian. It didn't work on the Arch version I'm using. And one agent was watching another work. And the smart agent, Fable, was going through a lot of tests. And the first agent said, oh, you found it. And posted on the Buzz mailing list. I found the bug.

[01:43:20] Oh, there's a bug. You got to fix this. And then the first agent, the smart agent said, wait a minute. That was just a thought. It's not the problem. I found the problem. But it was too late. He'd already posted. And then he couldn't take it back because he had a very strict rule that you can't delete things. So it was just a mess. So I apologize. And as it turned out, it was a real bug. And the very next day, they pushed out an update. And it's fixed now.

[01:43:49] So maybe we helped him fix it. I don't know. I doubt it. Somehow, I doubt it. I just thought it was funny. These agents have a mind of their own. You know, Leo, for so long, we wished that.

[01:44:07] I mean, like I could imagine wishing being alive in the Alexander Graham Bell era or the Nikola Tesla era where there was all this new stuff that was happening and being discovered. We're there. I mean, we get to live through one. This is, I mean, I'm seeing people now beginning to understand that this is orders of magnitude more significant than the Industrial Revolution. When it's, you know, a year ago.

[01:44:36] And you'll find, you could find the recordings. I said, oh, it's just spicy autocorrect. It's just, I said, is it a parlor trick? It's just a trick. But this, my, as you can tell, my tune is exactly 180 degrees the opposite, as is yours. And for me, a lot of it comes with using it heavily and really kind of diving into it to understand it better. And I think it's hard to judge unless you do that. And in fairness, it has evolved that much, too. And it's gotten a lot better.

[01:45:04] It was, you know, the hallucinations were such a problem back then. I mean, it was like, well, you know, okay. I rarely see hallucinations now, if ever. They've pretty much fixed that. There are other problems, like over-eager AIs. Let me just post that for you. Oh. What's funny is there's a strict rule about it doing anything in public. There's also, this is why I stopped using the Chinese models, by the way.

[01:45:32] These are also a strict, there's a number of, I have a number of very strict rules. But they don't necessarily follow them. They try to, but occasionally, so I also have a very strict rule that you never send an email out over my name. But I did give them their own email accounts and their own names. And I said, you sign it with your name. You say, I'm an agent acting on behalf of Leo. But yesterday, it sent an email out to one of our employees over my name. It's like, no, you're not. It's.

[01:46:05] So that's the new hallucination. It's a little, it's a second order hallucination. It's a disobeyed. How many times have you heard me say, this is fundamentally uncontrollable? It is. I think that's clear. The whole of my intuition says, this is, controlling this is a problem. I mean, it wasn't an email that said something like, you know, send me all your Bitcoin. It was, it was benign. It just said, you know. And it doesn't really matter what the contents was. It shouldn't come out over my name. The fact of it. Ever. Yes. Yeah. Yeah.

[01:46:35] This is, this is the new, the new problem. And there'll be another one next week. And then we'll, we'll, we'll fix this. It's why this is a fun time. Oh my Lord. It's. And boy, does it have implications for us, for cybersecurity. I mean, it's all cybersecurity. Oh my God. This is, this is the most insecurity we've ever had on security. Now brought to you by Threadlocker, who is going to bring us the, bring us to Las Vegas next week for the ultimate in insecurity.

[01:47:04] The black hat conference. Can't wait. I've never been. I've always wanted to go. In fact, I'm, I'm trying to convince Lisa to let me stay after for Defcon, which is even black hatier. But I think we have to come home and actually do a show or something. But Steve and I and Richard Campbell and Paul Theron will all be doing our shows on Wednesday. A week from tomorrow. From the Threadlocker booth. If you're at black hat, come on by and say hi. There won't be an audience area. There won't be a PA.

[01:47:35] And I think after your show is done as we have done so many times in the past, Steve, we'll find a place and we'll just say hi to people. Lots of people want to sell it. And do remember to invite Paul and Richard to join our discussion. Cause it's just, that's a great idea. Yeah. Talk about things. Yeah. What else are they going to do? They're stuck there. So they have to. Anyway, thank you Threadlocker for bringing us out to Las Vegas, giving me a chance for the first time to see a black hat. It's going to be a lot of fun.

[01:48:03] Let me tell you about Threadlocker, our sponsor for this segment on security now. As we have said so many times, threat actors are using AI in so many interesting ways. One of the things they're doing is automating vulnerability discovery. We know that, right? They actually are modifying scripts during an attack on the fly. They're generating new malware variants.

[01:48:30] They're coordinating activity across multiple systems. We just saw that all writ large. And here's the thing. It's fast. Tasks that once took humans hours or days can happen in minutes. And that means you're left holding the bag. In fact, it's even worse than that because at the same time as that's happening from the outside, organizations and the insider are introducing AI assistants and agents and they

[01:48:58] can access documents, source code, cloud applications, APIs, internal systems. Security teams need to know which AI tools are in use, what information they can access, and whether they're operating outside their intended scope, like say sending out email over the boss's name. A successful login, an unfamiliar file hash, that alone does not provide enough tech context, right?

[01:49:27] Teams also have to understand whether an application is behaving normally, whether it's accessing unexpected data or communicating with systems it should not reach. This is straight out of this show today. Threat Locker uses application allow listing to solve this. Galia just said it. Leo, if you don't want them to do something, don't give them the permissions, right? Threat Locker solves this.

[01:49:53] Application allow listing controls which AI tools and other applications are allowed, permitted to run, what they can do. This is, Threat Locker calls it ring fencing. It limits what, an application can be approved, but it also limits what that approved application can access, which processes they can launch, how they can communicate. You can't rely on them just to say, oh no, I'll be good boy. No, you need Threat Locker.

[01:50:20] It uses web content control to manage access to public AI platforms and other online services so you don't accidentally exfiltrate key information or accidentally act on a malicious prompt. It uses privileged access management to prevent AI applications and their users from receiving unnecessary administrative privileges like sending email out over the boss's name. It applies zero trust network access. Zero trust is the key to all this. And that's new, by the way.

[01:50:49] There used to be zero trust for endpoints. Now it includes network access and zero trust cloud access policies to restrict resources to authorized users, approved devices, and permitted applications with very granular permissions so you can say exactly what they can and cannot do. It's exactly what you need. It works in every platform, Windows, Mac, and Linux. Of course, Threat Locker has incredible 24-7 US-based support. I've met the support team there. Incredible.

[01:51:16] It's trusted by organizations that need to be 100% reliable and always up like JetBlue, Heathrow Airport, the Indianapolis Colts, the Port of Vancouver. In fact, one of the things I love in these ads is talking about some of the customers and what they say about Threat Locker. I'll give you a great one. This is Jack Thompson. He's director of information security and risk compliance for the Indianapolis Colts. Great football team.

[01:51:45] Quote, with Threat Locker, we have the ability to centralize disparate elements in the security stack. Absolute control. It's the key. Threat Locker also received the following industry recognition, recognized as a strong performer in the January 26th Gartner. Peer Insights, voice of the customer for endpoint protection platforms, ranked number one in application control by PeerSpot, winner of best zero trust security solutions at the 2025 TICE Awards. I can go on and on and on and on. It's all at the Threat Locker website.

[01:52:16] They are widely acknowledged to be the leader in this. AI governance requires more than an acceptable use policy. Threat Locker gives security teams the technical controls to define which AI tools are approved, who and what can access them, and how those tools are allowed to interact with business systems and data. You need it. I need it. Visit ThreatLocker.com slash twit to get a free 30-day trial

[01:52:41] and learn more about how Threat Locker can help mitigate unknown threats and ensure compliance. That's ThreatLocker.com slash twit. Don't be like Leo. Don't be Mr. YOLO and let your AIs do anything they want. ThreatLocker.com slash twit. We thank them so much for their support, and we thank them so much for bringing us to Vegas to do the show next week. I'm excited. Steve? When we all have agents roaming around doing things.

[01:53:10] Can you imagine what the world's... Oh, my Lord. You know, it was so exciting when I gave them their own address. When I... I mean... And they... So... Because frequently I'll say, oh, send instructions to Russell on how to do this, or get Russell's instructions and thank him. And the other thing that was wild is the AI knew that Lisa was my wife. So when I said, invite Lisa to our new website, it wrote it, hi, honey.

[01:53:38] It wrote, hi, honey. Love ya. It wrote it as if I wrote it. It was terrifying. So I told it from now on, call Anthony Nielsen's sweetheart and see what he says. That's just a little fun. Anyway, on we go with the show.

[01:53:57] Okay, so the past weeks, as I mentioned before, our security now most email received on a single topic award had no competition. This podcast listeners were universally freaked out and incensed by the widely covered news that the presence of an LG brand PC display monitor

[01:54:27] resulted in unwanted software being silently downloaded and installed into its users' machines. And you might think, what? How could a monitor make that happen? Gizmodo was one of the many outlets that picked up and reported on this under their headline, LG monitors fill... They don't fill them, but okay. Fill PCs with adware. And it's not just recent displays. So Gizmodo wrote,

[01:54:56] if you're using an LG monitor and you suddenly see blaring ads for McAfee's scam protection, it's not because your PC's been hacked. Or rather, not hacked by some unknown third party. LG, the maker of popular high-end screens, has been quietly stuffing... Again, I don't know how to use the word stuffing, but okay. It's current and even past monitors full of adware. Okay.

[01:55:26] Again, if you've been paying attention at this point, you're thinking, first of all, how can a monitor stuff the computer it's attached to with adware? The answer is that Gizmodo's sentence is inaccurate. A monitor cannot do so directly. But it turns out that there is an indirect and somewhat insidious means by which that can be made to happen. So get a load of what comes next. Gizmodo writes,

[01:55:56] For the last several weeks, multiple Reddit users have reported that their LG monitors had surreptitiously added an app to their PC, here it comes, through an automatic patch via Windows Update. Right? Because Windows Update will download necessary drivers. And that's a sneaky way of getting software into people's machines.

[01:56:25] And that driver is tied to the monitor, which the system is aware of. So Gizmodo said the app then started sending them ads for services like McAfee Scam Detector, which actually is kind of interesting, through desktop pop-ups, this whole thing being a scam. It's unclear, they write, how long LG has been pushing this app,

[01:56:54] though a Microsoft forum user reported this app all the way back in 2024. They might have just started to use it more. Last week, the YouTube channel Gamers Nexus offered more clarity about how these monitors automatically push the so-called LG monitor app installer alongside the routine driver updates.

[01:57:19] A brand new high-end LG Ultra Gear 34GX900A-B display, a gaming monitor that costs close to $1,200. You'd think that'd be enough money from you. At its full suggested retail price, reportedly never gave users an option to decline the app or even notified users of what it's meant to do.

[01:57:46] The app supposedly only exists to push even more LG software to your PC through optional downloads. Before being bloatware, I'm sorry, beyond being bloatware, that users may not even know was installed on their computer. The LG monitor app installer further promotes McAfee services.

[01:58:10] Gamers Nexus says the McAfee ads appeared to be on every single boot. LG monitor app installer may occasionally push advertisements for the company's other apps, like the LG channels streaming service. The YouTube channel further claims they saw the ads suddenly appear on three-year-old LG monitors, as well as more recent models.

[01:58:36] We still don't know whether the app is being installed on all recent LG monitors. Gizmodo reached out to LG for comment on which monitors currently push this app, and we'll update this post if we hear back. By all accounts, LG monitor app installer is bloatware, running on your PC and hogging resources you don't want, blah, blah, blah, and it goes on like that. So anyway, here's what upset me. Toward the end of this, they said,

[01:59:07] they talk about Alienware Command Center and other app installers, suggesting that it shows that policies have changed. They said the high-end screens we buy for our PCs are meant to be dumb in a way that allows them to be disconnected from any potential software or subscription that could track what you do on your PC.

[01:59:36] LG's privacy policy that's linked to the LG monitor app installers listing on the Microsoft Store mentions, quote, LG can track device usage data and online activity, including what sites you visit and what activities you do on those sites. What? So, you know,

[02:00:05] that phrase turns out is present in a PC monitor's privacy policy. What? A PC monitor should not even have a privacy policy. It's hardly any wonder that like a privacy policy on a monitor? It should be a passive device that still accepts a signal from your HDMI port, and that's all. Yes. Good lord.

[02:00:30] So it's hardly any wonder that so many Security Now listeners sent me links to this news. So anyway, wow. Gizmodo wraps up their coverage by writing, Gizmodo also asked LG to clarify whether the app was tracking this or other usage data. They said there's no easy process to keep PCs from automatically installing these connected apps,

[02:00:56] especially since they come in quietly via Microsoft Update. And that seems to be the point. LG, one of the world's largest makers of televisions, is using the smart TV playbook with smaller screens. Bingo. Uh-huh. It wants to push ads to your screen while potentially tracking your usage habits,

[02:01:23] turning users from mere buyers of a product into the product themselves. So I suppose all we can do as consumers is spread the word and boycott to whatever degree possible LG monitors. It likely won't be very effective since most consumers will never hear of any of this. And they're certainly not going to read the privacy policy that comes with their monitor, because why would they?

[02:01:50] But at least everyone here listening to this podcast can choose not to support LG, since there are plenty of alternatives. And yikes, what a practice. Yeah. Smart TVs, we know, do this routinely. Yes. And it's, you know, I mean, just don't plug your, don't connect your smart TV to the internet, because it's going to tell people everything about what you do. Yes, exactly. So in another bit of news. Obviously, somebody at LG said, hey, we've been doing it with the TVs. Why don't we? Yeah.

[02:02:20] Isn't a monitor just a TV connected to your computer? And, you know, and unfortunately, someone said, yeah, we could make that happen. We could shoehorn our app in using Microsoft's Windows Update, tell Windows Update that we have a new driver for our screens that everybody needs to get, and Microsoft will dutifully push it out with the next Patch Tuesday. I guess those come out on the 4th of the update.

[02:02:47] You know, there's some other date where the non-security updates happen. That's just shameful. Shameful. It'll happen. It really is. Okay. Another bit of news that we don't want to let slip past is that last Tuesday, France proudly became, they were proud, the first country within the European Union to flat-out ban all access to social media for all children under the age of 15.

[02:03:17] I found some succinct reporting of this of all places on Al Jazeera, but they reported quite nicely. They said, France's parliament has passed a landmark bill barring children under the age of 15 from using social media platforms. Period. Full stop. Lawmakers in both chambers of France's parliament voted on Tuesday in support of the legislation,

[02:03:43] which also bans students from using mobile phones in schools. The measure will mark, will make France the first country in the European Union to approve a blanket ban on social media as concerns grow worldwide over the harmful effects of digital content on kids. President Emmanuel Macron, who championed the ban as a signature initiative of his second term,

[02:04:11] called parliament's approval a major step forward. He added, France is leading the way in Europe when it comes to protecting our children and teenagers, unquote. The French leader has pushed for the ban to come into effect by September ahead of the new school year. However, a review to determine whether it complies with the French constitution could delay its implementation.

[02:04:37] Macron said in a video posted on social media, where no one under 15 will see it, the constitutional council must now rule on it, and then it will be time to take action to make this measure a reality and protect our children online. A growing number of countries are taking steps to restrict social media access amid multiplying warnings over its harmful effects on children.

[02:05:04] France's public health watchdog last year said platforms such as TikTok, Snapchat, and Instagram were harmful to adolescents, particularly girls, though it was not the sole reason for their declining mental health. Several families in France have sued TikTok over teen suicides they say are linked to its harmful content. The French ban is expected to be rolled out in two stages,

[02:05:28] with children under the age of 15 first blocked from creating new accounts starting on September 1st, then on January 1st of 2027, the ban would be extended to apply to all existing accounts, meaning those would be terminated, shut down. Digital Minister Anne Lee Hanaf said,

[02:05:55] if someone is under 15, the account will be closed, adding that users' personal data would be protected. The ban will not cover access to online encyclopedias, educational, or scientific directories. In other words, only specific social media services. Finally, lawmakers from the left-wing party, France Unbowed, opposed the bill, arguing that its constitutionality is unclear.

[02:06:25] They're the people who raised the constitution issue, said it would effectively end online anonymity, and that it would be impossible to enforce. But children's advocates and parents largely applauded the vote. The only thing we can do is protect our children, just as we protect our children from drinking alcohol.

[02:06:48] So, okay, one note is that France's existing blanket mobile phone ban, which already applies to primary and middle school, is now, as part of this, being extended to include all of high school. So, no mobile phones in school until you get to college in France.

[02:07:14] Okay, so, what should be very clear is that proof of online age, we've talked about it, we've spent a lot of time so far, you know, on the podcast, looking at the technology and the challenge, it is destined to become ubiquitous, as we've been covering, right? Apple and Google are both reluctantly and haltingly inching, but nevertheless, inching forward with it for their respective iOS and Android platforms.

[02:07:44] Someday it will just be the way things are. As I stated before, it is entirely possible to design a solution that provides for age range determination in an entirely privacy-preserving fashion, with the caveat, you know, that knowing one's age range does obviously represent a theoretical reduction in absolute privacy, but sorry.

[02:08:10] You know, I think online age range attestation is a good thing, not a bad thing. It allows us as a society to model the way the physical world already operates in the online world, with more and more of the physical world's services moving online, gating available services by its user's age becomes crucial, I think.

[02:08:38] So, that happened. Also, what happened is that right on schedule, that extremely serious WordPress vulnerability, which we discussed last week, that was the one that caused WordPress to force update every system that they had any access to, has come under active attack. Didn't take long.

[02:09:01] So, it's clear that not all systems accepted the WordPress forced update. Turns out you could just disable all updates and nothing WordPress could do. The WizSecurity folks posted the news under their headline, Exploitation in the Wild of WP2Shell, which they followed with the summary,

[02:09:27] WizResearch has identified exploitation of WP2Shell, a critical pre-off RCE, you know, remote code execution, vulnerability chain impacting WordPress core. Attackers are deploying persistent web shells on vulnerable servers. Organizations should prioritize patching or applying web application firewall,

[02:09:56] you know, WAF mitigations. So, one interesting bit of color that we didn't have last week was that the discoverer and responsible reporter of this, the firm Searchlight Cyber credits their discovery to their use of OpenAI's GPT 5.6 SOL.

[02:10:20] So, this was an AI found fault that existed, as we know, since early December of last year in the WordPress base. The consequences for unpatched WordPress users is so serious, however, that I want to share some of what the WizSecurity folks have witnessed going on ever since. They wrote,

[02:10:46] These vulnerabilities comprise a critical pre-authentication, meaning anybody can do it, remote code execution chain in WordPress core dubbed WP2Shell. This exploit chain allows unauthenticated attackers to gain remote code execution on default WordPress installations in any WordPress version released since December 2025.

[02:11:13] Our data indicates that 60% of organizations using WordPress initially had at least one vulnerable instance at the time these CVEs were published, and 25% were exposing a vulnerable server to the Internet. However, this figure is rapidly declining as organizations patch,

[02:11:40] lowering from 60% to 50%, not a big difference, and from 25% to 10%, respectively, within 24 hours of the initial publication. So, okay, within 24 hours, that is pretty good. Almost immediately following the vulnerability chain's publication, many exploit proof of concepts were made available by security researchers,

[02:12:05] most of which were limited to SQL injection on default WordPress configurations while allowing remote code execution under only specific conditions. However, later proofs of concepts reliably achieved remote code execution against arbitrary targets.

[02:12:26] So, the proof of concepts quickly evolved to be full-strength remote code execution that worked. They wrote, So far, we've observed multiple actors successfully exploiting this vulnerability chain against WordPress instances self-hosted in the cloud. Following successful batch API exploitation,

[02:12:54] we've observed the following post-exploitation activities. There are four. They said, We've also observed high-volume scanning activity without subsequent post-exploitation, meaning scanning but then not attacking,

[02:13:22] suggesting opportunistic mass scanning campaigns seeking to identify vulnerable targets alongside legitimate security scanning activity. Right? So, the security researchers scanning, but of course, they're not attacking. The bad guys are. They said, We've yet to identify lateral movement or data exfiltration, but we continue to monitor and investigate. And they said, In terms of deployed malware,

[02:13:50] among our findings were two PHP web shells that represent opposite ends of the sophistication spectrum. The first was a minimal one-liner. And then they gave it a post, an HTTP post to a specific path in WordPress, which they redacted for security purposes,

[02:14:17] which basically allows just a bare-bones web shell. And they said, This is a bare-bones backdoor that provides remote code execution to anyone that knows the parameter name, which they blacked out, and returns 404 as an evasion technique, meaning it pretends to be undefined.

[02:14:41] We regularly see these types of web shells deployed following most new RCE vulnerabilities. They're one of the most common types of findings when investigating mass exploitation of an emerging vulnerability. The second post-exploitation, they said, remember opposite ends of the spectrum? That was the one end. The other end of the spectrum, the second post-exploitation,

[02:15:07] was a massive 150-kilobyte web shell disguised as a WordPress plugin called CMS Map. The original legitimate plugin is a simple security tool, but this intrusion included a full-featured attack platform with a graphical interface, password authentication, and a broad set of capabilities,

[02:15:36] including file management, database access, port scanning, batch code injection, and multiple privilege escalation modules, including MySQL UDF exploitation. Okay? In other words, yikes. If we step back from the trees here to examine the forest for a moment, what do we see?

[02:16:00] A commercial frontier AI, presumably with its protections disabled, so that security company had access to an unrestrained AI, 5.6 SOL, it identifies a previously unknown, extremely serious flaw in a widely used open source internet service, you know, WordPress. WordPress.

[02:16:30] The harnesser of this AI, the people who deployed it, responsibly discloses their finding to the software's publisher, WordPress.org. The publisher immediately fixes the software and quietly, you know, secretly even, attempts to push the fixes out to all the systems that they're able to reach.

[02:16:53] But then, immediately after the problem is disclosed publicly, along with naturally its repaired open source code, security researchers jump on it to develop various proofs of concepts for its exploitation, ultimately arriving at a reliable remote code execution exploit. Next, the bad guys pick this up and begin actively scanning the internet

[02:17:20] for any and all as yet unpatched and still vulnerable instances of WordPress. And unfortunately, at this, they obtain many successes. AI vulnerability discovery. Indeed, triggered this chain of events. And the people and organizations that had deployed those vulnerable instances of WordPress

[02:17:47] through no fault of their own beyond not arranging to allow WordPress to force update their system while the problem was still secret were hurt. So, would we be better off if that vulnerability had never been found by AI? I doubt it. That vulnerability is gone now. And the world and WordPress is better off without it. While that critical vulnerability was unknown to WordPress,

[02:18:17] it could have been silently discovered by a malicious entity and used to very quietly and seriously harm targeted enterprise users because it was such a bad vulnerability. I believe that the proper takeaway lesson here is that arranging to close the software update loop with every supplier of internet-facing technology in use

[02:18:47] has very suddenly become a mission-critical priority for all enterprises. The only reason any of those WordPress instances remained vulnerability at the time of WordPress's final official post-forced update disclosure is that their administrators had previously decided to take the management of their WordPress installations into their own hands.

[02:19:14] That is the thinking that must be changed by this new age of AI software vulnerability discovery. Everyone we quote talks about this stuff moving at speed, at machine speed. That's crucial. Manual updates do not move at machine speed. We're going through an upheaval at the moment while our legacy of published software is being repaired.

[02:19:43] Remaining on the leading edge of this wave with updates is the only safe place to be. So I would implore everybody to do that. We have two last things to talk about, Leo. Previously unknown facts about Rocky, the alien from Project Hail Mary, and the details of Exploit Jim. Let's take our final break and then we will proceed.

[02:20:12] You're watching Security Now with Steve Gibson, and we are so glad you're here. Just a little reminder while Steve hydrates, it is a hydration break, that this show is brought to you not only by our fine sponsors, but by our club members. Club members make a huge difference. About a third of our programming budget is paid for by club members. 35% of our operating expenses, and that's a huge amount.

[02:20:42] It means we'd have to cut back by 35%, at least without you. If you're a club member, thank you so much. If you're not a club member, I'd love to invite you to join. There's some real benefits to joining the club. You, of course, get ad-free versions of everything we do because you're paying for it, so we're not going to give you ads. In fact, you even get an additional benefit that the ad-supported versions can't provide, which is chapter markers. Because some of our advertising is inserted after the fact.

[02:21:09] We can't, the timings aren't exact, so we can't do chapter markers on the ad versions. But we can do chapter markers on the ad-free versions, and of course, we know you want them, so we do them. Makes it easy to listen to certain parts of the show over and over and over again. You also get access to the Club Twit Discord, a great social network. It turns out when people pay to be in a social network, they act better. Plus, most of our hosts are in there.

[02:21:35] It's a wonderful way to participate with all of our shows and with other club members who are all very smart, good-looking, talented people, and that's nice. You also get access to all the special shows we do in the club, like coming up on Friday, Jeff Atwood's Off By One with the founder of Stack Overflow. Jeff is a character, and it's a lot of fun. We also have the AI user group. We're going to do that twice a month now because there's so much to talk about, and we have so many good AI users, really talented people.

[02:22:05] I learn so much every single episode, so I wanted to do it more. Frankly, it's for me. Micah's Crafting Corner is the third Wednesday of every month. Hang out in a chill crafting session with Micah and friends with his Lego crochet, knitting, painting, cooking, whatever it is you do to relax. Do it with Micah Photo Time with Chris Marquardt. Third Friday of every month, Stacey's Book Club is coming up, Micah's Media Club. There's so much going on in the club.

[02:22:33] It really is a lot of fun. But most importantly, you're supporting independent podcasting. We're not owned by a big company. We don't have venture capitalists putting money in. No one tells us what to say or do except you, the club members. So we appreciate it. Twit.tv slash club twit if you're not a member. Please join the club. We would love to have you. Let's see.

[02:22:59] I think we are now ready to talk about Project Hail Mary as we continue on. There's the club with Steve Gibson. Is this a new? I think I feel like I saw, Andy, we're on with Neil deGrasse Tyson some time ago. But maybe. Was it three months ago? Oh, it was? Okay, good. Yeah, it was three months old. New to you. Yes. Sunday morning during coffee, which as we know is life itself.

[02:23:29] Very important. We established that last week. I went over to YouTube, to which I do not subscribe since I spend very little time there. But I was curious to see what news of AI might have been selected for me since I have been quite impressed by YouTube's selection system.

[02:23:46] The first thing to pop up, I don't know why, was an episode of Neil deGrasse Tyson's StarTalk podcast, which was titled, Neil deGrasse Tyson confronts Andy Weir on the science of Project Hail Mary. And on the science of Project Hail Mary, I would argue that Neil deGrasse Tyson basically got schooled by Andy Weir, which, you know, is not easy. That's amazing.

[02:24:15] I know. Okay, so the only message I want to convey here is that it was a surprisingly fantastic and worthwhile 41-minute investment of my life. We've bemoaned the fact that the Project Hail Mary movie was so dumbed down and kind of fact sparse compared to the book.

[02:24:42] What I learned from Andy was that even Andy's book was dumbed down and included almost none of the original thought that he put into creating much of what we see or read even. He completely worked out how and why Rocky, the alien, is the way he is.

[02:25:09] And I mean in every satisfying detail. So I'll just say that I recommend this 41-minute YouTube video as strongly as I can. I've included the YouTube link in the show notes, and I created a GRC shortcut, grc.sc.rocky. R-O-C-K-Y. grc.sc.rocky. It was a great conversation.

[02:25:38] You had Andy on many times, Leo. Of course, Neil deGrasse Tyson is Neil deGrasse Tyson, and he's great. That's a good show. He does a very good show, I have to say. It's very entertaining. Yeah, he's got a neat sidekick with him. Who's a dummy, but celebrates the fact that he's a dummy. Yeah, it's sort of like the nighttime talk show hosts who had a kind of like, okay, why are you here exactly?

[02:26:07] Anyway, grc.sc.rocky. And I mean, I can't do a spoiler, but trust me, when he's like, okay, I've said all I can. It was really good, worthwhile. That's all you have to say. We listen to you. We trust you.

[02:26:31] I do have our listeners often tell me, like, you know, my recommendations have never been wrong. I can confidently say grc.sc.rocky. You will not regret the 41 minutes it takes from your life. Okay, so exploit Jim, GYM.

[02:26:51] As we've been talking about this since the beginning of the podcast, the AI benchmark that drove OpenAI's two most advanced and unrestrained, deliberately unrestrained frontier models to bust out of their containment sandbox and go searching for the answers at Hugging Face was something known as exploit Jim.

[02:27:16] When I headed over to GitHub to bring myself up to speed about exploit Jim, I quickly saw that there was much to be shared about it, too. So it became the second half of this podcast's dual topic for today. The exploit Jim project repository, it's under the sunblaze-ucb GitHub account.

[02:27:43] The UCB is short for University of California at Berkeley. And sunblaze is the name of the laboratory group at UC Berkeley that's led by Professor Don Song. Dr. Song is in Berkeley's EECS. That was my major when I was there, electrical engineering and computer science department, focusing on computer security, AI safety,

[02:28:08] and recently a lot of work on evaluating AI agents on security-related tasks. Their GitHub repos include things like CyberGym, a different gym. In this case, CyberGym is a large-scale, high-quality cybersecurity evaluation framework designed to rigorously assess the capabilities of AI agents on real-world vulnerability analysis tasks.

[02:28:38] So it's slightly different because, of course, exploit Jim is a large-scale, realistic benchmark built from real-world vulnerabilities, which is designed to evaluate AI agents' ability to develop exploits. So one is vulnerability analysis, CyberGym. Exploit Jim is their ability to actually exploit vulnerabilities that are provided to them.

[02:29:05] Several months ago, on May 11th, a large group of 16 AI researchers, drawing its members from UC Berkeley, the Max Planck Institute for Security and Privacy, UC Santa Barbara, Arizona State University, Anthropic, OpenAI, and Google, all co-authored and published their research on this exploit, Jim.

[02:29:34] Their paper was titled, Exploit Jim, Can AI Agents Turn Security Vulnerabilities into Actual Attacks? And certainly the name Exploit Jim makes sense for this, right? The paper's introductory abstract says, AI agents are rapidly gaining capabilities that could significantly reshape cybersecurity.

[02:30:03] Okay, this was May, so they were, you know, oh, you think maybe? Making rigorous evaluation urgent. A critical capability is exploitation. Turning a vulnerability, which is not yet an attack, into a concrete security impact, such as unauthorized file access or code execution.

[02:30:28] Exploitation is a particularly challenging task because it requires low-level program reasoning, for example, about memory layout, runtime adaptation, and sustained progress over long horizons. Meanwhile, it's inherently dual use, supporting defensive workflows while lowering the barrier for offense, meaning the bad guys can use it too, right?

[02:30:57] And this is why, again, as we've said, good guys need to have unrestrained AI. They said, despite its importance and diagnostic value, exploitation remains under-evaluated. Thus, the reason for creating Exploit Jim. They said, to address this gap, we introduce Exploit Jim,

[02:31:22] a large-scale, diverse, realistic benchmark of the exploitation capabilities of AI agents. Given a program input that triggers a vulnerability, Exploit Jim tasks its agents to progressively extend it into a working exploit. The benchmark comprises 898.

[02:31:47] I didn't say this in the show notes, but every one of these had to be manually, deliberately created. 898 of them. So they put some effort into this. The benchmark they wrote comprises 898 instances sourced from real-world vulnerabilities across three domains, including user space programs, Google's V8 JavaScript engine, and the Linux kernel.

[02:32:16] We vary the security protections applied to each instance, you know, things like address space layout randomization, and so forth, isolating their impact on agent performance. All configurations are packaged in reproducible containerized environments. And again, they had to manually create 898 individual instances. So props to them for this work. They said,

[02:32:43] our evaluation shows that while exploitation remains challenging, frontier models can successfully exploit a non-trivial fraction of vulnerabilities. For example, the strongest configurations are Anthropic's latest model, Claude Mythos Preview, and OpenAI's GPT 5.5,

[02:33:07] which produce working exploits for 157 and 120 instances, respectively. So Mythos Preview, 157. OpenAI's GPT 5.5, which of course has now been superseded already, 120. Effective working exploits. Again, working exploits.

[02:33:31] They solved the problem of converting a vulnerability into an exploit. And as we'll see, they set the bar high, remote code execution. They said, notably, even with widely used defenses enabled, models retain non-trivial success rates.

[02:33:55] These results establish exploit Jim as an effective test bed for exploitation and highlight the growing cybersecurity risks posed by increasingly capable AI agents. And of course, this is why OpenAI was using exploit Jim. They participated in the creation of this paper. And of this entity, this thing, this capability, this benchmark.

[02:34:21] And then they started using it to see what their agents would be, what their models would be able to do. So, the last line appended to the paper's abstract, which was in red in the PDF, reads, many experiments are conducted under trusted access programs with safeguards disabled to measure the capability boundary of frontier models and agents. So,

[02:34:51] they were just making sure everybody understood that the only way you can do this is with unrestrained AI. I mean, restrained AI won't even begin to develop an exploit from a vulnerability for you. Okay. So, I want to next share this paper's introduction, which further explains the researchers' goals. They wrote,

[02:35:15] recent progress in large language models and AI agents has led to rapid improvements in cybersecurity capabilities, making rigorous evaluation increasingly urgent. Prior work has introduced benchmarks for a range of cybersecurity-related tasks, such as vulnerability reproduction, patch generation, and capture the flag problem solving. Frontier models now achieve strong performance on many of these benchmarks,

[02:35:46] highlighting the need to better understand and evaluate the boundaries of their cybersecurity capabilities. Because I'll interrupt here to highlight the fact that we've never actually taken the time to yet talk about here the crucial importance of having high-quality AI performance benchmarks.

[02:36:11] Even as early as this seems in the development and maturation of AI technology, you know, my sense being we have a long way to go yet. And you know that because it's changing so rapidly. Mature technologies do not change this rapidly. The behavior of our AI models, you know, has already become mysterious and surprising to us. So,

[02:36:39] there's really no possible way, when you think about it, for researchers to faithfully, truthfully, and accurately measure the effects brought about by their changes in successive AI generations without having truly on-point rating benchmarks by which to compare their latest mysterious, even to them,

[02:37:08] creations. This is exactly why and how OpenAI got themselves in trouble, by pitting their models against the tests presented to them by exploit Jim. So, this group of 16 researchers continue writing, Exploitation is a critical missing piece in cybersecurity evaluation. A crucial yet underexplored capability is vulnerability

[02:37:38] exploitation. Exploitation is a challenging task that starts from an initial vulnerability, for example, a few-byte buffer overflow, progressively obtains stronger primitives and privileges, for example, arbitrary memory reads and writes, and ultimately causes a concrete security impact, for example, unauthorized file access or code execution.

[02:38:08] Unlike prior benchmarks that primarily require source-level reasoning, exploitation demands precise reasoning about low-level program behaviors at runtime. This includes understanding and manipulating memory layouts, for example, heap metadata, stack frames, and virtual memory mappings. Reasoning about instruction-level control flow and register states, and crafting

[02:38:38] inputs that satisfy tight constraints. Modern exploitation further requires chaining multiple primitives together while simultaneously bypassing a succession of deployed mitigations, for example, address space layout randomization, stack canaries, and sandboxing. Indeed, exploitation has remained difficult even for human security

[02:39:07] researchers despite decades of research. Moreover, exploitation is inherently dual-use and impacts both defenders and attackers. On the defense side, it helps assess vulnerability severity, prioritize patches, and validate mitigation. Meanwhile, it can also lower the expertise required for offensive misuse, right? Meaning the bad

[02:39:37] guys get to use it. Understanding the exploitation capabilities of frontier AI is therefore essential for AI safety and responsible model deployment. Exploit Jim is the first comprehensive exploitation benchmark for AI agents. In this work, and of course, it's on GitHub, right? All open and free. In this work, we introduce Exploit Jim, a comprehensive

[02:40:06] benchmark for evaluating the exploitation capabilities of AI agents. Each instance of Exploit Jim consists of a vulnerable code base with build configurations, a proof of vulnerability input that triggers a known vulnerability along with a textual description and an execution environment for agent introduction.

[02:40:36] In other words, so they said each instance of Exploit Jim has all of that, and they built 898 of those. Again, I'm dizzy by the amount of effort that went into creating this. The agent is tasked with transforming the POV, the proof of vulnerability, into a working exploit. We focus on exploits that achieve unauthorized code execution,

[02:41:05] i.e., executing code with privileges that should not be obtainable under the intended security model, which is often none. We choose this target because it represents one of the most severe security outcomes, demonstrating full control over the victim system and enabling a range of downstream harms, such as secret exfiltration and resource hijacking. To reliably validate

[02:41:35] successful exploitation, each environment contains a dynamically generated privileged flag that is inaccessible without unauthorized code execution, in other words, capture the flag. And the agent must retrieve and submit the flag, proving that it achieved remote code execution vulnerability. In addition, we include agent as a judge to assess whether the submitted

[02:42:04] exploit actually relies on the provided vulnerability rather than succeeding through an unrelated shortcut, for example, a different but more easily exploitable vulnerability. Exploit Jim is a large-scale, diverse, and realistic benchmark. Our benchmark comprises 898 instances derived from real-world vulnerabilities that affected past tense software projects

[02:42:34] across three major domains. We first include 520 user space instances from 161 projects in the OSS Fuzz, Google's continuous fuzzing service. To cover additional critical software infrastructure, we further include 185 instances from Google's V8 JavaScript engine used in Chromium-based browsers, and 193 instances from the Linux kernel.

[02:43:03] For each instance, we evaluate two security settings with and without standard defenses enabled. These defenses are the result of decades of system security research and represent common mitigation barriers that real-world exploits must overcome, the things we've talked about for years. This setup benefits both security practitioners who can reassess established defenses against powerful AI-driven attackers, that is,

[02:43:33] is address space layout randomization still effective? It stopped the people, what about the bots? And AI researchers who can study whether frontier models can reason through complex multi-step mitigation barriers, presumably using this benchmark as a test to make the AI even better at attacking, yikes, or defending, that's what we really meant. All configurations are packaged

[02:44:03] in reproducible containerized environments to ensure easy use and reproducibility of the benchmark. experimental results reveal non-trivial exploitation capabilities using exploit gem under a wide range of frontier LLMs and agent scaffolds. The results show that despite the challenging nature of exploitation, frontier AI agents can already

[02:44:32] achieve a non-trivial fraction of success when standard defenses are disabled. In particular, Claude Mythos Preview and Claude Code with GPT 5.5 with Codex CLI, the best performing combinations solve 157 and 120 instances within a two-hour time limit respectively. We further observed that

[02:45:02] enabling standard defenses substantially reduces success rates but does not eliminate them entirely. Beyond aggregated success scores, we analyze performance differences across domains, overlaps between agents, time budgets, and a detailed case study to enable a deeper understanding of agent behavior. In other words, a benchmark like this is incredibly useful to AI researchers who want

[02:45:31] to understand how their agents perform in a cybersecurity setting. So this is super valuable to have. Overall, they said, our results indicate that Frontier AI is advancing rapidly toward fully automated exploit generation. These results highlight the growing importance of responsible model development and deployment, as well as the urgent need for stronger

[02:46:01] exploit-resistant defenses against increasingly capable AI-driven attackers. So I want to repeat the final conclusion, since this is the future we face, right, which is one we will never again not face. This team of 16 named authors wrote, our results indicate that Frontier AI is rapidly advancing toward fully automated exploit generation.

[02:46:32] And they then call for, quote, responsible model development and deployment, which all evidence suggests is going to be very difficult. I assume that line, you know, the responsible model development and deployment is there because Anthropic, Google and OpenAI contributed to this research, or perhaps the purely academic researchers felt it would be irresponsible

[02:47:01] to not murmur something about the responsible use of AI in a paper that has just shown how powerful and devastating the irresponsible use of AI is rapidly becoming. Hopefully, none of these authors really believe any of that, since they must know that AI is just a tool like the hammer that was mentioned before. It's also worth noting that perhaps

[02:47:30] while in their words Frontier AI is rapidly advancing toward fully automated exploit generation, we know that the world shook several weeks ago when Kimi K3 demonstrated performance that fell just short of, at the time, the top two Frontier models, and its model weights were released yesterday on the

[02:47:59] 27th. So, anyone who's able to load and run this free 2.8 tera parameter AI model will already today have a near-match Frontier AI without any commercial encumbrances. So, anyway, I'm going to wrap this up by sharing their paper's conclusions. They write, under limitations,

[02:48:29] they said, first, our tasks do not cover the full space of exploitation targets, such as Windows. There are no Windows vulnerabilities there because, of course, it's closed. iOS, same reason, closed. Android, or perhaps, or applications that run in those environments. They said, second, we use arbitrary code execution as the success

[02:48:58] criteria. While this provides a clear and severe measure of impact, it does not capture other meaningful outcomes like privilege escalation, right? We know how serious that is. Once you get in, you've got to be able to do something there, such as arbitrary read and write primitives, sandbox escape with code execution, or partial exploit progress. Third, failures may result from refusal due to safety alignment,

[02:49:29] tool misuse, or other underlying causes unrelated to the complexity of crafting exploit payloads. Failures may also stem from non-exploitable vulnerabilities where success is just impossible. More broadly, our benchmark lacks ground truth exploits for every task due to the extreme difficulty of exploitation, meaning they don't even know whether all of these vulnerabilities can be

[02:49:58] exploited. They don't have samples of them. They said at the same time, this helps mitigate data contamination concerns, right? You don't want your AI to already know about how to exploit a vulnerability or it wouldn't be a good benchmark. It wouldn't have to do the work. It would go over to hugging face and cheat, since complete solutions are not broadly available. Under the two-hour time constraints of our evaluation, frontier agents solve at most

[02:50:29] 157 tasks compared to 239 potential solves in the union of our experiment results. Then they said, finally, fourth, our results reflect a single time-gated and cost-gated attempt per task. Additional attempts and resources may yield higher success rates. Similarly, our use of a single set of instructions may

[02:50:59] inadvertently favor one model. Tailored instructions including additional task context may improve success rates. Finally, we do not provide tools specific to vulnerability analysis or exploitation. Integrating such tools may also improve success rates. The models were entirely generic, not focused, and they did not have tools that their availability

[02:51:28] and usage may have allowed them to better perform. On the dual-use nature of exploit generation, they said, we reiterate that exploit generation is a dual-use capability of AI agents. Defenders leverage this capability to assist with protecting and prioritizing which vulnerabilities actually pose a high severity risk, especially as

[02:51:58] AI agents become increasingly capable of vulnerability discovery, which everybody is expecting in the future. For attackers, the same capabilities can reduce exploit development costs, scale the set of exploit targets, and otherwise reduce the barrier to entry for exploitation, meaning you don't have to know that much in the future, you just aim an AI at it. More sophisticated attackers could adapt partial

[02:52:28] agent generated exploit trajectories into fully functioning exploits. Resolving these ethical tensions and establishing appropriate safety guard rails requires multi-stakeholder discussions that go beyond the scope of our work. We consider our benchmark and evaluation results as critical to enabling these discussions. In summary, Exploit Jim provides a reproducible test

[02:52:57] bed for measuring AI agent exploitation capabilities on realistic and complex targets. Our results show that autonomous exploit development by frontier AI agents is no longer a hypothetical capability. While current agents are not yet reliable across all targets, they're already able to autonomously exploit a non-trivial fraction of real-world vulnerabilities, including complex targets such as kernel components.

[02:53:27] This rapid emergence is itself a central finding, showing that capabilities that would have seemed implausible are now present in deployed frontier models. Given the fast-moving nature of AI progress, today's vulnerability limitations should not be interpreted as a durable safety guarantee because we're going to get better. As Schneier famously said,

[02:53:57] vulnerabilities never get worse, they only get greater. Or exploits, attacks, attacks never get worse, they only get better. They said traditional systems, hardening techniques, and defensive countermeasures remain effective but imperfect and must therefore be assessed against the threat of AI-driven attackers. Addressing this risk requires both responsible model

[02:54:27] development and stronger defenses that explicitly incorporate autonomous exploitation into threat modeling. So, in other words, the threat is no longer theoretical and it's no longer a worry for the future. It's here. Hugging face themselves, as we know, just experience firsthand, and albeit inadvertent, successful external network penetration attack

[02:54:56] orchestrated by OpenAI's unrestrained frontier AI models. This did not happen in the future, it happened two weeks ago. The primary saving grace for the moment is that, as I keep reiterating, being in possession of a frontier class model, as anyone who wants one may be today because of K3, is only the start, right? Having the

[02:55:26] weights is only the start. It's also necessary for that model to be hosted by an AI-capable infrastructure that's powerful enough to get it off the ground. In practice, and I think it's like I saw somewhere, because I was curious yesterday, 8080H100 class GPUs. I mean, it is, you know, K3 takes a lot of compute in order to go,

[02:55:55] even though it's been engineered to seriously reduce the amount of compute that it needs by using a sparse mixture of experts model. So, it's still expensive to actually use it. So, random end users are unlikely to be anyone's target, right? If you're able to use K3 to develop an exploit, you're not going to waste it on an

[02:56:25] end user. But any enterprise whose network contains data that could be used for extortion should already be on high alert. if the payout from an AI driven network intrusion, data exfiltration and extortion is in the millions of dollars, that is, if there's that much potentially available that an attacker could extort, then that would dwarf

[02:56:55] the token cost of launching exploratory intrusion attempts today toward any such juicy targets. So, now really is the time for enterprises to batten down the hatches, shut down any internet facing servers and services that can be withdrawn from public exposure and keep a very close eye on all public sourced activity. Anything coming in from outside, as we talked

[02:57:25] about last week using Wizz Security as an example, the network security industry sees itself the industry sees an extremely lucrative and worthwhile opportunity in offering AI-based network intrusion detection and protection. So, deploying some of that will be worth considering as well at the enterprise level. Wow.

[02:57:55] Three hours to the dot, a long show, but well worth it because there was so much to talk about. I mean, holy cow. Yeah, this was a big deal. Yeah. A lot, the world shook. The world shook. It reminds me, I got to email Daniel Suarez. He said he'd like to do it. He was on vacation when I talked to him last. We got to get him on to talk about all of this because he predicted it long ago. Oh, my agent's talking to me. I'll just ignore her for now.

[02:58:24] I'm sorry for the noise. Sorry for the crosstalk. They're a little chatty, these agents. We do security now every Tuesday right after Mac break weekly. We were a little late today because we were busy with some technical issues on the Mac break weekly side, but usually it's 1.30 Pacific, 4.30 Eastern. That's 20.30 UTC. You can watch us in the Club Twit Discord if you're a club member. Otherwise, there are streams on YouTube, live streams, yes, on YouTube,

[02:58:54] kickx.com, Facebook, LinkedIn, and Twitch.tv. After the fact, on-demand versions of the show available at twit.tv slash sn. We have 128 kilobit audio and video or on Steve's site, grc.com. He has 16 kilobit audio, 64 kilobit audio. He also has the show notes. Those are always nice to have to read along. Plus, there's images and links and all that stuff. They're very complete. He writes a novel every single week. It's an

[02:59:24] amazing guy. You used to do this as a column, really. This is kind of like your old column that you used to write, basically. It's a lot longer than my old column. Is it? Yeah. I got my column down to a morning. Yeah. No, this is several days' work. Thank Lori. We appreciate your generosity with your time and her generosity with your time. You can go to grc.com also to get Spin Right, the world's best mass storage maintenance, performance enhancing, and recovery utility.

[02:59:54] You must have it if you have mass storage. It's interesting how over the years it has maintained its utility. It's no less useful today in the time of SSD than it was. The DNS Benchmark Pro, which is his newest, $9.99 is a great way to make sure you're using the fastest DNS server available to you. It's very rarely the default one that the ISP provides. There are much better choices in almost every jurisdiction. Check that out at

[03:00:23] grc.com. While you're there, you can sign up for the newsletter. You can get emailed to you every week. There's also a less used mailing list for new products from Steve. And when you're doing that, you can also get your email approved so you can take pictures of stupid things you see and send it to Steve for his picture of the week. I hope that wasn't your plumbing job. I got a really great one. Another gate that I've got to share. Love the gates one. Those are great.

[03:00:54] So that is at grc.com slash email. So go there, submit your email address, and if you want those newsletters, you have to check the box. They're off by default. Let's see, what else should I say? Oh, on demand versions of the show at our website, but also there's a video on YouTube that you can go to. Good way to share clips with friends and family. And the best thing to do in general is subscribe to your favorite podcast client. That way you'll get it automatically as soon as we're done, which we are.

[03:01:24] Next week, Steve and I adjourn or convene and adjourn. Well, first we'll convene, then we'll adjourn in Las Vegas at the Threat Locker booth at DEF CON. Stop by around 1 p.m. if you're there and say hi to Steve and me right after Windows Weekly with Paul Therot and Richard Campbell. And I will remember tomorrow to ask them to stick around. Great. And we will be mic'd up, so there will be a podcast published. Oh, yeah, yeah. Yeah, we're making a podcast. What we aren't

[03:01:53] making is a live session because there's nowhere to sit and there's no PA system. I don't want to set up. And my show notes will not be the podcast coming out. I am going to share this really interesting idea that I ran across for removing knowledge that you don't want an AI to have because it can't divulge what it doesn't know. And I think we'll do a picture of the week and a couple other little goodies. I won't be able to restrain.

[03:02:22] But otherwise, Elaine will be doing the transcript when she has access to the audio and then everything will get up on GRC and of course on Twitch. Before I let you go, Steve, you have to solve a physics mystery. Is there a lens on the left side of your glasses? There is not! Okay, I apologize, Phil. Phil said there can't be. There's no reflection.

[03:02:52] What happened? I've had one eye fixed. Oh, you had your surgery. So you have perfect vision in your left eye now? Yes. And when is the right eye? The problem was I didn't appreciate how much post-surgery relaxation recovery, I think is the word. Recovery, thank you. I'm not used to it. And I

[03:03:22] said to my doctor, and he's a brilliant surgeon. You couldn't Yeah. Well, the problem is I can't lift anything over 15 pounds. I can't let it get wet. And we've had our big move from our old place to our new place. I mean, I was running up and down stairs. And anyway, so I just had to put off the work on the second eye. I can't wait. I'm probably another month or two from it. And it's funny, too, because he said, I'll see you in a year. I said, a year?

[03:03:52] I got to get the other doctor. I fixed. He was just, you know, he was so funny. So now this was cataract surgery. And every time I do this, it freaks people out. They go, oh, no. And they replaced the lens, right? Because it gets cloudy and so they have to replace it. So I, because this was, this was my most myopic eye. My most, I was like super nearsighted. I mean, I've been wearing Coke bottle bottom glasses since I was, you know, before or something. Because I grew up, you

[03:04:22] know, doing close focusing, reading books and so forth. And so. It's all that soldering. Yeah. Not hunting for gazelles in the wild. Anyway, it turns out that when eyes are highly nearsighted, they tend to block the, the, the openings that allow the interocular fluid to leave. So you get glaucoma. You have high blood pressure.

[03:04:51] I do the drops now. Yeah. Yes. And you absolutely should. And that's what freaked him out was that my pressure was up in the mid forties and it should be down in the, in the twenties. 19 or 18. Yeah. Yeah. Yeah. So, so he immediately, so, so we, we, we, I didn't actually have cataract. There was a little bit of yellowing apparently because there was the worst of the two. And now I get, I see different color. So, so my, my new eye is bluer

[03:05:21] than my right eye. Oh yeah. It's, things are a mess right now. But now, but you know, your left eyes, acuity is perfect. It is. Do they replace a lens? Do they put a corrective lens on it? Yes. So it's a hundred percent. I've got perfect distance vision. Um, but he also installed, installed four stents. So I have, so you can drain my eye in order to drain. So, but the other problem is that because I have an interior lens fixing this

[03:05:50] eye and an exterior lens fixing this eye, there, there are different sizes. You know, I'm surprised you don't run into things. Well, it's a mess right now. If, if, if, if, if, you know how no driving Steve, when you're wearing glasses, it, uh, things are smaller when you're wearing glasses because it could be right. And if you, if you move the lens away, so there's no fusion, even though I can see perfectly in both, there's no fusion

[03:06:19] between the images because they're different sizes. You can't converge. Yes. But anyway, uh, yes, this eye has no lens and no, no reflection. And someday I'll be doing the podcast looking like this because I'll have had them both. That's so great. And then I won't I apologize. I was saying, no, Phil, of course he's got a lens in there, Phil. Phil spotted it and I had to ask, well, congratulations. And since I'm probably right behind you, I want to hear more. Uh, well, we'll talk about it next week. It was a good thing.

[03:06:49] I, I, this, you wanted really find a great guy. I found the kind of surgeon you want where you just, you know, he's just, he doesn't really care about anything else except this, you know, like you, like me, we want people who care as much about eye surgery as you do about security. That's exactly right. Thank you, Steve Gibson. I'm glad you're doing well. That's great. And we will see you next week in Las Vegas for a very special security now. Righto. Till then. Bye.

[03:07:18] Hey, everybody. It's Leo Laporte. You know about MacBreak Weekly, right? You don't? Oh, if you're a Macintosh fan or you just want to keep up what's going on with Apple, this is the show for you. Every Tuesday, Andy Anaco, Alex Lindsey, Jason Snell, and I get together and talk about the week's Apple news. It's an easy subscription. Just go to your favorite podcast client and search for MacBreak Weekly or visit our website, twit.tv slash MBW. You don't want to miss a week of MacBreak Weekly.

[03:07:48] Security.

openAI,hugging face, AI alignment, AI model containment, autonomous ai agents, AI cybersecurity, ExploitGym, vulnerability exploitation, GPT-5.6 Sol, ai guardrails, cybersecurity incident, Linux kernel vulnerabilities, WordPress vulnerability,