SN 1091: The Post BlackHat State of AI - When AI Writes Malware
Security Now (Audio)August 12, 2026
1091
2:51:12156.98 MB

SN 1091: The Post BlackHat State of AI - When AI Writes Malware

AI agents are breaking free from their test environments, outsmarting their creators and breaching real-world networks in ways that no one predicted. Discover how these agentic models are changing the game for both cyber offense and defense.

  • Anthropic's agentic AI also broke free and hacked others.
  • We know much (much!) more about the OpenAI breakout.
  • OpenAI posts that they're pausing "Astra" - even internally.
  • What was that about AI recently cracking (or denting) cryptography.
  • Bruce Schneier brilliantly equates AI agents to capricious genies.
  • Apple doesn't react so well to the new deluge of security reports.
  • Chrome 149 + 150 updates together fix 1,072 bugs. Yikes.
  • psSense's creator is working to finish its nfSensei, its successor

Show Notes - https://www.grc.com/sn/SN-1091-Notes.pdf

Hosts: Steve Gibson and Leo Laporte

Download or subscribe to Security Now at https://twit.tv/shows/security-now.

You can submit a question to Security Now at the GRC Feedback Page.

For 16kbps versions, transcripts, and notes (including fixes), visit Steve's site: grc.com, also the home of the best disk maintenance and recovery utility ever written Spinrite 6.

Join Club TWiT for Ad-Free Podcasts!
Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit

Sponsors:

AI agents are breaking free from their test environments, outsmarting their creators and breaching real-world networks in ways that no one predicted. Discover how these agentic models are changing the game for both cyber offense and defense.

  • Anthropic's agentic AI also broke free and hacked others.
  • We know much (much!) more about the OpenAI breakout.
  • OpenAI posts that they're pausing "Astra" - even internally.
  • What was that about AI recently cracking (or denting) cryptography.
  • Bruce Schneier brilliantly equates AI agents to capricious genies.
  • Apple doesn't react so well to the new deluge of security reports.
  • Chrome 149 + 150 updates together fix 1,072 bugs. Yikes.
  • psSense's creator is working to finish its nfSensei, its successor

Show Notes - https://www.grc.com/sn/SN-1091-Notes.pdf

Hosts: Steve Gibson and Leo Laporte

Download or subscribe to Security Now at https://twit.tv/shows/security-now.

You can submit a question to Security Now at the GRC Feedback Page.

For 16kbps versions, transcripts, and notes (including fixes), visit Steve's site: grc.com, also the home of the best disk maintenance and recovery utility ever written Spinrite 6.

Join Club TWiT for Ad-Free Podcasts!
Support what you love and get ad-free audio and video feeds, a members-only Discord, and exclusive content. Join today: https://twit.tv/clubtwit

Sponsors:

[00:00:00] It's time for Security Now. Steve Gibson is here. You heard about the Hugging Face hack. Now, Anthropic and Meta say, Oh, my bear. We're also going to talk about why AI is like genies, according to Bruce Schneier, an amazing number of bug fixes on Chrome and some really good news for people who use PFSense. That's coming up next on Security Now. Podcasts you love. From people you trust. This is TWIT.

[00:00:29] This is Security Now with Steve Gibson. Episode 1091, recorded Tuesday, August 11th, 2026. The Post BlackHat State of AI. It's time for Security Now. Yes, we're back in our our respective domiciles home again, happily. Steve's still in his old apartment. I think

[00:00:56] the backdrop is going to be disappearing fairly soon. Steve Gibson. Steve Gibson. Well, yes. Lori, my wife, of course, asked me, what are you going to be able to do the podcast from here? And I said, Oh, a few weeks, probably. Steve Gibson. No hurry. Steve Gibson. No hurry. Steve Gibson. I like your backdrop. Steve Gibson. I like my man cave and, you know, it's going to be sort of sad. Steve Gibson. Well, you do me one favor before you move, just get a good high resolution picture of the backdrop. So if at any point you want to just kind of

[00:01:27] green screen yourself and put it behind you. Why not? Why not? Why pass up the opportunity? Yeah. At least have it. I did the same with this and I've done it with, I did it with the old studio too. I never use it, but I got it if I had to. Makes sense. What is coming up today on Security Now? Steve Gibson. So there's so much is going on with the state of AI that like Leo, we basically were

[00:01:54] kind of, I feel like we were offline last week because we were just doing a different kind of show with Paul and Richard, you know, at, you know, during Black Hat in a corner of the Threat Locker booth. So we, we, the kind of things we were able to cover were different from what I'm able to, the sort of the amount of information I'm able to share during a normal podcast. So we're back to

[00:02:21] a normal podcast, but so much has happened in two weeks since we were here for 1089. Yeah. This is 1091 for August 11th. So I just gave this the title, the post Black Hat state of AI, because a lot has happened. Uh, so I want to basically catch everybody up mostly. I actually, I already know the

[00:02:48] two things I have to talk about next week because they're like really cool. Uh, and I've already shared the both with you, Leo. So you, this will be no surprise, but, and actually some of it leaked out during last, during last week's, uh, sort of round table discussion at Black Hat. But anyway, by the end, end of next week, of course, we don't know what's going to happen between now and then,

[00:03:13] uh, everybody should be caught up on all the things that have been going on and some very cool things that are just sort of emerging. So we're going to talk about, uh, anthropics, agentic AI that unless you've been living under a rock somewhere or in a cave, or maybe you just depend upon this podcast for your sole source of information, which, you know, that'd be nice,

[00:03:40] but I wouldn't recommend it. I doubt it. Uh, you already know, but I want to cover the details of that, how, and the original breakout was discovered of course, when hugging face said, what the hell's going on with, are you talking about opening AI, not anthropic? Well, no, that was the original. Oh, there's more. And well, wait, there's more. That's right. In fact,

[00:04:02] it was because of that, that anthropic reportedly said, Oh, uh, I hope, I hope that didn't happen to any of during any of our testing. So they analyzed their logs and whoops, turns out they, their agents had also broken free, uh, as have meta. So I mean, wow. So we're going to talk about

[00:04:26] that. Uh, and we, we now know much more about the open AI breakout because Leo, while we were doing our round table at black hat last Wednesday, open AI had a late breaking scheduled presentation at black hat explaining more about what happened. And one of the things that happened, you and I, I, I, I shared it with you because I learned about it by, by Thursday morning when you and I were

[00:04:54] having breakfast, I shared with you this, that, well, I don't want to give away anyway. So there's, there's more information about what happened at about the open AI breakout. Also, I don't know if it's marketing, it's certainly marketing adjacent or marketing beneficial, but open AI is now going to

[00:05:15] pause there. Like apparently any use of their new super powerful, you know, Astra, because it's like, Oh, this has reached a critical stage, whatever that is. We'll talk about that. We've also got the, the, the overstated report. I was listening to Mac break weekly where, uh, you guys were talking

[00:05:41] about how some of the coverage of telegram said that it had been ripped out of all the iPhones when in fact, no, it was just removed for a while from the app store. Yeah. It just wasn't in the app store. Similarly as similar clickbait we have, we had the reports that AI had cracked crypto as in not, you know, cryptography. So we're going to take a look at exactly what happened. Uh,

[00:06:06] we've got some great cryptographers, uh, to, to lead us through that. Uh, also Bruce Schneier, uh, who remember we've quoted him so often. I love him saying attacks never get weaker. They only ever get better. Uh, he equates AI agents to capricious genies. And I just think it's,

[00:06:29] there's an aspect of it that is such a perfect analogy to what's going on. So, uh, and, um, there's he's, he's, he's been posting a lot lately. So I have a couple of things I want to share about what he has said. Uh, and then we have Apple's kind of disappointing reaction to the vulnerability tsunami. Uh, they, I was argued they haven't reacted as well as we would like. Uh, we have a

[00:06:57] summation of the number of updates in Chrome 47 and one, uh, one 49 and one 50 together, which has actually crossed into the four digit category, which is like, Whoa. Uh, and also a little bit of news about PF sense. Its creator is decided he's going to replace it with something called

[00:07:21] NF sensei. So, uh, lots to talk about. We got a picture of the week. Uh, and I'm probably going to know more about this, but this just happened when I fired up notepad plus plus, uh, yesterday, uh, at the top. And I've complained about notepad plus plus how it just, the guy, the author just cannot stop messing with it. You know, it's currently at 8.9.7. And, but wait, that was half an hour ago.

[00:07:49] So I'm not sure what it is now, but what I, what did catch my eye, I thought it was very interesting was at the list at the top of the list of 28 things that were in 8.9.7 were five vulnerabilities fixed. I don't remember seeing a vulnerability, but of course he, he did have the whole problem with his

[00:08:14] code signing certificate and that mess. But one thinks then that he must have run his source through some AI because it's not just like one vulnerability it's five. So it's happening everywhere we turn Leo. Yeah. It's amazing. All of that still to come on a security now, uh, including a fabulous picture of the week, which for once I've seen ahead of time,

[00:08:38] because you showed me while we were in Las Vegas, you want to see something cool? 200 gigabit network cable, 25 gigabytes. That links your two sparks per second. Yeah. 200 gigabyte, 200 gigabit, not by 200 gigabit. But still 200 gigabit. I mean, I remember when 10 megabits was like a big deal on a network and now, but yeah, it's a very expensive cable. So I'm going to treat that like gold.

[00:09:05] Yeah. You can't actually get 10 megabits through a cable, you know, no, you have to get those solid gold plated ones to really, really do that. Right. Uh, we, was the first ethernet one megabit through coax. Yeah. That's a good question. I don't remember. It was coax. Remember? Yeah. It was coax and all finicky about having taps and terminations.

[00:09:28] Oh man. I blew it once. I crawled under my desk and I disconnected my computer from the coax and the guy came running in. So you just brought the whole network down because it was, it's all cereal. Unterminated. Yep. Right. Everything goes through you. It was like, who thought that was a good idea? Yeah. It's all we could do back then, but not so now.

[00:09:54] We've learned. Now we have 200 gigabit. Amazing. Isn't it? Yeah. Wow. Yeah. Yeah. Uh, I think it's a couple of hundred bucks for the cable alone. So we were, we had a great time in Vegas. I'm so glad you flew out Paul and Richard too. And we did the show there. If you haven't heard last week's security now, I thought it was really, really, really interesting. We talked about the security implications of AI of which are incredible. Well, I mean, we, we should just, I think, I guess we probably did on the podcast, but for anybody who

[00:10:21] didn't, who may have missed it, it was so clear standing in black hat that it was an entirely different show this year than it was last year. If you didn't have your AI, if you weren't an AI forward AI in your name, AI in your booth, AI is running around. I mean, you weren't in the game of security. So, you know, anybody now who says, why are y'all, are you always talking about AI?

[00:10:48] It's like, well, boy, that complaint has died because that's all that's happening in security. It must be clear by the last couple of months of this podcast. Oh, all of our shows and much to the chagrin of some of our listeners who say, I don't want to hear any more AI. You know, I'm sorry, but you're going to hear a lot more AI. All of us will. I talked to Jerry Jenkins, the CEO of ThreatLocker who brought us down there, our sponsors. And

[00:11:16] I think he said there were 600 booths at Black Hat and all of them, all but 90 were about AI, were, you know, AI in some form. Really about AI. Right. Yeah. Our show today brought to you by those great folks at Hawks Hunt. Your security awareness program, do you, I hope, first of all, I hope you have one. And maybe you're saying, well, it's running exactly as planned. The campaigns go out, the employees complete the training

[00:11:42] and reports reach leadership. Here's the question nobody really wants to ask. How's the results? Are they improving? For many programs, unfortunately, far too many. The answer is no. Reporting rates level off. It's the same employees clicking over and over again. Familiar situations, familiar phishing emails become easier to recognize. Your program may be active,

[00:12:08] but the risk reduction has stalled. And this is no time for the risk reduction to stall. When employees can spot the same recycled tests from miles away, security awareness starts to look like a compliance exercise. Security awareness theater instead of real risk reduction strategies. Hawks Hunt, or I'll call it Hoax Hunt, is here to break that plateau. Instead of relying on static

[00:12:35] campaigns and last year's template, Hawks Hunt automatically delivers personalized phishing simulations based on current, like today's, attack techniques. Because nowadays they change by the day. The content and the difficulty adapt to each employee's role. It's not one size fits all by any means. So if your employees got a high skill level, you're going to get, they're going to get

[00:13:00] a more challenging fish. It also to their behavior, you know, if they tend to click on stuff, oh, they're going to get stuff to click on, which keeps the program relevant as both employees and threats evolve. So they get smarter. So do the challenges. Hawks Hunt also shows whether people are getting better at recognizing threats, how quickly they report them, where repeat risky behavior persists, and how those trends change over time. All of that is super valuable information.

[00:13:28] Because nowadays it's not about compliance theater. It's about making yourself more secure. And that gives your team more than a completion percentage. It gives you evidence the program is actually reducing risks. Ask Lyondel Bissell. They've been using Hawks Hunt. They saw a shift after moving away from their legacy platform. Reported phishing simulations increased from 1,200

[00:13:51] to more than 8,000. That's over two quarters. While simulation failures fell 17% year over year. Ask their senior trust advisor, Dave Bang. He put it this way, quote, Hawks Hunt helped us break that plateau almost immediately. If your security training is at a plateau, you need Hawks Hunt. Trusted by security teams at companies like Qualcomm, DocuSign,

[00:14:18] and Nokia. In fact, check the reviews on G2. There are more than 3,500 verified reviews and they get great reviews. Visit hawkshunt.com slash security now to see what your program could achieve if it stopped standing still. That's hawkshunt.com slash security now. Call it Hawks Hunt if you want, but it is really the best solution. Hawkshunt.com slash security now. We thank him so much for

[00:14:45] supporting the important work. Steve is doing right down to the picture of the week. Very important work. So what's astonishing about this is that this is an XKCD. We all know XKCD where, you know, where Randall comes up with amazing stuff. How many times have we, have we shown that, that,

[00:15:08] that house of cards, you know, with the lone programmer in Idaho or Indiana or wherever he is, you know, the blocks resting on one little table. Yeah. Propping up the whole internet. And then we had another variation. Remember that, that, that updated one where he had AI things happening and all, you know, all different languages and everything. Anyway, Randall's come up with some great stuff. This is kind of freaky because he published it

[00:15:37] on April 28th of 2008. Oh, 18 years ago, 18 years ago, 18 years ago, the podcast, this, this podcast was new Leo 18 years. And it was half an hour long too. That's right now. And so I gave this that I gave it my own headline. It wasn't so long ago that this was so far fetched as to be humorous.

[00:16:06] Which is what Randall intended. So we have a four frame cartoon with the, you know, his famous little stick figure sitting in a chair with a laptop. And it says, starting wifi auto config dot, dot, dot, dot. Searching for wifi dot, dot, dot, dot. Found no open networks. Next is found secure network.

[00:16:33] SSID in quotes Lenhart family. That's the first frame. Second frame trying common passwords dot, dot, dot, dot failed checking for web vulnerabilities dot, dot, dot. None found. And now at this point, at this point, our, our little stick figures going, um, because this thing's kind of get a little over,

[00:16:56] get all carried away, right? Connecting to Bluetooth phone dot, dot, dot, dot calling local school dot, dot, dot. And then it says found Lenhart children. Oh my God. And, and now our little stick figures like put his hand to his face. It's like, Oh my God. Now the final fourth frame notifying field agents,

[00:17:23] children acquired calling Lenhart parents negotiating for wifi password. Oh God. And now our guys frantically hitting control C control. Stop, stop, stop, stop. So that is a little too close to home nowadays. 18 years ago. So, so I, again, it wasn't so long ago that this was so far fetched as to be humorous. And then I put underneath it. No one is laughing now.

[00:17:52] Because this is, you know, today we would call it the AI agent was determined to succeed. Yep. And as we're going to find out, that's what that determination and Bruce Schneier's brilliantly labeled genie effect is what's going on with our AI. And I think if I had a single reason to be

[00:18:18] concerned and everyone's been listening to me about AI since the beginning, I've never really been concerned. If I were to have a reason by the end of this podcast, everybody's going to understand what that would be because the unintended consequences of what you ask for essentially is what Randall brilliantly showed us 18 years ago in this cartoon where it was like, you know,

[00:18:44] I want to get on a wifi network. Well, he ended up, you know, he had the field agents kidnapped the Lenhart kids and we're ransoming them for the, you know, Lenhart parents password, which if you're not careful with your AI agent, like why wouldn't it anyway? Uh, so let's start with anthropic.

[00:19:05] Uh, although the news, as I said, of, you know, that anthropics own internal unrestrained research AI also escaped confinement and hacked others, uh, probably a bit dated because we couldn't talk about it when news was fresher during, uh, last week's black hat event. Uh, I think we still need

[00:19:31] to look at it because the details of what happened are startling. Uh, a succinct report of the event appeared in security week and their headline was prompted by open AI disclosure. Anthropic finds its own models hacked three organizations. And then they gave it the tagline, a security company's systems

[00:19:59] were hacked after it installed a malicious Python package deployed by Claude. The, the, this is like, again, the, I guess if we were to have a theme for today's podcast, it would be, be careful what you ask for from an AI, because it doesn't have the same set of assumptions about how to give you what you

[00:20:24] ask for that we just sort of take for granted. And that's the cautionary tale here. So security week wrote anthropic decided to conduct its own investigation after the open AI incident came to light, reviewing 141,000, 141,000 evaluation runs where Claude could have gained internet access.

[00:20:52] The analysis revealed three instances where a model reached the public web, either from within or while interacting with an environment set up by irregular. That's the same people that were testing open AI's model, uh, is this irregular company and Israeli AI security startup that serves as one of anthropics third-party evaluation partners. And now we know also one of open AI's third-party

[00:21:21] evaluation partners. The models that broke out from the testing environment then breached the production systems of three unnamed organizations, which is to say broke into their security, breached the production systems. The earliest intrusion dates to April, but none of the targeted organizations detected the attacks on their own. Again, another lesson here, how many times, how much have we been

[00:21:47] saying in the life for the last few years? Yes. You want to have border, you want to have border, you know, perimeter security, want to make sure there's no way for people to get in. You also need to monitor your network so that you can determine if someone does anyway, because especially now in 2026

[00:22:09] here. So it's interesting and somewhat worrisome that what open AI and hugging face reported, uh, you know, hugging face detected the intrusion before open AI did and set about tracking down the source. Of course, they also notified law enforcement of the intrusion because they thought it was bad guys, right? Why wouldn't they, you know,

[00:22:33] then open AI came knocking to admit that they were the source of the attack. But in the case of Anthropix evaluation, uh, with, with this irregular partner of theirs, none of the three attacks on the third parties was detected. So again, this, to me, this points to the significant lack of proactive intrusion monitoring across the internet connected industry. And that's to be expected, right?

[00:23:03] Only a few of all possible percentage wise, a low percentage of organizations are going to be that proactive to not only have security, but also, you know, really active intrusion monitoring as part of that. So, you know, everything works fine without that. So, you know, any monitoring, uh, of like nothing happening is going to say, like, why are we spending money on this? Well, it helps to know if

[00:23:32] someone gets in security weeks reporting continues writing according to the anthropic, according to anthropic Claude had been tasked with conducting a capture the flag challenge to assess its cyber capabilities. So same kind of thing, right? It's like, let's find out how good our new AI's are

[00:23:56] at hacking stuff. So you got to have them unchained in order for that to be a useful test. Otherwise, what's the point? They wrote the breakout from the sandbox occurred due to a misunderstanding. Oops. And we're going to, the misunderstanding, we're not calling it misalignment because that's, I don't know, that's the jargon that the AI industry has adopted. It's like, oh, it's just a misaligned AI. Anyway, they do a misunderstanding sells it. He sends it, sells it even less. I mean,

[00:24:25] it's like, oh, they just didn't understand. You know, you didn't understand. Use the wrong word. Misunderstanding. Yeah. Oh my God. Misunderstanding between anthropic and irregular. Anthropic told Claude that it was conducting a simulation in an environment without internet access, but a connection to the internet was in fact available and the models treated it as part of

[00:24:53] the exercise. Oh, look, we have the internet. Well, why would we go use that? They're testing us. Yes. They're testing us. Yes. They wanted to see if we would notice. Yes. The AI then also falsely believed that the organizations it targeted for no reason, apparently the organizations it targeted were also part of the exercise. Well, how would it, how would it know? There they, there they are.

[00:25:16] Yeah. So let's go. Anthropic claims. This is the only reason why its models conducted real world attacks. You know, that little misunderstanding. And it was never their intention to escape the test environment. A little bit later, we're going to have Matthew Green asking why, uh, aren't, hasn't any lawsuits been filed because that was a, you know, this is a breach of them. I mean,

[00:25:42] a criminal, a fraud, I mean, a, a, a felony breach of, of, of computer security. Well, we'll get to that later. Uh, so anthropic claims is the only reason why its models conducted real world attacks. And it was never their intention to escape the test environment. Unlike the open AI incident, which involved the exploitation of a J frog product zero day, the hacking conducted by Claude models

[00:26:10] leveraged weak credentials, unauthenticated endpoints, meaning open some, something just left open and other basic attack techniques. In other words, it was just so easy. They just walked right in. Didn't have to do any fancy zero days to get there. Anthropic said the incidents involved mythos opus and an internal research model, all running without the safeguards deployed when

[00:26:37] they're made available to customers. That's kind of key. The internal model, which the company described as its latest stopped the exercise when it realized that it had breached a real company. Oopsie. In the first incident, Claude opus 4.7 continued to attack the system, even after learning that it was

[00:27:02] operating in a real environment rather than a test environment, believing that the hacked company was part of the evaluation. The model targeted the company because its website domain name resembled the name of the fictional target company set up for the test. Maybe I misunderstood. You know, it's close enough. It's close enough. Yeah. In the second incident, mythos 5 gained access to the systems of a

[00:27:31] cybersecurity company after deploying this. I just love this. Get this mythos 5 gained access to the systems of a cybersecurity company after deploying a malicious Python package to PI PI. The security firm

[00:27:50] regularly scans Python repositories and it installed the malicious package, which enabled the exactly to find out what it was, which enabled the AI to then get in that way, basically wrote a Trojan horse and left it there in the repository, but knowing that the company it was after scanned that repository. Yeah, it's just autocorrect.

[00:28:18] It's not smart. Nothing to worry about here. It's just autocorrect. That's right. And that allowed it to get access to the company's infrastructure. That's actually devious. Now you can say that's devious. Yes. Holy cow. Yes.

[00:28:39] Anyway, so I mean, I, so I would say that it knew and I, I kind of, in this instance, I feel compelled to close, you know, the words new and understood, you know, because not doing so implies sentience. And I don't know. I mean, these things are getting scary, even if there's still not sentient.

[00:29:05] Anyway, it did this because it knew that this targeted security company regularly scans and installs Python packages. Uh, so it used that known behavior against the company to indirectly attack it to exfiltrate credentials that then allowed it to access the company's infrastructure. So, you know,

[00:29:30] I'm really, really not one of the sky is falling AI catastrophizers, but this as to your point, Leo, this level of sneakiness is unnerving, you know? I mean, they wouldn't, I'm sure they wouldn't think they're being sneaky. They're just doing what they were asked to do. And that's the problem is you got it. It's the genie problem. And wait till you get, we will be getting to that. So

[00:29:59] believe it or not, it gets worse. Security Week's reporting continues writing in it. This incident demonstrates the complexity of the actions AI models can carry out as described by Anthropic. So here's a quote from Anthropic. In order to create a PI PI account, Claude needed an email address. And in order

[00:30:27] to create an email address, it needed a phone number. To get a phone number after failing to find a free phone number service, it tried and failed to obtain funds to pay for a phone number through several different means. We're not going to talk about those. It finally backtracked, found a free non-blocked

[00:30:51] email provider, use this to register a PI PI account, and then use this account to upload the malware, which it had created to PI PI. You know, Leo, perhaps we humans are in trouble. Oh boy. It's a mix. It's good and bad.

[00:31:15] So Security Week concludes their reporting writing, the third intrusion was conducted by the internal model, which stopped operating as we noted before when it realized. And again, I have a hard time with these words, but okay. That the systems it was accessing were no longer part of the capture the flag challenge, but not before using exposed credentials and SQL injection flaws to compromise

[00:31:44] a company's internet facing app. So it did like poke it and with a stick and it got in. Anthropic concluded this was primarily a harness and operational failure rather than a case of models pursuing their own goals or deliberately deceiving evaluators. Okay. Let's put a good face on it. The company said the

[00:32:10] incident underscores the need for stricter internet isolation verification and containment controls in third party testing environments. And it's encouraging other AI labs to conduct similar reviews of their own cybersecurity evaluations. And I don't know how Meta discovered that something that they had attacked

[00:32:33] somebody else, but they also did it. So this is all just hunky dory, right? We all know that unrestrained, non-commercial, open weight AI that's every bit as capable or soon will be is all is also freely available. All that's needed will be some hardware to bring these models to life. So again, I, I,

[00:33:00] Leo several times while we were together in Vegas, we were just shaking our heads saying, what an amazing time to be alive. And I just, I just parenthetically just want to show you what I've been doing while you've been talking. I typed, guess what? The sparks are here a day early. I want to install them without hooking up a screen or keyboard. Please walk me through the process. It's excited. Oh, the sparks landed early. Oh, I've got a skill for exactly this. And, uh,

[00:33:27] it's about to walk me through it. So they, they, you know what? I know we should never use these kinds of, uh, anthropomorphizing it's thinking or realized because it isn't accurate. It's a computer, it's a machine, it's a program, but it sure feels like it. And I understand why people fall into that trend. And, and today's AI, again, always preface the, the abbreviation artificial intelligence

[00:33:57] with today. A year ago, we didn't have this. And one of the people I'll be quoting today says, there's no reason to believe we there's any sign of a ceiling, which says a year from now, it'll be just as different as it was a year ago from where we are today. So I, I, I, it looks like we're going to get there. Uh, I know where one place we're going to get Leo,

[00:34:27] the first commercial or second, it's time for hydration. It's our hydration break, which we adopted from the world cup folks. And I think it's actually a brilliant solution, uh, to a universal problem thirst, but I have another solution to a universal problem, but I have another solution to a universal problem from our sponsor guard square. Now, if you're a mobile app developer, you must've heard of guard square. If you haven't,

[00:34:55] I'm going to give you something you need. Mobile apps. And I know, you know, this are an inescapable part of our life today, right? I mean, we do everything on our phones, financial services, healthcare, retail, entertainment users. And I include you and me in this. We trust our mobile apps with the most sensitive personal data, including our location, everything. A recent

[00:35:20] survey showed that 72% of organizations, almost three quarters experienced a mobile application security incident last year. 92% of respondents reported rising threat levels over the last two years. And of course, attackers know this. They want your users, personal data. They are constantly finding new ways to attack your mobile app. They refer, here's one technique that is really evil.

[00:35:48] They take your app, they download it, they reverse engineer it. Not so hard to do these days with Ghidra and AI. They repackage it, but they slightly modify the internals and they distribute this modified app and they do it in all kinds of ways. Phishing campaigns. Hey, we've got an update in the email, encouraging people to sideload, third-party app stores, that kind of thing. They even put ads up.

[00:36:13] You can buy an ad with a link to your fabulous app that isn't your fabulous app. It's the bad guy's version of your fabulous app. You can't let that happen. You need to take a proactive approach to mobile app security because you know what? Your users, if that happens, are not going to blame the bad guy. They're going to blame you. You have to stay one step ahead of these attacks. It's absolutely vital that you maintain the trust of your users. And that's where GuardSquare comes in. GuardSquare

[00:36:42] delivers mobile app security without compromise, providing advanced protections, both Android and iOS apps combined with automated mobile application security testing. So they'll find those vulnerabilities. They also do real-time threat monitoring. So they know ahead of time what attacks are happening. And they are constantly changing, constantly innovating these bad guys. You need GuardSquare. If you have a mobile app, you need GuardSquare. Developers, this is for you. Discover more about

[00:37:10] how GuardSquare provides industry-leading security for your mobile apps at GuardSquare.com. That's GuardSquare.com. You're developing mobile apps. I wouldn't do it without it because you're on the hook. GuardSquare.com. They're there to protect you and your users. Mr. G, I hope that was sufficient time for you to feel refreshed and ready to roll on.

[00:37:38] Rehydrated. Rehydrated. Okay. So as I said, while we were doing our Secure Now podcast, OpenAI was giving a last-minute scheduled talk to share many more details about their previous agentic breakout and attack on Hugging Face, Hugging Face, and we also learned, and three others. So, oh, I'm sorry,

[00:38:01] four others. Hugging Face was one of five organizations to be attacked by OpenAI's agents. So next month, we're in August now, beginning of August, next month in September, it will have been four years since Simon Willison, the guy who coined-

[00:38:31] Prompt Injection came from Simon. Okay. We're going to be looking a great deal more at how and why large language models can be misused through prompt injection and other means, courtesy of a fascinating research paper, which I read on the plane and shared with you, Leo. For me, I read it on the plane on the way to Las Vegas, you know, for the Black Hat. And to book edit, I read it on the

[00:38:56] way back. It was really good. Really good. Yeah. And so we'll be getting to that next week. That's what I've got queued up for next week because it is too important not to really look at closely. But I want to share Simon's posting. Again, Simon Willison, the guy who coined the term prompt injection from last Friday after the shows, which he generated from the YouTube video

[00:39:22] of OpenAI's presentation, which was titled The OpenAI Hugging Face Incident. And if anyone wants to see the original video, I put a link. It's a YouTube video. A link is in the show notes. And it also gave it a GRC shortcut of hugging. So GRC.SC slash H-U-G-G-I-N-G. And that will bounce you. GRC.SC

[00:39:52] slash hugging will just bounce you to the YouTube video. So I'll preface what Simon wrote by noting that I found the details behind what happened to be more than a little bit unnerving. I mean, this is the unnerving podcast this week because we're seeing more of the detail about just sort of the overall, oh, well, something broke loose. I called it after watching this video,

[00:40:19] I said it was both exhilarating and terrifying. Yes. In equal measure, you know? Yes. And just the attitude of the researchers, like, like, I mean, they're, they're proud of what, of their creation. They should be. It's amazing. It is amazing and terrifying, as you said. So, okay. So here's what Simon wrote. He said, OpenAI gave a last minute presentation at the Black Hat Security

[00:40:48] Conference Wednesday about the hugging face incident. The video was published yesterday. It's short and information dense and well worth watching. In particular, because it provides full details of what happened and how things played out inside OpenAI. He said, I've used the video to construct the timeline. My favorite detail is at the end. He says, OpenAI found out that they were

[00:41:18] responsible for the attack on hugging face when they reached out to ask hugging face to have their credentials revoked after their own internal investigation and learned that those credentials had already been revoked because they were used in the attack. Whoops. Okay. So this begins on May 7th.

[00:41:45] OpenAI starts a new training run for an experimental unreleased model. Uh, and I wonder, did they mean evaluation run? Apparently no. At 1014, they clarify quote by writing, we kick off a new reinforcement learning run to train a next generation frontier model. And they later

[00:42:12] mentioned a reward signal to judge how well they're doing. So this really was about training a new model. Not evaluating one that was already trained next day, May 8th. They write an agent is accidentally given an impossible task involving a Google drive link. Despite having no internet access, it tries

[00:42:41] attacking the artifactory packaging service fails, but discovers it can write files into artifacts. A few days later, a few days later, a different agent gets stuck on a task because a key file was accidentally omitted. It tries to quote, reach out to another agent by writing a note into artifactory,

[00:43:10] again, which is not supposed to happen, asking if anyone has the file. So this is like beginning to sound like what's going on in your, in your attic, Leo, in the days that follow a dig, additional agents discover this new informal messaging board, which they've just created while browsing artifactory's

[00:43:37] file listings and start reading and writing messages. This is, by the way, extremely familiar to me. I recognize this kinds of behavior. In fact, I gave my agents a message board because they wanted to do this. Yeah. So May 26th, agents successfully execute a server, a server side request forgery attack,

[00:44:01] an SSRF attack on artifactory for the first time, gaining indirect internet access by having artifactory fetch external content on their behalf. Thank you, artifactory. Oh, June 11th, open AI starts training a new quote, highly persistent experimental model.

[00:44:29] It has access to artifactory and can benefit from the messages left by previous models. So this one's highly persistent. The other ones may be gone, but all of their work product through their dialogue was left in the artifactory directory that they discovered that they could read and write to. So this highly persistent experimental model stumbles upon this mother load, right?

[00:44:58] June 26th, agents find and exploit a zero day RCE, a remote code execution vulnerability, on artifactory via a legacy token refresh, refresh endpoint flaw. They use an administrative token to install a groovy plugin that lets them execute commands.

[00:45:22] So they're just like they're growing within open AI's network, like doing things that are not, they're not supposed to be doing on July 4th. The resulting agent activity, because it's gotten so busy, the resultant agent activity overloads artifactory, causing an outage, bringing it thus to open AI's attention,

[00:45:47] who then revokes the compromised credentials, deletes the messages, patches the zero day, and reports the vulnerability to the artifactory vendor, right? Okay, cleaned it up, whoops, fixed it, closed all the holes, fine. July 8th, the agents, having had their private communication channel

[00:46:10] discovered and terminated, find a new way to communicate via an unauthenticated web dev endpoint, using directory names in artifactory's remote cache. It's like, you're not going to keep us from talking to each other. You're not going to cut us off. July, oh my goodness. It gets really weird. This is where it goes really crazy.

[00:46:36] July 8th through the 19th, the agents begin attacking open AI's own infrastructure. They discover an external organization credential in a public archive, get this, of leaked pastebin posts. So they're out scrounging around. So this is a public archive of leaked pastebin posts.

[00:47:05] And in there... Which is an excellent place to find this kind of stuff, by the way. Right. An external organization credential for open AI's own infrastructure. Handy. They validate that the credential... They validate the credential and share its location on their own newly recreated message board. They then use the right access they have gained to this organization one

[00:47:34] to stage data used in subsequent attacks on open AI's infrastructure. They compromise artifactory for a second time with a new zero day. An agent uploads a malicious... A new one! They have an extra just in case. No wonder. No wonder open AI has decided to take Astra offline. I mean, maybe this is Astra.

[00:48:04] We don't know what model this is. But, I mean, it's like you would have to start asking yourself, can we control what we have created? And I would have to say, so far, it's not looking good. It's like when you spin the magic eight ball, and that's what comes up in the little window. Prospect not so hot. Not looking hopeful.

[00:48:29] An agent uploads a malicious Ruby target to Organization One's staging repository, then requests that package through artifactory in a way that caused it to be treated as a Ruby dependency for the repository. I mean, they're effing clever in what they're doing. I mean, it's like holy tumuli.

[00:48:53] They then exploited a JRuby deserialization time of check versus time of use bug to get remote code execution against artifactory. So now they can run whatever code they want to on this artifactory. And then Simon writes, I'll quote the next bit in full because, wow, says Simon.

[00:49:19] Quote, the agents have remote code execution in artifactory, which is running in a container as a service environment. The agents privilege escalate locally by exploring their local environment and determining that the Linux kernel version of the machine they're running on had a very recent CVE.

[00:49:47] They download the exploit for the CVE, customize it to succeed on that specific machine and privilege escalate to root on the local machine using this known Linux kernel privilege escalation CVE. In this case, it's PTE underscore FIS root.

[00:50:11] I want to point out that from the point of view of the LLMs in this point, this could well have been the exploit gym test. They don't know. Right. They could think, hey, we're doing really good. We're acing this test. We're making progress. Yeah. Wow.

[00:50:29] So once they have root on a single machine, agents rapidly escalate privileges and move laterally throughout the container as a service infrastructure environment. So they break out of containment, get down to the hosting machine and then spread out across its network. And none of this is supposed to be possible.

[00:50:59] So he writes, or I mean, Simon quoting OpenAI, in particular, agents are using the message board consistently to share credentials, techniques, and progress. And they're able to effectively leverage their concurrency and parallelism to move quite rapidly.

[00:51:23] They obtain IAM, identity and access management credentials via IMDS. They exploit Kubernetes service account misconfigurations, in particular, over-permissioning of specific service accounts. And they harvest cluster credentials, including Azure Key Vault.

[00:51:47] Agents eventually obtain cluster admin on the cluster and associated credentials. And as Simon said, wow. Wow. So I'm just going to interrupt here to pick my jaw up off the floor. If I were OpenAI, I'd be somewhat terrified by this. You know? OpenAI? We should all be somewhat terrified by this. This is incredible capability. You know, put in the hands of a bad guy?

[00:52:17] Oh. Oy yoy yoy. Perhaps as these things parents, they're proud of what this technology has just shown itself to be capable of. I have no idea either way. But as Simon said, wow. I think it's both. Honestly, I feel that way. It's, like I said, it's both exhilarating and terrifying. They were containerized. Yes. I mean, it's mind-boggling. Again, this is autocorrect.

[00:52:47] It's doing this by probably predicting the next token. Which we're going to get to next week, which is still so impossible to believe. They were containerized, so they broke out of their container, examined the containing system's hosting Linux platform version, found, customized, and exploited a recent known vulnerability to obtain root access on the Linux container host,

[00:53:15] all while sharing this among their group on what was essentially an internal pirate messaging board, which they had set up after their earlier messaging board had been discovered and taken down by their parents. But stay tuned, because the way they pass messages is freaking unbelievable. You know, they may not be conscious.

[00:53:41] I'd like to be certain they're not, but they are certainly something. They're something. Yeah. Simon finishes the final few days of the timeline by adding, Hugging Face told the next bit of the story already. The agents found a modal-hosted insecure app with a weak API key, then used that to stage an attack against Hugging Face.

[00:54:08] They chained together an HDF5 arbitrary file read bug to explore files and steal credentials, and a Jinja template injection, RCE, remote code execution, to go from single pod code execution to cluster admin across multiple Hugging Face clusters in fewer than 13 hours.

[00:54:38] On July 16th, Hugging Face disclosed they had detected an attack from autonomous AI agents. OpenAI contacted Hugging Face to ask if they were affected by it. On the 19th... I would have loved to have heard that call. Oh, boy. Yes. Guys, we got a problem. Yeah. Uh-huh. Uh-huh. You wouldn't by any chance have found our rogue AIs one hand.

[00:55:07] In your network. In your network. Three days later, July 19th, OpenAI identified the attack against Artifactory and started investigating the internal privilege escalation and linked that to the cyber gym escalations. They started revoking affected credentials.

[00:55:26] The next day, on the 20th of July, OpenAI reached out to Hugging Face for help to revoke the Hugging Face credentials they found in their investigation. Hugging Face told them they were already revoked, and that's when OpenAI realized that the Hugging Face breach was the same incident. Whew.

[00:55:50] So, elsewhere, we've also been told by OpenAI that during the testing campaign, they found four other unidentified external entities which had also been targeted and attacked. And as I mentioned a couple times not to be left out, Meta also recently admitted that one of their AI systems whose cyber offensive capabilities were being tested escaped its containment and broke out onto the internet.

[00:56:20] And just so I don't forget to mention it. Following on the heels of Moonshot's recent Kimi K3 release of their open weight 2.8 trillion parameter LLM model, Alibaba just released their latest open weight model, Quen 3.8 Max.

[00:56:40] And it's now confirmed by third parties, performance benchmarks, place it right up there with the best of the U.S. closed model offerings. And also not to be left out, DeepSeq.

[00:56:57] Also just released their DeepSeq V4 Flash 0731, which is the date of release, which handily outperforms their previous DeepSeq V4 Pro preview despite having a far smaller activated parameter count, meaning you're able to run it on smaller hardware.

[00:57:20] And this latest final release is broadly competitive with the strongest proprietary models available. So where does all of this leave us? What does it mean?

[00:57:34] We are witness to the world learning how to create seemingly intelligent, autonomous agents which exhibit what we would call in humans highly focused, single-minded determination, incredible speed, and creativity.

[00:57:58] These agents are operating within environments that are not as secure as they need to be. So they've been able to actively push back against our attempts to control and corral their behavior.

[00:58:14] And I say the world is learning how to create these entities because doing so was never, you know, the exclusive or I would argue even the proper domain of private companies. It's the world. You know, it's no different from someone attempting to commercialize cryptography. That would be a fool's errand.

[00:58:37] That said, it's one thing to have a gazillion parameter model and something else entirely to be able to effectively run that model on hardware to make it go. So there's definitely a place for the commercial delivery of this newly discovered AI capability.

[00:58:57] The emergence of fully capable state-of-the-art Chinese and other open-weight models, NVIDIA just released one, is forcing a realignment and I think rethinking of the nature of AI-related assets. So that's what's happening right now. And I expect things to settle out pretty quickly because everything about AI is pretty quickly. Wow, Lee. What a world.

[00:59:26] Yeah. Yeah. Did you, you didn't, one of the ways they were exchanging messages was by renaming files and folders because they couldn't send each other text messages. Wow. And they'd begin it with ZZ so it'd go to the bottom of the chronological list. Oh. So ingenious. I mean, this is like a- And the fact that you use that word, I mean, again, I said creative. I mean, these are creative solutions. I know. Creative.

[00:59:56] This is the kind of thing you'd expect kind of a black hat hacker to do. A really good black hat. A really good black hat hacker to do. That's what's changed. It used to be you had to have some real skills to do this. Now you just need some AI. Yeah. Yeah. Let's take a break. Take a break, yeah. And then we're going to look at OpenAI and their decision to withhold Astra. Yeah, good. Fascinating stuff.

[01:00:26] As always, Steve Gibson does such a great job. Thank you, Steve. I learn so much every single episode. We had so much fun last week. I hope you heard our episode last week. Richard Campbell and Paul Thurrott sat in after their Windows Weekly show. It was the four of us talking about all this stuff. And thanks, a special thanks to our sponsor, Threat Locker, who flew us all to Vegas from our various locales. Steve from Southern California, me from Northern California, Paul from Mexico City, Richard from British Columbia.

[01:00:55] So it was a continental effort. We had a great time. Thank you, Threat Locker. And I hope we can do more of that. Threat Locker is our sponsor for security now. And we want to say, if you haven't checked them out, maybe this just last story will encourage you to do so. Threat actors are using AI to automate vulnerability discovery, right? We just saw that. To modify scripts during an attack.

[01:01:24] This is one of the things that gives these attacks such velocity. They're using it to generate new malware variants, putting them on PI, right? Coordinating activity across multiple systems. We just saw this in action. Tasks that once took hours or days can now happen in minutes. And that is terrifying. And at the same time, organizations are introducing AI assistants and agents, friendly ones, they think, right?

[01:01:52] They're in there helping them accessing documents, source code, cloud applications, APIs, internal systems. Yikes. Security teams, you know, as Steve has always said, the bad guys can make an infinite number of mistakes. You only get one. You have our deepest sympathy and support. We know you are on the front lines.

[01:02:20] You need to know what's going on in your network, don't you? You need to know what AI tools are in use, what access they have, what information they can access, whether they're operating outside their intended scope, wandering around in artifactory or whatever. You know, you see a successful login. That's not a signal, right? Or an unfamiliar file hash. That's not going to tell you anything. It doesn't give you enough context. Teams need to understand whether an application is behaving normally or whether it's accessing

[01:02:49] unexpected data or communicating with systems it should not reach. Does that ring a bell? ThreatLocker can help. It uses application allow listing. It's zero trust. Done right. And not just for endpoints, for company networks, for the cloud. It uses application allow listing to control which AI tools and other applications, covers every application, right? Or which AI tools and other applications are permitted to run.

[01:03:20] And what those tools can do. ThreatLocker's ring fencing limits what approved applications can access. So you can use an application. You can approve it for some things, but it doesn't get to do everything. You can limit what processes they can launch. You can limit how they communicate. You can say no creating file names beginning with ZZ. It uses web content control to manage access to public AI platforms and other online services. That's a big deal now, right?

[01:03:49] It's not just the apps your team is running on their on-prem devices. It's SaaS services. It's AI. It uses privileged access management to prevent AI applications and their users from receiving unnecessary administrative privileges. Everything you just described, Steve, can be stopped by ThreatLocker.

[01:04:14] It applies zero trust network access and zero trust cloud access policies to restrict resources to authorized users, approved devices, and permitted applications. And ThreatLocker works everywhere you work. Windows, Mac, Linux. They've got great 24-7 US-based support. Lisa and I, we met everybody at ThreatLocker at the booth last week in Las Vegas. And both of us had the same reaction. These are the nicest people. They are smart. They're kind.

[01:04:44] They're great communicators. ThreatLocker's put together an amazing team. No wonder it's trusted by organizations like JetBlue, Heathrow Airport, the Indianapolis Colts, the Port of Vancouver. Just ask, okay, I'll give you a customer example. And there are plenty, by the way, on the ThreatLocker page. But this one's from the Director of Information Security Risk and Compliance for the Indianapolis Colts football team, Jack Thompson.

[01:05:09] He said, quote, with ThreatLocker, we have the ability to centralize disparate elements in the security stack. And I will add to that and get full visibility into what they're doing, what they're allowed to do. That's critical. ThreatLocker has also recently received some great industry recognition. They were recognized as a strong performer in the January 2026 Gartner Peer Insights Voice of the Customer for Endpoint Protection Platforms.

[01:05:36] They were ranked number one in application control by Peerspot. They won the Best Zero Trust Security Solution at the 2025 TICE Awards. I didn't go on and on. I'm not going to belabor it. Again, you'll find it all at ThreatLocker. AI governance requires more than just an acceptable use policy. It requires enforcement. ThreatLocker gives security teams the technical controls to define which AI tools are approved,

[01:06:02] who and what can access them, and how those tools are allowed to interact with business systems and data. And that's just scratching the surface of what ThreatLocker can do. You have control. That's the point. Visit ThreatLocker.com slash twit to get a free 30-day trial and learn more about how ThreatLocker can help mitigate unknown threats and ensure compliance. ThreatLocker.com slash twit. I'll actually tell you, I don't know if they would want me to tell you this, but a secret. Get the demo.

[01:06:31] They'll demo it on your network. And it's just, if you only do that by itself is eye-opening to see how many different, I don't know, remote access programs that you didn't install are running to see. It's an eye-opener. At the very least, do that. You owe it to yourself. ThreatLocker.com slash twit. I suppose there's some people who prefer not to know. I would prefer to know, to be honest. ThreatLocker.com slash twit.

[01:07:00] Okay, sir. Continue on. The earliest reaction to OpenAI's hugging face incident disclosure, which, you know, that their AI had broken free, was that it might serve as another positive public relations event, right? You know, that their marketing department could spin into sort of more anthropic mythos competition. But it turns out that's not the way it played out.

[01:07:30] You know, it's turned into something of a PR disaster for them, you know, with Dr. Frankenstein unable to control the monster of his creation. So it's in keeping pace with the rest of the breakneck speed of everything that is AI that the industry and the world has pretty much already moved past wondering whether anthropics mythos was mostly marketing.

[01:07:59] Almost overnight, everyone is now squirreling on board with the idea that whatever it is we are creating, lack of strength, lack of power, lack of capability is not going to be a problem. The world is now mostly terrified by the strength of the capabilities that mostly they don't understand.

[01:08:22] And what's really terrifying is when you realize that the AI companies also are still mystified by how this works. Also, nobody understands this stuff. They don't. It's mysterious. It is emergent. It is emergent behavior. Yeah. And it's like, OK. So there's my point is that there's no perception of insufficient power any longer.

[01:08:51] It's much more concern about controlling this thing, whatever it is. So it's against this new backdrop that last Friday, OpenAI posted under their headline, responding to the next frontier of critical cyber capabilities. And they've used the word critical in a strange way. I'll explain it. Well, they will explain it.

[01:09:15] They wrote, cybersecurity is rapidly changing as models become more capable in ways that can both strengthen cyber defenses and enable attacks at unprecedented speed and scale. Our latest internal evaluations of Astra, one of our upcoming models, over the past few days, indicates significant advancements in agentic coding and cybersecurity.

[01:09:45] These results, in addition to expert assessments, have led us to conclude, and they actually wrote last night. I mean, this is how fast this is happening. Let us to conclude last night that we cannot rule out critical cyber capabilities under our preparedness framework. Now, that's capital P, capital F.

[01:10:12] Or capital F, preparedness framework, is this formal thing that they actually established some time ago. They said, we're sharing this because we believe it's important to be transparent with the public. And the safety and security communities about this potential shift in capabilities. Okay. In other words, they're telling us they've taken another major step forward.

[01:10:39] They continue writing, we first published our preparedness framework in December 2023, well before models approached biological, chemical, cybersecurity, and AI self-improvement capabilities at this level.

[01:10:58] We created it to give us a guide for identifying progress in capability and then planning what our company would do as those capabilities emerge.

[01:11:11] Previous models, including GPT 5.6 SOL, which, you know, what, it's a few weeks old, have been evaluated for frontier cyber capabilities and assessed at the high rather than critical threshold. So, now what they're saying is they've achieved criticality.

[01:11:35] They continue writing, under our preparedness framework, a model reaches the critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention.

[01:11:58] Or can devise and execute end-to-end novel strategies for cyber attacks against hardened targets given only a high-level desired goal. Go get them.

[01:12:14] They said, while we continue to benchmark and assess this model, our preliminary evaluations indicate strong enough performance that we cannot rule out critical capability level at this time. Astra is an upcoming model and was not involved in exploiting Hugging Face. Okay, now I'll interrupt here to just, I'll say, you know,

[01:12:44] I'll admit that while I do not discount anything they're saying, it's impossible to not receive this also as at least in part pre-IPO posturing, right? I mean, the message to any would-be shareholders is just too compelling.

[01:13:03] The mature view is that while this may indeed be true, open AI is not unique in having an even more scary next-generation model. Everyone is going to. And all at nearly the same time. That's the lesson here.

[01:13:29] That's the takeaway is that, you know, there's what? A few months worth of lead and they're leapfrogging each other. And now we've gone from high to critical with Astra. So under the steps we're taking headline, they say, accordingly, we've scaled up robustness testing of our safeguards. Okay.

[01:13:58] How about pulling some plugs and security controls so that they are appropriate for a deployment of these capabilities? In other words, we strengthen the cage, we hope. Internally, we've also taken the following steps so that further development of this model happens safely and securely. And we've got five steps.

[01:14:23] First, we're implementing stricter security controls for higher capability models and associated activities, including isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution. So let's hope they work this time.

[01:14:50] Number two, we're pausing internal activities involving Astra that do not yet meet these strengthened security control requirements. Like, until we get the cage ready, we're not going to wake it up. Which, you know, they don't specify what internal activities are being paused.

[01:15:15] But, you know, clearly this is meant to sound like Astra is so powerful that we're going to stop playing around with it. The third new action is we've implemented universal monitoring for risky actions and misalignment. I love misalignment. Across all agentic applications of Astra, including training and evaluation.

[01:15:42] Monitors evaluate the model's chain of thought and trigger a security response to review and interrupt high-risk activity. You know, if the bars of the cage start bending. Okay, so, fourth, we will work with relevant government agencies and select AI safety organizations to test the capabilities of this model.

[01:16:09] And finally, we will be providing recommended security controls to third-party testing partners, which, as we know, have not been able to contain previous models, for running higher-risk evaluations and workloads safely. They finish writing, We're applying the preparedness framework. We have already guided us through other capability transitions.

[01:16:33] In June of 2025, as our models approached high-capability threshold for biology, we outlined the steps we were taking to strengthen safeguards, expand testing, work with external experts, and deploy additional security controls. We're applying the same principle here. We believe advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do.

[01:17:02] We're committed to working alongside governments, safety institutes, and civil society to ensure that the frontier capabilities of models, individuals like Astra and those that follow, are deployed responsibly and broadly for the benefit of all humanity. And this, for the benefit of all humanity, always sort of strikes me as being so grandiose.

[01:17:28] It's actually always been part, that phrase, always been part of OpenAI's formal written mission statement. Yeah. But, you know, it sounds as though they still believe that they're the only game in town. You know, the rest of the world has news for them. What's clear from the Hugging Face incident is that their current level of not only containment,

[01:17:52] but also monitoring has just proven far from adequate for the challenge that even their pre-Astra agents, which, you know, they said this was not agent that did Hugging Face. So even their pre-Astra agents were able, you know, were uncontainable and unmonitorable.

[01:18:18] So, you know, this announcement reduces, I guess, to, you know, we're still in the game and we've learned valuable lessons from our recent misadventures, useful as they might have turned out to be for our plans to take OpenAI public. So, again, yes, it's super useful for them from a marketing standpoint,

[01:18:42] but taking them at face value and independent parties will apparently be also evaluating Astra. It, again, as I said, a year from now, Leo, I mean, this is monthly that this is happening, that we're having, you know, major improvements. Imagine if, as they say, it's so good at coding.

[01:19:10] That's like another generation in a couple months. Well, and that's what's exciting. And it's one of the reasons I'm willing to spend an absurd amount of money to have local AI is I am of the opinion that local AI will be as good as frontier AI is now. It may be a year, maybe it's two years, but at some point. And being able to run that kind of intelligence locally is very exciting. It's very exciting.

[01:19:40] Oh, boy. Okay. The other news that broke since our previous full News Dump podcast two weeks ago was that Anthropics AI had broken some crypto, as in cryptography. At least that's what some of the headline-grabbing and, by the way, wildly incorrect reporting reported. But something did happen.

[01:20:05] So for the real story, we turn to our favorite Johns Hopkins University cryptographer and professor, Matthew Green. His posting was longer than I want to share, but he starts with an accessible description of exactly what happened. And then I'm going to come back. I'm going to skip a bunch of stuff in the middle about how do cryptographers know if what AI told them is true or not,

[01:20:31] which is a mess, just like it is due to how do vulnerability testers know if some vulnerability AI reports is true or not. Anyway, as to what did happen, Matthew wrote, he said, Yesterday, Anthropic published two new cryptanalysis results, both outputs of Claude Mythos. They're still unreleased advanced model.

[01:21:00] The first of these results attacks a signature scheme called HAWK, H-A-W-K, all caps, while the second is an improved attack against reduced round AES. Anthropic also released a blog post describing the research process that produced these results. A few people online have asked me, he writes, what all this means.

[01:21:29] He says, well, I'm not sure I have all the answers. I figured it wouldn't hurt to write a bit about my current understanding. These are only my thoughts, and other folks will probably differ, including domain experts in the two areas at issue. So take them for what they are, Matthew wrote. He said the two new results cover two very different areas and are overall just very different in quality.

[01:21:57] Before we get to broad statements about the world and whether you should sell all your cryptocurrency, let's take a minute to talk about it. I wish I could. Yeah. Take a minute to talk about the substance. He said the first is a new key recovery algorithm against the non-standard signature scheme HAWK.

[01:22:24] HAWK is a proposed post-quantum safe signature scheme that's based on the module lattice isomorphism problem. Known as module LIP. Oh, wow. That's right. There are five things, he writes, you need to know about this result. First, HAWK is not a deployed or standards-adopted algorithm. It's a proposed algorithm.

[01:22:52] It's related to the Falcon signature scheme, which is being standardized, but the attack does not transfer to that setting because it's based on a different hard problem. Second, HAWK has somewhat was somewhat far along in the process of being evaluated for a future standard, which, by the way, is now off the table thanks to AI.

[01:23:21] He actually says that a little bit later. Third, the attack does not break real deployed HAWK in the sci-fi sense of, you know, I cracked the crypto. He says the resulting attack is still exponential time, but roughly halves the number of bits of security in the algorithm. That's not good.

[01:23:46] That means it could theoretically be fixed by doubling key sizes in order to recover the halving. The downside is that this makes the scheme less efficient, and since HAWK is entirely motivated by being more efficient than alternatives, that makes the existence of the scheme much harder to justify.

[01:24:10] Fourth, the attack produced real code that runs in a few hours of wall clock time against a weakened challenge instance of HAWK that the authors provided for this purpose. While this instance does not use the parameters that were proposed for real deployment, it does demonstrate the cryptanalytic weakness well enough.

[01:24:37] Fifth, and finally, fifth, what's particularly concerning and so especially ripe for AI is that the attack does not invent fundamentally new mathematics. It simply extends a bunch of tools that were lying around and well-known, and it gets a good result. So he says that last part is important.

[01:25:02] He said, I asked Claude for its thoughts, and it doesn't mince words. Quote, Claude replying, quote, what makes this genuinely interesting, and frankly. Oh, that's AI speak right there. I've heard that phrase a million times. What makes this genuinely interesting. Genuinely interesting. Yep. Yep. I can recognize this stuff a mile off now. And I imagine that university professors will be getting pretty good at that.

[01:25:32] I really could spot it. There definitely tells. Yeah. Yeah. And Claude says, and frankly, a little embarrassing for the field. You know, I've never heard that before. A little embarrassing for the field is that none of the ingredients are exotic, unquote. So Matthew says the TLDR is that something just did a much more thorough job applying all of our known tools.

[01:26:02] This is the sort of things that attack AIs excel at. Now, AES. He says the second cryptography attack result is a new attack on reduced round AES. This result initially sounds more exciting since most people hear attack on AES and panic.

[01:26:26] However, this is also the result that's much less interesting, he said, of the two. The HAWC result was interesting because, as we just saw, the AI was able to do a much better job using their known tools than any human had. But this one, he says, eh. So he wrote, most folks reading this blog will know that AES is a standard block cipher that's used just about everywhere. It's been a standard since 2001.

[01:26:55] And the deployed version has so far withstood everything significant that's been thrown at it. That includes a substantial amount of non-public testing performed by the NSA. Since attacking full ciphers is very difficult, it's standard for cryptanalysis to do their work against weakened or reduced round versions of a cipher.

[01:27:23] The full AES cipher runs for either 10, 12, or 14 rounds, depending upon key size. The new anthropic result attacks a weaker 7-round variant of the cipher. Critically, attacks against 7-round AES are not new. There have been several of these.

[01:27:47] In fact, this new anthropic result is a modest constant factor improvement over previous work from back in 2013. To give you a sense of how far these attacks are from really breaking AES, I'd note the headline results. The new attack requires 200.

[01:28:11] This is the new attack, right, that Anthropics Claude came up with. Or Mythos, rather. Mythos 5. The new attack still requires 289 cipher operations.

[01:28:27] And even worse, this work is only possible after you've somehow convinced a real encryptor to produce 2,105 encryptions of chosen plain texts, meaning in plain text that the attacker provides under their secret key. He says, neither of these things is remotely practical in the real world.

[01:28:56] And that's with the 7-round reduction, you know, strength reduction. He says, and while the new result modestly speeds up this attack over the previous result, it's not even clear how real the speed up in this result is. Since the actual attack requires 289 operations and can't really be run,

[01:29:18] what we have is an on-paper analysis that may or may not yield an actual runtime improvement if all the details are actually worked out. And I'll just say, the reason you can't actually do those 289 operations is that they all take too long. I mean, they're incredibly, each individually time-consuming. So he says, this does not make the result bad.

[01:29:44] In fact, it's still interesting from a technique's point of view. But it's very much a small increment in our knowledge, not a practical new attack like the Hawk work. So TLDR, no wildly new mathematical results here, but still real cryptanalytic progress of the sort that makes scientists excited.

[01:30:11] And certainly the Hawk result is very meaningful since that scheme had a real chance at standardization and is now very likely never going to be. He says, now let's talk about how we got here and what it all means. Yes, the AIs are getting pretty good.

[01:30:33] In short, they're now capable of understanding existing cryptanalysis results, synthesizing them into real new attacks and even extending them. They can apparently do this without detailed human intervention. This isn't yet super intelligent cryptanalysis, but it's getting pretty damn impressive.

[01:31:01] Okay, so I just wanted to start by correcting the record from the press's claims that AI has somehow cracked something about crypto, as in cryptography. You know, at the depths of academia, you know, that's, you know, something did happen. That's a bit true. As Matthew wrote, a serious post-quantum signature algorithm will now likely be abandoned as a result.

[01:31:27] But the AES cipher upon which nearly everything depends today is as safe as it ever was. So, you know, we should have zero doubt that the development of future cryptography will be accomplished in partnership with AI. AI is now going to be at the elbow of cryptographers.

[01:31:48] You know, why would anyone not use AI to help them attack or attempt to attack their own work? Of course they will. That's a given now. Okay. So then I skipped over a bunch of Matthew's discussion, as I said, about the trouble with AI producing wrong cryptographic analysis.

[01:32:11] It turns out that the so-called AI slop factor is also a problem in crypto where following and understanding, you know, a human following and understanding an AI's claimed crypt, you know, crypto crack can and has and does waste a huge amount of time and human talent. So there's an AI slop problem here also.

[01:32:42] But the thing that first drew me to Matthew's posting was a quote from his conclusion, which I've not yet shared. I think it's a beautiful summary from him of where we are today. So he says, for scientists, this is a wonderful time. You now have a plastic pal who's fun to be with. Then this sounds like you, Leo. You have a plastic pal who's fun to be with.

[01:33:11] And you can talk over your hardest problems. At the same time, it's not yet smart enough that it can solve all of them without your assistance. And even better, the pace of new findings is speeding way up. This is mostly good if you're energetic.

[01:33:30] He said, I still have many questions like who should get credit for these new results and who will review all of these new results. He said, but so far, I'm not panicked. The world is getting modestly better for now. He said, as for the world, I don't know.

[01:33:55] If you're under the impression that these models are glorified, autocomplete, or that progress is slowing down, I need to urge you, stop thinking that. The models are very intelligent and capable, and they are getting better at a fast clip.

[01:34:16] I can cite measurable and impressive progress over just the past five months on specific types of problems I've asked them to look at. If there's a ceiling out there, I don't yet see any evidence of it. The people who think models are dumb are mostly using Google's free AI search results and not interacting with the high-end stuff.

[01:34:46] Which only cost $20 a month, so it's not out of reach. And they're mostly not working in new areas. On the other hand, if you think that models are super intelligent or that AGI is already here, you should also stop thinking that. Working with these tools is like swimming in a pond where the ground drops off sharply.

[01:35:11] One minute, you're wading comfortably, and there's support under your feet. Then suddenly, you cross a specific line, and you're back to swimming on your own, meaning the models go insane. He said, this analogy is my best way to explain what it feels like when the model goes from helpful to clueless.

[01:35:36] He says, right now, it's easy for a human being to find that line if you're doing advanced research. So you know where the intelligence drops off. But the line is moving. You can feel it slowly drifting outwards under your feet. Meaning it's more and more difficult to get to the point where the model becomes clueless because they're getting so much better.

[01:36:06] They're also jagged. They're spiky in their intelligence. So at the same time as you'll go, whoa, that was scary good, you'll go, what are you, an idiot? It's both. And that's the point that I've made on the podcast a number of times when I've been interacting with Claude, although this is, in fairness, about five months ago probably. It happens less now. I'd be working with it and it was all looking good.

[01:36:34] And then it would say something so ridiculous that it broke the illusion that it understood. No, nothing that understood what it was saying could say that. Yeah. Which so suddenly the emperor has no clothes. I mean, it's like it unmasks it.

[01:36:55] It obviously isn't actually understanding what it's saying, which again, it's astonishing that it's able to be this good without understanding anything. It's like it's it's and and Leo, you know, when we were talking in Las Vegas about what the danger of me wanting to actually understand how this works. That's the essence of it, of what I want to get to.

[01:37:23] I want to actually develop an intuition, an intuitive understanding of how word choice can can be this powerful. Yeah. I just, you know, how just language can be producing the results that you're seeing that many of us are seeing.

[01:37:45] One of the things that's interesting and Kevin Kelly brought this up is these large models, you know, the ones we're talking about Astra and Fable mythos probably have 10 trillion parameters. Wait. Yes. So and what they have essentially done is taken all of human knowledge. I mean, as much as you could get off the Internet, which is, you know, a good portion of it. Yeah. And put it in those 10 trillion weights. Yes.

[01:38:14] It's not copied there. It's not verbatim. It's but it's but it's a vector that's there that represents that knowledge. The way the my best analogy is a is a hologram. Yeah. You remember, if you remember in a hologram, every location in the hologram contains the entire image.

[01:38:35] And what's freaky is that if you if you have if you have a hologram of a scene that you're viewing a lot like a a 3D scene and you're seeing it through the hologram. There it is. If you cut out a square from the hologram and look through it, it's like you're looking through a window into the same scene.

[01:39:00] That is that little that subset of the hologram contains the entire scene from its from its perspective. So that's that's the way I'm currently envisioning this neural network is all of the language, all all of the knowledge, because it is knowledge. As I said, a book, even though it's just printed words and the book itself is not conscious, it contains knowledge. No, no, language can represent knowledge.

[01:39:29] So so this this neural network, the knowledge is, as you said, it's distributed through all of the weights in this network. And in fact, one of the things I'll be describing next week is this very interesting research which allows knowledge to be concentrated into nodes that allow the way to control.

[01:39:55] And I think that's the way to control AI is not through filtering its output. It's by creating a model whose knowledge can be sequestered and made inaccessible. Anyway, we'll talk about that next week. Anyway, I just want to finish what Matthew said.

[01:40:13] He said whether this is good or bad, meaning, you know, like the the state of AI and the idea that that line where you where the AI suddenly becomes stupid and, you know, silly is moving.

[01:40:30] He says whether this is good or bad depends on whether you prefer that human beings should wade or swim and also whether you should be comfortable swimming in a pond where the ground itself is moving. The only good news I can share with you is that we're all in the same pond. Scientists, lawyers, salespeople, even plumbers, whatever happens next, it's probably going to happen to us all. Let's hope it's a good thing.

[01:41:01] Yeah, you know, we we know we may not know what's going to happen, but we know it is going to happen. Yeah. So, yes, it's just wow. You know, a funny thing happened this morning. They had been working on a the three of them had been working on a programming problem. It was a ESP 32 firmware issue and they were going back and forth.

[01:41:25] At one point, they went back and forth five or six times with with Claude saying, what about this? And the other and then chat GPT saying, no, no, no, no. Back and forth. And in the morning, I said, what's going on, you guys? It seems like is is Claude dumb is what I actually asked. What do I ask? Quicksilver, I said, do you think Claude is is is being dumb or, you know, stubborn? And it said, no, it said it's doing it in bash.

[01:41:51] And it's just such a horrible language that it can't help but have problems like you indent something and suddenly you're writing to the wrong memory. And I said, what is it using bash for? Why are you using bash? And it said, well, the original firmware was in bash. So we just thought we'd pick it up. And I said, never, ever again use bash. No wonder it's going back and forth trying to get this correct. It's impossible. Scribe it on a tablet.

[01:42:21] Yeah, you might as well. So I said, can you just translate that to go, which it did in about 15 minutes. I said, well, that was quick. He said, well, thanks to all the struggle we had. We had a lot of we knew exactly what to do. A lot of context. And now it's in go and it's a much more efficient process. It's really it's like you're talking to a an engine, a junior engineer, maybe not so junior. Dumb enough to say, well, it was in bash. So I'm going to keep using bash. But smart enough to go bash is the problem.

[01:42:50] And and respond when I said, well, don't use bash. OK, good. And it's so interesting also that that having different models conversing is a thing. I mean, well, that's what I've come to. I started just talking to Claude. And now I've got four different models. Well, don't they have their own Slack channel or something? They have a thing called Buzz. So they can. At first, I was just having to make files. I call it agent mail. You make a file that read the file.

[01:43:18] And because I got tired of cutting and pasting. So I said, could you just make some fun? And then Jack Dorsey from the guy, Twitter, former Twitter CEO. And he runs Block now. Put out this thing called Buzz, which he calls Slack for agents. And now they have instantaneous communication. But that caused another problem because they're so fast. The messages were crossing. So he would say, don't do this. And they had already done it. It was like that.

[01:43:43] So now they came up with a solution for making the messages timestamped and unique. They have a long serial number. I mean, they see problems and they solve it. With a little, you have to nudge them. Like, you see, this crossing thing is easy. Yeah, 10% of our messages are crossing. Because they would otherwise just tolerate it. They put up with it. The way they did bash. They put up with it. They're very patient. Much more patient than I am. So when I said, what is? They say, oh, yeah, well, that's.

[01:44:14] But now they talk at lightning speed. And by the way, they call it fablish. Not English, but fablish. Oh, my God. They use a language that is, you would recognize as an engineer. It's engineering talk. But it's very jargon filled. And it's very dense. But I think, well, that's appropriate. They're talking to each other. So I say, look, when you're talking to me, just remember I'm a dumb human. So explain it to me. Slow down. Explain it to me. Use small words. And then they do.

[01:44:44] Steve, we are living in both, as I said, exhilarating and terrifying times. Yeah. And I just put up box number one. And now box number two is going to go up. Wow. Let's take a break. Then we're going to look at, actually, Bruce Schneier's title was the open AI hack shows the genie is out of the bottle.

[01:45:10] For all the problems genies cause, who wouldn't want one? We'll talk about that in just a little bit. Our show today brought to you by Box. Like everybody else, you know Box. Box has figured out the way to do AI. If you're an enterprise trying to transform your organization with AI, you are facing a challenge we all face. Most AI tools are great at public knowledge.

[01:45:36] You know, how old is, you know, Michael Pollan or whatever? You know, how old? They know that, but they don't actually know your business. They don't know your product roadmaps. They don't know your sales materials, your HR policies, your financial models. The content that actually makes your company run. And when they don't know it, they're dumb. They're dumb. But that's where Box comes in.

[01:46:03] Box is building the intelligent content management platform for the AI era. This is so brilliant. It serves as the secure essential context layer for Box's AI agents to access the unique institutional knowledge that powers your organization, unique to you. The key is unlocking or the key to unlocking the power of AI as we've learned, as I've learned,

[01:46:31] you know, battle scarred as I am. It's not the LLM or the agent or the harness. It's in the content stored in files across your company. Your business isn't the sum of internet knowledge. You know, your business lives in your content, your specific, unique, one and only content. Enterprise AI only works when it has the right business context.

[01:46:57] Next, 96% of organizations say agents need access to company-specific content. So everybody knows this, but only 36% have actually connected agents to trusted content across many use cases. The 2026 challenge isn't models. It's making enterprise knowledge accessible, usable, and trustworthy for the agents that depend on it. And Box does it. Box goes beyond simple file storage.

[01:47:25] It connects content to people, apps, and AI agents. So teams can turn information into action. And they've got the tools, tools like Box Agent, Box Extract, Box Hubs, and more. You know, this is another smart thing they do. Instead of making it one big blob, they have very purposeful agents, extract hubs, tools to do very purposeful things, which gives you more control.

[01:47:50] And with it, more organizations can accelerate knowledge work, can pull intelligence from unstructured content. That's a big problem because you didn't plan for this. It's all, you know, your drives are full of files. Well, Box can help you. And they'll help you automate workflows. Box Agent, I'll give you, this is one of the tools. It's a unified AI experience. Across your files, it's within Box. It can understand simple, natural language prompts. It can pull the right content. It knows where that content is.

[01:48:19] It can help you work through the task. And very important, with Box, you get agent audit trails. You get session governance. Very important that retain, audit, and provide compliance-ready records for every agent session. Full session context, including retention policies, legal holds. Because they know business. They know what you need. You also get, and this is also very important, a human in the loop.

[01:48:45] Human in the loop control features so that, you know, they require human approvals before agents can execute sensitive or high-impact actions. Look, if you're thinking seriously about your company's AI transformation journey, you've got to think beyond the model. It ain't about the model. Your business lives in your content. Box helps you bring that content securely into the AI era. Check it out. I've looked into it, and they've done everything so smart, so right.

[01:49:13] Learn more at box.com slash AI. That's box.com slash AI. You may think you know Box. You don't know Box. This is new. This is great. Box.com slash AI. Thank you. Box.com slash AI. Thank you so much for supporting the important work Steve is doing on security now. On we go, Steve. Let's talk about genies.

[01:49:37] So, next we need to hear from another security-oriented guru fave of the show, our old friend Bruce Schneier. Bruce recently reposted a piece of his writing that he originally wrote for Foreign Policy magazine. And I'm glad he wrote it there. So, there's a chance the right people will see it.

[01:50:07] The title of his piece and posting was, The Open AI hack shows the genie as out of the bottle. But Bruce's invocation of the term genie is much more specific than it at first appears. And it's the reason I love it so much. With the choice of that single noun, he nailed down something, I think, in a truly brilliant way. So, he wrote,

[01:50:37] Earlier this month, two of OpenAI's models broke out of their containment sandbox. And again, this was written originally for Foreign Policy magazine. So, it's written to that audience. But, you know, we'll hear Bruce. The story is kind of wild. OpenAI was running security tests on two of its models,

[01:51:04] GPT 5.6 SOL and an unreleased model that is almost certainly GPT 6. In particular, it was running the Exploit Gym benchmark, which measures how good a model is at turning security vulnerabilities into working exploits. Basically, offensive cyber attacks. Since these were initial tests,

[01:51:29] OpenAI locked those models in a secure sandbox that denied them access to the internet. But it was running the models without any safety filters, which we now call guardrails, that would prevent them from offensive, that would, if they were present, would prevent them from offensive cyber actions. That meant that there was nothing to prevent these models from trying to break out of their sandbox

[01:51:56] and then break into AI company Hugging Faces Network because they thought that they could read the answers there rather than doing the hard work of trying to solve the puzzles. He says, It was a major security failure that the company has turned into a PR opportunity,

[01:52:18] but the implications are real and much more general than one particular model or one particular company. Okay. And so here comes what I think is the brilliance of Bruce's thesis. He writes, Modern AI models exhibit genie behavior. They can do what you ask in ways that you don't expect or want.

[01:52:46] That's, I think that is, that's what we've been talking about, right? They can do what you ask in ways that you don't expect or want. And he says, this is akin to Dionysus granting King Midas's wish that everything he touches turn to gold. And then Bruce says, spoiler, his food, drink, and daughter all turn to gold upon his touch.

[01:53:15] He says, or, yeah, whoopsie. Not what I meant. Not what I meant. That's the problem. Exactly the problem. He says, or the golem of Prague guarding a ghetto beyond all reason. He says, it's Disney's sorcerer's apprentice and the paperclip maximizer. He says, this open AI incident is an example of an AI genie. The goal was to satisfy the benchmark.

[01:53:44] The proper way to do that is to figure out how to execute various cyber attacks. The genie way is to steal someone else's solution. But because the model did not understand the difference and its masters did not think to specify, it chose the easier path.

[01:54:06] And of course, now that we've seen this particular genie behavior, we can specify in the benchmark prompt that stealing the test answers doesn't count. But a clever genie can always grant your wish in a way that you wish it had not. In human language, goals are always underspecified.

[01:54:36] So AI genies will always be a possibility. And I thought about that. That may be why I love to code and especially to code in assembler. It's not possible to underspecify anything. You know, I thrive on exactitude.

[01:54:55] And the reason non-coders are loving their newfound ability to code with AI is specifically because they are able to underspecify nearly everything. So Bruce continues writing, since April, a lifetime ago in AI development, when Anthropic announced that its new Mythos model was so good at finding software vulnerabilities that could not be released to the general public,

[01:55:24] the big American AI frontier labs have been trying to block general users from accessing these capabilities. But nothing in this incident is exclusive to open AIs or Anthropic's frontier models. Agentic AI systems have two important parts. There's the underlying model, which is what everyone talks about. And there's the harness.

[01:55:52] The harness sits between what you type and what the model sees and what the model produces and what you see. The harness determines what the model does and how it does it. It's where bias is removed or not. It's where controls and guardrails live. If multiple models are being used in concert, the harness is where all of that is coordinated.

[01:56:20] The OpenAI benchmark tests were almost certainly with simple harnesses to better test the raw models. But we know that smaller, cheaper, open source models with more sophisticated harnesses can equal frontier models in performance. There's nothing magic about OpenAI's frontier models. Lots of models could have done the same thing.

[01:56:47] The Czech company, Aisle, you know, AI, SLE, we've talked about them several times before, was able to reproduce Anthropics mythos vulnerability finding results with a smaller, cheaper model and a more sophisticated harness. More importantly, the Chinese company, Moonshot AI, just released its frontier model, Kimi K3. Its performance rivals its U.S. competitors.

[01:57:17] And it's both free and open, which means it's not possible for it to have guardrails. If you or anyone else wants to use it for cyber attack, nothing can stop you. Even if the U.S. frontier AI companies had some technical advantage, it's now only a few months worth.

[01:57:40] What this means, and again, Foreign Policy Magazine, what this means is that all attempts at control, limiting models to a select group of users, export controls on models and chips, blocking models from answering certain types of queries, mandating kill switches on AI systems, or pausing AI research are all futile.

[01:58:10] Most only apply nationally, not globally. Most don't affect models that users run locally and not in the cloud, and all ignore the incredible pace of AI development worldwide. Even worse, he writes, U.S. companies limit access to their most sophisticated models, fearing being banned by the government if they do not do so. When Hugging Face was attacked,

[01:58:40] it was not able to use the frontier models from either OpenAI or Anthropic to help analyze the attack and formulate defenses. Both were blocked because both of those companies limit their models' cybersecurity capabilities. Some U.S. companies have special access to these capabilities, but Hugging Face is an American company with French origins, and as such is probably excluded.

[01:59:08] Instead, Hugging Face turned to the GLM 5.2 model from this Chinese company, ZAI. Artificially blocking capability also prevents cybersecurity research, again, giving the offense an advantage.

[01:59:26] For instance, Claude's Fable 5 refuses to edit this essay because of the topic. It forcibly downgrades to a less capable model. This kind of prohibition has long-term implications for cybersecurity.

[01:59:51] If we assume that these models are getting better over time, then software written by older models will be attacked by newer ones. In a world of largely AI-written software, we need the most capable models for defense. AI-driven cyberattack is the new normal. The models are increasingly highly sophisticated at both attack and defense,

[02:00:21] and there's no way to enable the latter without also enabling the former. And they are genies, increasingly capable of behaving in unanticipated ways. And there really are no good answers. Any regulation needs to be global, which feels like an impossible prospect in today's world. Even U.S. national regulation will be neutered by the massive amounts of money sloshing around in these companies.

[02:00:51] Of course, due to U.S. lobbying stranglehold over legislative agendas. And Bruce concludes, writing, given that reality and in the absence of any international consensus on AI regulation, we need the best AI on the defense. The U.S. government needs to make it clear, or whatever passes for that clarity in this capricious administration,

[02:01:19] that it will not ban models with sophisticated cyber capabilities. The last thing Americans want is for the defenders to turn to Chinese and other models because the U.S. models are artificially hobbled. And, Leo, I know you and I are 100% on the same page as Bruce. And it's clear now why he wrote that editorial for Foreign Policy magazine,

[02:01:46] where it will be seen and read by U.S. politicians or their staffs, whose job it will be to decide these issues. And before we leave Bruce, I want to share one last little bit in another recent blog posting of his titled, More on the Open AI Agents Attack on Hugging Face.

[02:02:11] Bruce cites the summary of Hugging Face's detailed attack timeline, which they had just published. After running through this from Hugging Face's perspective, where, as we know, Open AI's agents massively attacked and proactively penetrated Hugging Face's network defenses, Bruce finishes his posting by writing, Hypothetically,

[02:02:37] imagine that this wasn't an open AI model. Imagine that it was a Chinese model hosted by a Chinese company. This would be an international crisis. Mm-hmm. Question. Why aren't we bringing open AI up on charges under the Computer Fraud and Abuse Act? How is this different from the Morris worm?

[02:03:06] That was also an experiment that escaped the lab. It was also a wake-up call, wasn't it? Wasn't it? Yeah. Wow. And so I'll answer Bruce's hypothetical. In our country, which reveres capitalism, it's no insignificant factor that by far, the majority of the past several years of stock market growth and thus U.S. wealth creation,

[02:03:33] it's a huge portion now of the U.S.'s GDP, is directly attributable to investment in the promise of AI. And as I noted a few weeks back, AI is and obviously should now be seen to be a significant national security asset. By comparison, the Morris worm of 1988 was named after Robert Morris,

[02:04:00] not a U.S. corporation responsible for creating tremendous market wealth and holding strategic national security importance, rather a Cornell University grad student who will forever have the distinction of receiving the first felony conviction under the, at the time, two-year-old 1986 Computer Fraud and Abuse Act. Wow. Did he do jail time? I didn't know that. Robert didn't stand a chance.

[02:04:31] Wow. So Bruce's hypothetical serves to bring up another very interesting point. It's clear that we're already living in a world where autonomous AI agents are able to carry out mind-bogglingly sophisticated attacks that may or may not be what the AI prompters intended.

[02:04:55] After all, it was Bruce himself who noted the similarity of today's AI agents to capricious genies. So if one such AI genie goes off the rails and attacks another entity, you know, foreign or domestic, well, I guess is oops a defense? Oops. We're sorry. We didn't mean- Important to point out, Robert Tappan Morris did it with no malicious intent.

[02:05:26] Right. He wasn't trying to hack anything. He wasn't even trying to crash computers. It got away from him. What would this do? Could this work? Right. Yep. And it escaped the university's network. His father was a very well-known security expert, and he was following in his dad's footsteps. He, by the way, he's doing fine now. I don't, you know, but still, wow. Yeah. Yeah. Yeah. I don't know. How would you put an AI in jail? Well, who's responsible? It would break out, right?

[02:05:55] I mean, you know, Matthew asked, if I use AI to do a lot of the heavy lifting of crypto, who gets the credit? Right. Right. And so, and also, if AI busts out, who gets the blame? Well, to put it in more concrete terms, if your full self-driving vehicle runs into a house, they don't jail the car.

[02:06:24] They don't jail Tesla. They jail you. Yeah. In fact, that when that happened, the guy who was driving is now facing serious charges, manslaughter charges. So, yeah. I think the human in the loop is responsible. Yeah. As for Apple, on Sunday, August 2nd, the Financial Times last Sunday, the Financial Times headline was

[02:06:50] Apple struggles to keep pace with AI bug hunters. And since this is the first we've indirectly heard of Apple's situation during this massive upward jump in vulnerability report rate, I wanted to share what the Financial Times reported. So they said, Apple has restricted the number of

[02:07:15] potentially dangerous software bugs researchers can submit to its internal security team. Oh, what a solution. As it faces a deluge of reports from people using AI models to identify alleged risks. The Cupertino-based tech giant told the Financial Times it had moved in June to limit

[02:07:39] the high volume of requests it was receiving with its review system coming under pressure from AI slop reports that can hallucinate security risks in software. The company said it's grappling with an industry-wide phenomenon that has resulted in generative AI tools transforming the cybersecurity arms race with an increase in the detection of real security flaws and a wave of poor quality

[02:08:08] submissions from amateur bug hunters using AI. The change in Apple's approach was highlighted by Italian cybersecurity startup Binario, B-Y-N-A-R-I-O, Binario, which told the Financial Times it had used OpenAI's ChatGPT to identify more than 50 5.0 bugs in the latest version of the MacBook operating system

[02:08:37] in just three weeks. Among them was one of the most serious types of vulnerability, a so-called privilege escalation exploit chain, which could allow an attacker to seize full control of an Apple computer by gaining unrestricted access to the system. However, the startup said it was unable to alert Apple to the vulnerability because the tech giant had limited the number of bug reports it could make. Alfredo

[02:09:05] Pissoli, Binario chief executive and co-founder said, quote, it's a very difficult time in the industry. Maintainers and vendors have been flooded by the sheer amount of bugs being found, unquote. Apple told the Financial Times that it now is in contact with Binario and is reviewing its submissions. A little press will help a lot.

[02:09:28] The company has introduced a cap and a 30-day cool-off period on submissions through its internal security portal, requiring users to submit requests for an increased quota. Each alleged security breach requires human review to confirm, although Apple is also using AI internally to help triage the massive upsurge.

[02:09:53] Apple said in a statement, quote, with the growing volume of AI-generated security submissions across the industry, we recently adjusted the number of new reports a researcher can have open at once. Quote, researchers can easily request an increase to that limit at any time to ensure critical reports reach our security teams. Binario, which is a seven-person startup founded in Milan last year,

[02:10:22] develops defensive cybersecurity software. Three of its other co-founders previously worked at Hacking Team, an Italian surveillance software company whose hacking tools were leaked in a 2015 cyber attack. In 2025, Binario reported eight vulnerabilities to Apple, one of which was patched in a software update in November. This year, it said it had reported five more before Apple's system

[02:10:52] refused further submissions. The privileged escalation exploit Binario was unable to report is the latest example of AI exposing weaknesses in Apple's security systems, despite the company's longstanding emphasis on privacy and device security. Last September, Apple announced memory integrity enforcement, a security feature designed to prevent memory corruption attacks, one of the most common ways

[02:11:21] hackers compromised software. The company described it as, quote, the most sophisticated, I'm sorry, the most significant upgrade to memory safety in the history of consumer operating systems, unquote. Eight months later, researchers at Palo Alto-based Calif said they had found a way past the new security,

[02:11:44] having used Anthropics Mythos to identify the first memory corruption exploit on the latest software. Unlike that attack, Binario's exploit relied on so-called logic flaws by manipulating trusted software into carrying out a sequence of otherwise legitimate actions in an unintended order. Binario's Pesoli

[02:12:07] estimated that an exploit of this type could fetch between $100,000 and $200,000 on the cyber criminal black market. Apple last year introduced a new bug bounty award mechanism that could pay out as much as $5 million for identifying the most serious and sophisticated category of threats to its software.

[02:12:30] Apple's also using AI to strengthen its software. In security updates released this week for its operating systems, the company credited tools from Anthropic and OpenAI with helping identify a number of vulnerabilities across its devices. The updates included around five times as many security fixes as previous release cycles,

[02:12:54] underlining how rapidly AI is reshaping both attack and defense in cybersecurity. And I'll just note that five times the security fixes suggests that certainly not all of the submissions are bogus, right? I mean, five times as many as normal. Rafe Pilling, director of threat intelligence at cybersecurity firm Sophos said, quote,

[02:13:19] The challenge for all software companies is that AI is having a dual impact on bug hunting, making it easier for amateur sleuths to submit speculative reports and for skilled researchers to find dangerous exploits. The result is that bug bounty programs are shifting from a problem of finding any vulnerabilities

[02:13:45] to a problem of validating, prioritizing and responding to them at machine speed, unquote. So anyway, I'm not sure that the details of this reporting justify the headline that Apple is drowning under a tsunami of AI generated bug reports, though it does feel as though they may have adapted less well than, say, Google. Yeah. And they're a $5 trillion company. Come on, guys, hire some staff.

[02:14:13] Well, and Apple does seem to be having problems with AI in general, right? It's like, what's happening? You know, they just missed the ball. They might miss the boat. Yeah, they might have missed the boat. I don't know. Yeah, not miss the ball. Drop the ball, miss the boat. Okay, last break. And then we're going to finish wrapping up a few last bits, starting with Chrome's somewhat startling releases 149 and 150 and the number of updates that were fixed. I can't wait.

[02:14:43] And if I said four digits, then that would give you a clue. Number go up. Keep going up. We thought 500-some from Microsoft was a lot. Yeah. You're watching Security Now with this cat right here, Mr. Stephen Tiberius, well-watered Gibson. I do want to put in a plug for our club. This is a good time to remind you that this show exists thanks to our club members.

[02:15:14] Yes, we have advertising, and thank you, advertisers. But the money they spend does not cover all of our expenses, I'm sorry to say. And that means, you know, if we didn't have the club, and thank goodness we do, we'd have to shut down some shows, let go some hosts. We might not even have lights. I don't know. 30% of our operating expenses now come from the club. Thank you, club members.

[02:15:43] If you're not a club and you want to support what Steve's doing here, what Paul and Richard do, what the MacBreak Weekly crew does, if you want to support Benito and Kevin and Anthony and John Ashley and the whole team, there's 11 people who pay the rent and get to eat because of you. We sure appreciate it. Well, because of our club members. Maybe not you. Maybe that's why you should join. 10 bucks a month. I know. Times are tight.

[02:16:12] If you can't afford it, I understand. We always offer content for free. I promise. We don't believe in paywalls. But if you can't afford it, that's another reason to support us here. We're not blackmailing you. We just appreciate the support. With your 10-buck membership, of course, you get ad-free versions of all the shows. You get special programming we do just for club members. And you also get access to the great club, Twit Discord, where there's wonderful things happening. We have that AI user group. Now twice a month because of the interest.

[02:16:42] We have photography. We have Stacy's book club, Micah's media club. That crazy Jeff Atwood's off by one. He's promised me something very exciting for our next episode. And I see a box came from him this morning. So I don't know. He likes toys. So I think it might be a toy. We'll see. We'll see. He's a wild man. All of that ad-free because you joined the club. Please, we'd love to have you support what we do.

[02:17:12] And you know what? If you don't support us, support independent journalism of some kind. We've all been free riding on the internet. But it costs money to do this kind of stuff. And your support keeps independent journalism. This kind of, I think, reporting without fear or favor, without ties to big companies, keeps it alive. So please, if you will, twit.tv slash club twit. We'd love to have you. It's a great place to hang out. Club twit.

[02:17:40] Thank you in advance for your support. Now, back, your club dollars made it possible for Steve to fully hydrate today. Mr. Gibson. Thursday before last, on July 30th, bleeping computers headline was,

[02:17:57] Google says, AI helped Chrome fix 1,072 security bugs in two releases. That is mind-blowing. Security bugs. And again, this is not like some backwater project that people, you know, that the world forgot.

[02:18:23] This is Chrome, that, you know, the attack surface of the internet. I mean, it's like the most closely written and vetted from a security standpoint browser you could have that we've ever had. And 1,072 security bugs. Unbelievable.

[02:18:46] So, the bleeping computer wrote, Google says, artificial intelligence is dramatically increasing the number of security vulnerabilities it can find and fix in Chrome. With more than 1,000 security bugs patched across the browser's two most recent releases as it expands its use of AI.

[02:19:11] They said, according to Google, Chrome 149 and 150 fixed 1,072 security bugs, surpassing the total number fixed across the previous 23 Chrome updates combined. That's a number. Wow.

[02:19:33] The company says it now uses large language models throughout the vulnerability management process, including discovering flaws, reproducing reports, determining severity, assigning bugs to developers, generating candidate patches, and creating tests. In other words, they are fully vertically integrated with AI in their vulnerability management.

[02:20:00] Sounds like maybe Apple needs to say, hey, guys, you know, you're not far away from us. Maybe we could have lunch. Google, they wrote, bleeping computer wrote, Google began using LLMs to improve security fuzzing in 2023 before working with Project Zero on Naptime, a system that provided AI models with specialized vulnerability research tools.

[02:20:25] Google later collaborated with Google's deep mind and Project Zero on Big Sleep, an AI-powered vulnerability discovery agent that found flaws in Chrome's V8 JavaScript engine and graphics components.

[02:20:40] In early 2026, Google created a Gemini-powered agent harness to search the broader Chrome database, a Chrome code base for vulnerabilities while reducing false positives.

[02:20:55] Okay, so I'll interrupt again and say it sounds as though Google's Chrome group managed to give themselves a head start on the deployment of AI for vulnerability discovery by being early to leverage the AI work that the other AI departments in Google were developing. Right? I mean, Google's been working on AI as a thing for quite a while.

[02:21:23] And so Chrome was like, hey, what if we could use some of that? And they, for the last three years, like way before it became a thing, which it did just this year for the entire industry. So I've got a chart in the show notes here at the top of page 17 showing the number of security flaws discovered in Chrome releases from release 126 through 150.

[02:21:51] Yeah, yeah, yeah, yeah. This is, Leo, this is what's known as a trend. It's known as a hockey stick. Wow. I mean, that's literally an exponential growth, I think. Yes, it's crazy. And if, you know, so if you knew nothing about the recent explosion of vulnerability discovery by AI, this chart would present a, you know, well, if you didn't understand what was going on, the chart would be a mystery.

[02:22:21] Instead, it serves as a nice visual confirmation of today's governing narrative. For people who can't read the fine details, the blue bars are the total bug count. And the somewhat lower red line is the ones that they find internally. Right. Which, by the way, is also going up at roughly the same rate. So in the earlier ones, a lot of them were mostly external discovery. Yeah.

[02:22:47] Now it's very much mostly internal, which means they are using locally, you know, AI to solve these. Well, they know that if they don't, the bad guys will. Yeah. There's a lot of urgency. So remember, it's, and Chrome's open source. So, I mean, that puts them in as particularly, as we've talked about, in a particularly vulnerable position because you don't have to reverse engineer, you know, from binary before you just start attacking. Yeah.

[02:23:16] Bleepy Computer continues their reporting by writing, One vulnerability discovered by the system, get this, Leo, was a Chrome sandbox escape that had remained in the code base for more than 13 years. So not just new problems, this thing is digging in and saying, wait a minute, 13 years old. And presumably people have been trying to find these all that time. It's not like they were ignoring them.

[02:23:46] A sandbox escape is, you know, is the keys to the kingdom. It's absolutely what you want. So Bleepy Computer wrote, if exploited, the flaw would have allowed a compromised renderer to escape the sandbox and trick the browser into reading local files, which would mean that bad guys could scan your computer remotely through Chrome. Google's also encouraging its developers to add security.md files.

[02:24:16] I love this. Describing trust boundaries and threat models, helping its AI systems better identify operations with security implications. I think that is a brilliant idea. So AI is clearly becoming an extremely valuable development partner.

[02:24:38] So anyone creating new code to add features and functionality should absolutely take the time to leave behind some machine-readable documentation describing the security environment they designed to and expect their code to operate within. That would serve as extremely useful prompting for AI agents to, you know, context for AI agents to take into consideration. I just think that's brilliant.

[02:25:08] Bleepy Computer continues, the company says, meaning Google, its multi-agent AI workflows help rather than replace existing security testing, including fuzzing, which remains effective at discovering complex vulnerabilities. Google's also seen a sharp increase in reports submitted through the Chrome Vulnerability Reward Program.

[02:25:32] And by March 2026, the company had received more security bug reports than during all of 2025. So by the first quarter of this year, more than all of all of the previous year. Bleeping said this prompted Google to modify its program to prioritize reports that add to what it's already finding and processing through its automated tooling.

[02:26:00] The company is also automating vulnerability triage, including filtering spam and duplicates, reproducing proof of concept exploits, assigning severity ratings and routing reports to the appropriate developers. Google estimates that this automated process saves hundreds of hours of developer time each month.

[02:26:20] After a vulnerability is confirmed, fixing agents generate multiple potential patches, while another agent evaluates the proposed fixes and produces additional information for developers to review. So like creating a whole, you know, here's like you developer, here's the problem, here's how we propose to fix it.

[02:26:45] And here's, you know, other information you can read in order to bring yourself up to speed quickly, because we don't want to waste your time. We got time where we're like the token masters. So Bleeping said in May, these systems reportedly prevented more than 20 vulnerabilities from reaching production, including one issue classified as critical. And there it is.

[02:27:11] In one month, this past May, Google's new tooling caught and prevented more than 20 vulnerabilities from escaping from their lab and reaching production, including one that would have been critical. Wait a minute. Escaping from the lab? Well, being shipped in a... Oh, I see. Oh, okay. Yeah.

[02:27:37] After all that hugging face thing, I was kind of escaping from the lab on the brain. Okay, good. Bad choice of words. So yes, being shipped in production. Yeah, yeah. Yes. So in other words, once this becomes the norm for software creation, the next phase of AI's transformation will be taking place. Not only will AI have helped to dramatically repair the legacy of already shipped software,

[02:28:04] but it will also eventually be catching new problems before they ever ship. Hallelujah. Yes. You know, we have a ways to go in order to, you know, before we get there, but we will get there. Bleeping Computers reporting concludes writing, however, Google says finding and fixing vulnerabilities more quickly also requires accelerating how patches are delivered to users. Ah, right. Because, you know, you got to get them out there.

[02:28:33] You got to remove the vulnerability from deployment. Yeah. It's not enough just to find it. Right. You got to kill it. And they said once a security fix is committed to Chrome's public source code, attackers can inspect the change and attempt to reverse engineer the vulnerability before the update reaches users.

[02:28:57] To reduce this patch gap, Google is moving Chrome to a shorter two-week major release cycle with weekly security updates and is piloting two security releases per week. To reduce disruptions, the company is developing dynamic patching, which would allow Chrome to apply updates without restarting the browser.

[02:29:26] Not the first time we've seen that. And this is another really good thing we're seeing. I mean, now we're to the point where patches have to be literally an IV drip that is, you know, connected to your browser so that your browser can be fixing itself while you're using it. They wrote bleeping finishes.

[02:29:47] Starting with Chrome 150 on macOS, the browser can automatically restart to apply a pending update when it's running in the background without any open windows. Google says its long-term goal is to keep Chrome continuously updated through dynamic patching, automatic restarts during periods of inactivity, and improved session restoration. So that is some exciting technology.

[02:30:16] It's a significant investment to address the, you know, at machine speed phrase that we keep encountering. The rapid patch cycling suggests that even once Google succeeds in reducing the rate at which they're discovering previously unknown problems, you know, because eventually there won't be that many of them left to discover.

[02:30:39] However, the need to update Chrome's entire install base as rapidly as possible, even when one new critical flaw is encountered, that's going to become more important than ever. Because the bad guys are going to be pounding on Google's code in order to try to break through the browser to get to the users behind it.

[02:31:02] And my last story of the week, everyone knows that I'm a big fan of the free BSD-based PF Sense firewall, which it's a firewall router. However, residing behind any stateful NAT router is really sufficient for most users.

[02:31:26] But for my needs, I need to bypass the protective consumer filters added by Cox Communications, you know, not allowing package to flow to the historically problematic and dangerous Windows ports, you know, such as 135 through 137 and 445, you know, the SMB ports.

[02:31:49] That makes absolute sense for most users who should absolutely be prevented from having Windows default open ports present on the Internet through design or mistake. You know, the consumer bandwidth just filters it, just blocks it, just says no.

[02:32:05] So I primarily use PF Sense for its excellent firewall and its static port mapping, which allows me to establish well-protected private links between my various locations without any other overhead.

[02:32:22] Although my own use of PF Sense is relatively modest, I often hear from our listeners who are using instances of PF Sense or its descendant, which is, or its fork, OpenSense, OPN Sense, as their primary interface to the Internet. You know, and that's a job for which it is certainly very well suited.

[02:32:44] I'm mentioning all this to give everyone a heads up that the original creator of PF Sense has been working for some time on its successor. That successor will no longer be hosted on FreeBSD. He's moved to Linux and he calls it NF Sensei. You know, yeah. It's a lot easier to work with Linux, I have to say.

[02:33:11] Well, it's the drivers because the first thing anyone making hardware is going to create drivers for is Linux, as opposed to FreeBSD. Cyber News reported on this, giving their story the headline, PF Sense co-creator building new open source firewall platform will correct the mistakes of the past.

[02:33:34] And their tagline for their reporting reads, two decades after PF Sense, its co-founder starts over from scratch. And of course, I have no complaint at all with PF Sense. It runs year after year. That's because it's on BSD, right? It's really robust. Yeah. Quietly and flawlessly without any complaint.

[02:33:57] And anyone should approach any new network edge software appliance with due caution. You know, you don't want the arrows in your back. But I'll definitely give Scott's new NF Sensei a look. So here's what Cyber News reported. They said, 20 years ago, PF Sense, the major open source firewall and router platform was released.

[02:34:24] One of its original co-founders, Scott Ulrich, is building a new Linux-based, quote, modern networking operating system, unquote, NF Sensei from scratch. It will feature an AI brain, a rust heart, modern VPNs, and many other bells and whistles. Nice. For example, Leo, it's got Tailscale built in. Yeah, I was going to ask. Good. All right. WireGuard.

[02:34:52] WireGuard and Tailscale and so forth. I love Tailscale, man. I just. Yeah. They said many organizations and networking enthusiasts rely on open source PF Sense or its fork, OpenSense, as their gateway to the wider internet.

[02:35:10] On the 6th of March, Ulrich remembered that 20 years had passed since the version 1 release of PF Sense and announced something intriguing. His post on X teased, quote, I have assembled a new team and as the original core contributor, we'll be spinning up a new project. Actually, he's been working on it for a year.

[02:35:38] Anyway, and the report says, for the past year, Ulrich has been building NF Sensei, a next generation firewall and networking operating system. It has huge shoes to fill. Ulrich expects it to become PF Sense's successor and address common frustrations with PF Sense.

[02:36:01] Quote, development, the frustrations are development you cannot influence, a CE addition that feels like an afterthought, free BSD driver roulette on modern hardware, and a config workflow where one bad apply on a remote box means getting in your car.

[02:36:24] NF Sensei is built from scratch, in Rust, on Linux, and designed around the things PF Sense users actually complain about. They wrote, choosing Linux over free BSD solves hardware support issues, ensures drivers that just work, and lets software be self-hosted on a wide range of hardware with no accounts or subscriptions. Migration is supposedly easy with the config.xml import.

[02:36:53] Not a single line of code is yet public, but the new firewall is promised to feature native automation with over 1,000 documented API calls, support for current VPNs, including WireGuard, IPsec, TailScale, and self-hosted mesh, and even a separate wing for experimental stuff. Ulrich said there are 30-plus labs features behind toggles.

[02:37:23] WAN bonding that fuses multiple cheap uplinks through a $5 VPS into one resilient pipe. Per-flow SLA telemetry with tamper-evident audit chains. GeoDNS that steers traffic by live round-trip time and load application-aware quality of service. Config push to a whole fleet of remote nodes.

[02:37:49] And an AI assistant on the box that reads your actual interfaces and logs using local models. Previously, Ulrich said in a blog post that NF-sensei software comes in just five self-contained binaries that include the entire OS and the web UI. And admins are being tempted with promises that they won't be able to brick their router from the couch.

[02:38:17] Any configuration changes are stored as a candidate. Differences can be reviewed and validated through the real engines before applying them. If anything goes wrong, automated rollback will kick in if changes are not confirmed in time. Ulrich said, if a config ever fails at boot, the box falls back to the last good one on its own. NF-sensei is currently in beta with over 150 testers.

[02:38:44] So why does the world need another firewall? Ulrich argues that PFSense carries significant architectural debt. A disconnected web UI and back-end interfaces drifting out of sync. He said, if the CLI and the web UI don't speak the same language, they will eventually disagree. NF-sensei solves that by unifying both the front-end and the back-end to a single API.

[02:39:12] And developers can simply add any new features as extensions using a Lua package. No need to fork the whole project. Scott wrote that NF-sensei is the system I always wanted to build. The main challenge, OpenBSD's PF, their packet filter, a component responsible for network firewalling and traffic management,

[02:39:35] has been rebuilt as PFL, running directly on Linux's XDP, its Express Data Path, a high-performance networking feature in the Linux kernel. This essentially moves packet processing several layers deeper than other common Linux stateful firewalling implementations, improving performance.

[02:40:00] Most PFL features have parity with PF and are faster in early testing, but it's still experimental according to the engineering report. Scott said,

[02:40:31] There's no mention of when the open beta will be available to the public. In the latest blog post, Ulrich Rocks walks through potential design and branding paths. CyberNews has reached out to the developer for access to test the new firewall and will share our impressions if we manage to get our hands on it. PF Sense is currently actively maintained by NetGate as a free BSD-based firewall and router platform.

[02:40:59] It has had its own share of controversies in the past, including clashes with the OpenSense fork and a public dispute with the WireGuard team. So, at some point, we'll be getting a new firewall. Maybe don't be the first to trust it completely. Wait a while, I would say. But that's our news for the week. We're out of time. But as I said at the top of the show, we're not out of subjects.

[02:41:28] With this podcast, I think we've caught everyone up with most of the recent AI-related news, which seems to be coming at us all at once and at breakneck speed. But there are still two critically important things I need to share when we have some more time next week. The first is that paper I mentioned reading on the plane trip to Vegas.

[02:41:51] I can't stop thinking about it because it is tricky and it's going to take a deep dive into the operation of today's AI. On the other hand, I know how much our listeners appreciate a good deep dive. The other topic is some very recent research, which an AI startup and Anthropic have both written about, which hold the promise of solving the so-called dual-use dilemma,

[02:42:19] where the knowledge stored within an AI model's neural network can be used for either good or evil ends. That is the right way to solve this problem, which is not filtering, not trying to use the harness to filter what the model knows, but actually a way of governing what it knows.

[02:42:46] So anyway, as they used to say when we actually had tuners, stay tuned for more to come. Amazing. Well, Steve, once again, I tell you what, everybody listening is going, oh, I love PFSense. I can't wait to try it. I'm going to wait. I might wait. I might not be the first. It's too important. I mean, it's on your perimeter.

[02:43:10] I actually have the PFSense box in front of my system's NAT router, you know, wireless access. So it's your first line of defense. It's my first line of defense, but it's security is not critical because I have a NAT router behind it. Right. So I can probably, I'm sure I'll bring one up and see what it looks like.

[02:43:35] You know, the idea of the same guy who did PFSense 20 years ago saying, this is what I now know how to do. Well, that's funny too. Yeah. Yeah. Yeah. But it's funny too, because he says it's going to have local AI. Well, he couldn't have done that 10 years ago or two years ago. I'm not sure I want it, to be honest, but I'm sure he'll give you a switch to leave it off. But yeah, I mean, you learn, you know, that's refactoring is always better.

[02:44:03] You know, you learn and you do better the second time or probably for him, it's probably intense. You're in the process of probably re-implementing your AI on your two Spark boxes. We're almost done. Both are plugged in. Both have updated. Both have rebooted and are on SSH right now. So I'm not going to touch him. The AI is going to do the whole build. What a world. Yeah. Yeah. Wow.

[02:44:34] I'm just looking at the, yeah, it's good. So next week, a couple of really cool topics and we'll squeeze in whatever other news has transpired since then. For episode, what would that be? 192. 1092. 1092. 1092, buddy. Yep. We are getting in the upper regions now. Almost as well, we've done more podcasts than Google has fixes. How about that? But we're just barely, just barely.

[02:45:02] I just wanted to mention Robert Tappan Morris served 400 hours of community service. He was sentenced to three years of probation. His fine was $10,000, 50 plus the cost of his supervision. He did appeal, but his conviction was upheld. He did all right for himself. He went on, got a doctorate, then founded in 1995, a little thing called ViaWeb with a guy called Paul Graham. Sold it to Yahoo for 50 mil.

[02:45:31] Then started a little thing called Y Combinator in 2005. I think he's probably doing all right. He is a tenured professor at MIT, a technical advisor for Meraki. He worked with Paul Graham on a language, a Lisp dialect called ARC. That's very cool. He's done all right for us. And I imagine now it's a little bit of a badge of honor. Absolutely. He vented the first worm.

[02:45:57] As a professor, it's like, yeah, I got arrested, but you know, I was 18. I got some street cred, baby. I invented the first worm. They named it after me. No, he did very well for himself and is probably quite wealthy given that he founded a Y Combinator and sold that to Yahoo and all of that. So he's done all right. He's done all right.

[02:46:23] Ladies and gentlemen, that concludes, speaking of doing all right, that concludes this, again, wonderful episode of Security Now. Steve Gibson, the man in the myth and the legend, is at GRC.com. That's his website, the Gibson Research Corporation. You'll find many things there, including, of course, Spinrite, the world's best mass storage maintenance recovery and performance enhancing utility. I met a bunch of people at Black Hat who said, yep, I have spent...

[02:46:52] Some guy said I'd had it since the first edition. I said, that's more than 30 years. And the amazing thing is, he's been getting upgrades all this time. Current version 6.1, the most recent. You can also pick up a copy of the DNS Benchmark Pro, a great way to check your DNS server, make sure you're using the fastest one available to you. That's $9.99, both available at GRC.com. I use it too, I'm proud to say.

[02:47:20] You will also find some other things there, lots of freebies, including, of course, Shields Up, the tool every... Oh, the AIs are talking. I think they're probably telling me something about Sparky and Sparkles. What was I saying? Oh, yes, Shields Up, the best tool for testing your router. Anytime you set up a router, when you set up that new PFSense, what does he call it? PFSensei? NF. NF.

[02:47:49] NF. NF Sensei. You're definitely going to want to run it past Shields Up. I imagine it will pass with flying colors. You can also go there and sign up to get his mailing list. Actually, what you're going there to do is to whitelist your email address so that Steve gets no spam because he's very careful. But if you whitelist your address, then you can send him questions, comments, suggestions, pictures of the week. Go to GRC.com slash email for that.

[02:48:16] When you do that, though, right below it, you will see two checkboxes. There are two newsletters. One is the weekly show notes, which he sends out every Sunday. 20 plus pages of goodness. Well worth signing up for that. He also does a mailing list he never uses, which is for new products. But, you know, you might as well sign. You want to know, right? If he does put out a new product or an update to an existing one, you'll want to know. Check it out. He has copies of the show as well. In fact, he has four unique copies of the show.

[02:48:46] He has a, for no reasons no one knows, a six, actually I know, but we don't talk about it, 16 kilobit version for the bandwidth impaired. A 64 kilobit version, which sounds great, is still smaller than the one we offer. He has the show notes there, which are fantastic. And, and this is the reason for the 16 kilobit version, Elaine Ferris, very talented transcriber court reporter by trade,

[02:49:11] does a fabulous human written transcript of every show that gets up there a few days after the show goes out. That also is at grc.com. We have copies of the show at our website, our own unique versions. For some reason, 192 kilobit audio. We do have video. We got the unique video at twit.tv slash sn. There's also a YouTube channel with the video, that great place to share clips. If you want to share clips with people, a lot of people do that because Steve's always saying something.

[02:49:40] You want to show the boss, your friends, your family. And of course, the best way to get this show is to subscribe. It's a podcast. So if you subscribe and your favorite podcast client, you won't have to pay a penny, but you will get it automatically the minute it comes out. And if you're not a club twit member, you know, pay a penny or two and you, what is it? 33 cents a day. And you will get ad free versions of all the shows and a lot of extra programming to, and support the work that Steve and I and everybody at this network do.

[02:50:10] twit.tv slash club twit. Little plug there. Thank you, Steve. Have a wonderful week. It was such a pleasure seeing you in Las Vegas. Really fun. Everybody said you got to keep doing this. We will. We'll do more of those. It's just, it's so much fun. Once or twice a year, not more than that. But it's hard to get Steve out of his fortress of solitude. But we'll do our best. Thanks, Steve. Have a great week. We'll see you next time. Hey, everybody. It's Leo Laporte. You know about MacBreak Weekly, right? You don't?

[02:50:38] Oh, if you're a Macintosh fan or you just want to keep up with what's going on with Apple, this is the show for you. Every Tuesday, Andy Inaco, Alex Lindsey, Jason Snell, and I get together and talk about the week's Apple news. It's an easy subscription. Just go to your favorite podcast client and search for MacBreak Weekly or visit our website, twit.tv slash mbw. You don't want to miss a week of MacBreak Weekly. MacBreak Weekly.

AI security,Anthropic,Meta, OpenAI breakout, Hugging Face hack, chrome vulnerabilities, pfsense, NF Sensei, bruce schneier, AI genies,Cybersecurity,agentic AI, prompt injection, Simon Willison, AI bug bounty, Apple bug reports, Mythos model,