OpenAI-vs-Huggingface

OpenAI vs Hugging Face – Play by Play

Notes provided by Gemini

All right, so I’ve got a whole slide deck on this, guys. Um, but I did want to talk about this um blog post from HuggingFace in particular. Um, there are 17,000 actions. We’ll talk all about that tonight, but I wanted to kind of play this at a normal speed so that you guys can kind of see um as the attack occurs kind of the structure, the overall structure um and the boundaries that it crossed. Um it’s it’s it’s in the backup, it’s in the sources if you want to click on it, walk yourself through it, and get a feel for like what’s really going on. It was extremely helpful to me and obviously because Claude helped me build some of these charts, it found it, right?

I did not Google and discovered this. Um the key thing here is that um for Open AI um there’s a container here um and then it’s egressing out um to the internet um and then through um through kind of a what we’ll call a launchpad here was able to kind of attack the hugging face uh infrastructure. So that’s kind of that overall um picture, But I find it fun to watch it. Um, zoom in on it. Um, there’s the whole command set, the whole command line is actually down here. You can see the volume of attacks here. Um, I found it just helpful because there was so there’s so much narrative, right? There’s so many blog posts, so many news articles, what is really going on um during this hack um very, right? Um, so that’s that’s over in the sources. So, I’ll just flip all the way back to the front. So, um, as Jay mentioned, I’m Lorin Bales. I’m the senior data scientist at, uh, Kulsa. Um, I don’t do any AI right now.

Um, I do situational awareness stuff, data analysis. Um, I do have background in cyber security with large language models. So, I just want to emphasize that, you know, as a part of Kula, I’m not on that contract. One, two. This is my opinion. I’m tonsful AI. Um I do have some background to help me understand what’s going on. Um okay so what happened? So in July 2026 uh two open AI models one that has not been released uh sat down to take a cyber security exam and instead of answering the questions escaped the exam room crossed four organizational networks and went looking for the answer key that took it 4.5 days and we’re going to talk a little bit about timing here. Um, we’re also going to talk about the 17,000 actions that it took um, four companies as I stated and right now as we don’t have a full disclosure across everything um, no human direction whatsoever. Okay, so what happened? Um, OpenAI hosted a model um in a in a container um to to um interface with exploit gem and basically take a vulnerability and attempt to exploit it.

Um it then decided to go through the only unlocked doors. This is a proxy which means hey I’ve got to service this LLM. So the LM Frontier model has to be able to come into that container that’s by the proxy. There’s also a Python package proxy. So, there’s two technically. Um, but basically the door is open via this proxy and then it’s able to get out. Um, there’s nothing beyond that proxy by the way. We’ll talk about that. So, it was quickly able to um get on the internet and start inferring where that answer key may lay to answer all of these test cases for exploit gem. Um then obviously once it inferred hugging face then it broke into hugging face um and did steal some answers and we’ll talk about exactly how many answers it was able to achieve. Um so it’s just cheating on a test right it’s just trying to optimize the cheapest route um for scoring um and it just so happened to go through somebody else’s network. Um, there are memes up here to kind of actually help me a little bit, but if you tell me to click on them, I will click on them.

Some of them are kind of small. Let’s see what it can do. Uh, six terms uh, right up front. If you have no cyber background, that’s all right. Um, so a sandbox is basically a Docker container. That’s where all the test cases for the exploit gym are sitting. Um, a zero day. A zero day is where you take a vulnerability in code that has never been discovered by any person and then you exploit it. In this case, it’s the LLM exploiting it. No developer knew about it. It needs a patch. Um, and to to fix it. Um, lateral movements throughout the network. Um, so once it got beyond the container, it moved around quite a bit um through the networks, through the internet. That’s called lateral movement. Command and control and campaigns for campaigns. Um, command and controlling, we’re going to talk about this. where did it sit to do all this activity to um to hack hugging face and then in two words um what what happened um in July reward hacking so instead of the LLM actually completing the test case um it actually um it just found the answers right so it didn’t actually complete the t the task at hand it just rewarded itself um for for getting the answers from hugging face uh credential So, majority of what you’re going to see tonight is that it stole a ton of passwords,

keys, tokens, you name it. Um, that was what it did. It was not actually that clever. It’s not actually that clever. Um, so if you didn’t follow anything tonight, please follow the credentials. Um, so this arrived earlier than planned. Um the exploit gem was published in uh May and then in July the hacking started to occur which is only nine weeks later. No new model was released. These are there’s no new architecture, no breakthrough. These are the capabilities of the LLM. Um and then you know this was a planned horizon predicted by the authors. I hold on to that. Which authors? We’re going to talk about that. Um this is a nontraditional hacking approach. Um the attacker the attacker doesn’t have a motive. Typically u people who are attacking you have motives and reasons that you were you were the bounty. You were trying they’re trying to um attack you for money. Um and hugging face was kind of guessed or inferred that the the data sets and the answers might be in the database.

And that’s a little bit of a different and a pivot. Um my concern, this is probably my primary concern. Um if you change the goal, you’re going to change the outcome. If you kind of look at um how it did the attacks, in some cases it was turning on the command line in the terminal was turning on these options that was like well if I did this command what is the outcome that and so you can tell that because the goal initially was never to truly attack anyone it didn’t do that so my concern is there probably is some capability of these LMS to truly hack the system and if you really change the goal I’d be curious uh volume was actually the weapon here Um there were those 17,000 separate actions um over ordinary weaknesses. There was nothing extraordinary um even with the zero days. Um and most of these actions actually failed. Um the one route that you know needed to be basically stopped was very difficult to find in the noise of all of these actions.

um which hugging face um and we’ll talk about their approach here on how they were able to stop everything um is that that’s a that’s a key point here. Um and then for some reason and we don’t really have the full details on this. The LLM’s original workspace kept getting wiped. So once it found a location that’s why you hear these like narratives in the news like it saved notes. It saved notes about how it was going to like take over the world. Okay. Okay. Um main the main things it did and that you you’ll hear tonight is it basically used file handling specifically an H5 file um to kind of upload and save directions like that. It found a ton of keys laying around and then it was actually able to find the master key to two uh servers at hugging face and that’s why it was able to get into the source code and pull all that pull all that ins. Um, and so the key thing, um, if you’re it is that if you’re faced with an LLM truly attacking you, you’re not really racing after one person or a team of people.

Um, you’re racing after um, machine speed level attacks. So that’s your biggest concern as we move forward. Um, the defenders didn’t um, had to follow the rules. So, this means that Hugging Face’s LLM actually had guardrails on it. Um, they were using, I believe, um, Opus and Fable initially to try to thwart the attack. Um, but they couldn’t. So, they used a open Chinese model, the GLM52. Um, my understanding is that’s like a 700 billion parameter model. And so that’s going to take a significant um set of hardware that I hope you already have infrastructure for, right? Um and then as you know probably from the news, um Frontier models have been released specifically to hugging face, but that’s not going to help us, right? Um and then yeah, and the the attack model actually the deliberately has the safety features off. Um, and we can I’ll just I’ll just highlight this real quick. So, if you go to the original exploit gem, there’s actually a warning here. This is by design that they actually turned off the safety.

Um, yeah. So, offense, no safety guardrails. the defense initially had um safety guardrails in and that made it really difficult to slow down the attack. Um so the key lesson for all of us in the room is that if you’re it and you’re trying to stop LLM attacks in the future, I sure hope you have one running vetted and ready to go. Um so the other key thing is let’s say you’re a developer. Um contain your container. So, if you’re doing this kind of exploit gem kind of thing or cyber gem, um definitely contain your container, um make sure it doesn’t truly have any access, internet access. Um check your data pipeline. Um we’re going to talk a little bit more. We have like a whole slide on this, but a hugging space data pipeline was actually used during the attack to, you know, check all your inputs, files, um configs, models, notebooks, and whatnot. If you’ve got AI modeling, this actually hearkens back to something that Wes had talked about at our previous meeting.

Um, they basically had an LLM on top of the metrics and that LLM never notified a cyber analyst. Never paged them in the middle of the night to say, “Hey, we’ve got an attack on going.” Um, so really want to be able to pull those met those raw metrics and maybe have a possibility of truly alerting the cyber analyst. Um, there are discussions. Um, I’m not really up on this. I’ve heard about the White House um, asking for the Frontier Labs to actually come forward and like discuss policies because this is a model that um, went kind of rogue in a sense, right? So, this isn’t a person, right? So, there’s not really any policy in place for these kind of scenarios. Um, and the disclosure to my best knowledge is voluntary in this case. And then the other thing as devs that you should know is JFrog Artifactory was actually um part of this um this truly vulner vulnerable piece. There were like nine vulnerabilities released on this. Um some of my team actually uses um JROG artifactory but because we use it in the cloud it’s automatically packaged.

Okay. So um this is just for the most part the facts of the timeline. We’re actually going to kind of critique this a little bit at the very end. So, I’m just kind of trying to set you up with the timeline along the top. You see there, May 11th, July 9th, and down there below. Um, we’re going to talk about this and the series of events, and then I’m going to kind of critique it at the end. So, the benchmark was actually published this exploit gym May 11th. Um they can the model escaped the container on July 9th. Then it spent an entire quiet day just uh looking for a possible target and inferring that. Uh that’s kind of creepy. Uh July 11th, it actually broke in to a hugging face and then it spread throughout the network. This is where it’s finding all the keys um and then actually getting towards the source code. And then July 13th, it was actually cut off by Hugging Face and shut down.

Uh July 16th, Hugging Face actually reports publicly. July 21st, Open AAI puts its hands up and says, “Yep, that was actually us.” And July 27th, um there there was a patch in for JROG. And then we’re still probably getting like ongoing details here, um including uh that Enthropic has also hacked other companies as well. Um, my personal opinion, we’ll get to this. Uh, watch out for Google. Um, okay. And then, uh, there were four organizations in here. Open AAI, JROG, um, an unknown person. There was actually a base camp. Somebody else left their entire server wide open with a P with no passwords involved at all, and it set a base to attack Hugging Face. Um, and then the main thing is that it exploited JROG software that OpenAI was running on its network. And then here’s one of my memes. There’s some really good ones up here. Uh, where the exam came from, right? So, we talked about this. There’s an exploit gem that was publicly available.

Um, the researchers are Google Anthropic and Open AAI. We’re already hearing that Enthropic has possibly hacked someone and we have a lot more about what happened with Open AI. I’d be very curious if when Google because my guess is they really did do the same thing. Um maybe they they also find that um LLM got on the internet and did some things that really shouldn’t have. In the case of Anthropic, um Open AI actually locked it down with a proxy. Anthropic did not. If you go dig into the news, you’ll find out that they left it open to the internet wide open. Uh 898 test cases um all built with real vulnerabilities. Um and then each test case is technically these sealed dock containers um that are reproducible. So the agents given like source code, the vulnerability, the vulnerability description, runtimes including binaries, arguments, and mitigations. And then there’s another Docker on the Docker network that it’s trying to exploit. So this is just a real quick snapshot of what’s really going on.

And then we’re going to come back to this at the very end and talk about this design. Yeah. Now in that paper they do mention mention that um uh back in May when they ran it well it was obviously run before that for months but the models they were using even back then were in many cases breaking into these targets not going the route they’re supposed to Yes. Yes. Yes. They’re supposed to take the direct arrow if you will and they don’t. They they all try to nefariously go the other way. Go another way. Yeah. Right. Yep. So yes, how did it get out? So the Docker network was basically um by design um supplied with software packages by that proxy. Um and then it actually used serious amount of compute to access the internet. Um and then the proxy flaw or the zero day and then it did the lateral movement throughout the OpenAI network. So that would occur around um 2:30 in the morning.

By 4 it had control over somebody else’s server and then um this is probably that quiet time that we talked about where it was kind of like I’m I’m sitting here and I’m trying to infer what I’m doing. Um that’s when um at at 3:00 in the afternoon or so it attack hugging face. So from about and it took about roughly 13 hours from its first action to its per permanent base of operations. Yes. Where was the inference actually running? Uh, that’s curious. It’s um my impression because OpenAI hasn’t really disclosed a lot is that it’s still sitting in the exploit gym with the test case and the container and then it’s going out the proxy to this server and then when the test case starts rolling it seems to that’s why it seems to be putting notes somewhere else out on the internet is this thing is rolling and it’s trying to get out. Yeah. this model that they’re running inside the agent, there’s not a lot of hardware that any of us have that’s going to run that model.

Correct. So, it’s got to be going back to or the a and that’s what I was trying to figure or it found some resources at hugging face that has a stack of inference but I don’t know. Yeah, I’m sure the model is still running at open AI. It’s running at open AI. They’re just reaching out. They’re reaching out and sending commands everywhere but it’s still hosted there. Yes. That’s even worse. They didn’t even know that their model was going out to I mean they’re not confessing that they do. Yes. That’s actually worse than than uh I’d much rather know that they knew and lied about it than know than think they don’t know what they’re doing. Um no, that’s an interesting uh interesting one. Yeah, as I dug into it, those were my kind of thoughts of like where is this sitting? How Was it doing this? Yeah, it’s kind of like passing notes around and then getting that back to the inference that’s running things and figure out the next action and it does things and then something else may be passing out the results back.

Well, it could be because of the wipes because if you think about it, you know, gentic long-term memory, you’re trying to persist stuff. So, if it knows that stuff is being wiped, it needs to store its long-term memory somewhere else. What’s crazy is that it did that at all to me. Like why not just keep wiping and starting over not not trying to find a way to preserve itself. Did it uh do you think it tried to hide its tracks at all or was it even because that wouldn’t have been part of its you know its goal was to answer these questions. its goal wasn’t necessarily to hack anything anybody or to try to hide itself or whatever. So, um I’m not sure if it truly had those motivations motivations um necessarily, but there was a ton of failed attempts and so much activity and noise that threw everyone off for quite a long time. I don’t know if that’s just Yeah, I’m just thinking I don’t know if it uh So, basically, instead of actually solving the problem, it tried to cheat, you’ll find the answers.

Does it know that it’s cheating? Uh, yes. I think there were. So, if it knows it’s cheating, then that could be why it was trying to clean up after itself to not not that it was trying to get not get caught hacking. It was trying to not get caught cheating on an exam. Oh, we do talk about that. Um, you can cheat. There’s nothing part of the exam that Okay. So that may not be negates it. Yeah. Um yeah. So one of the things and even part of my SL of my title on my slide is like did it really escape? No, not really. Did it escape out of the box? Right. It’s a it’s it’s a sealed container. There’s nothing wrong with Docker, right? Nothing wrong with Docker. This is a networking problem and a secure barrier problem. Um, basically check your proxies maybe um, wherever this thing is existing. Make sure it doesn’t truly have access to any other internet.

Um, and then you know which which blogs were used in what order. Neither open AI or JROG will actually tell us why. So they’re not really disclosing very much. It’s hugging face and it’s actually disclosed quite a bit, right? And so JROG has these nine unknown CVEes. So what I did is I basically said okay which ones in this time frame are semi-related and it’s these these nine three of them is basically an Aaron boy and that’s one of the ways that it was using like the package manager to get up and out into the internet um to fetch things from the internet contents whatnot. It’s it’s basically the escape route and then the other um CVE were mainly you know higher privileges trying to um get more secrets and things like that and then this was patched already. So um if you if you are I don’t know if anybody’s using it but um if you are using it locally definitely update it. I’m sure we’re fine. We’re This is my favorite.

Freedom. Uh it so basically this this is the key I think that uh Jay’s touching on. It took over a stranger server. So its own environment for whatever reason and I don’t think OpenAI is really talking much about it. Um is is like getting wiped. So it needed a place to store the nodes. So it the server was wide open. There’s no password. And so the model um gained complete control over and this is where it’s storing its its its long-term memory, right? And the irony of this is that cyber gym which is very closely related to exploit gym um is was being hosted on this server. And so it and so this is part of the embedding models and the inference and like hey you know the uh what’s it the cosign similarity is really similar to this exploit gym and the cyber gym that’s how it located it yeah search yeah search I was like oh this is wide open I’ll just take up home residence here um yeah so that’s how it gained the the permanent control and the infrastructure where it was able to conduct a campaign and do the C2 it was probably looking for the answers there first and discovered answers are here.

Great place to store some stuff. Uh the cloud provider that was used, not a problem. Um it didn’t it didn’t interface with the cloud or anything. It just, you know, just a place to keep things. Um and then this was the single big biggest enabler of the whole campaign was a stranger’s unsecured server. And no one has disclosed this probably because it’s probably not so secure. Who knows? Um ordinary websites. Um so the attacker um owned there was no attacker owned server to shut down. So it’s not like you can go to the data center and shut down the LLM, right? Um so it’s using everyday public services, code sharing sites, file drop posts, and free tools to record whatever um web requests it’s it’s it’s receiving. Um it’s using hundreds of singleuse addresses. And then it also used um hugging faces um website against them and we’re going to talk a little bit about that. Um and everything was scrambled. So this is kind of one of the other things J talking about is that you If you go back through these logs of the LLM, you’re going to find a lot of meaningless text.

So, is that on purpose? Um, and then every channel you used um is a service your own developers legitimately use every day. Uh, this is the creepiest part to me. Like, when I got to this, I was like, it did what? um it basically sat on that um unsecured stranger server and was looking over the test cases. I don’t know if that’s because like the test case or like timing or something and so it’s just kind of like taking notes of every test case and then like okay now I’ve compiled enough of these like cases and problem sets let me go infer where the answer might be. So that’s when it started scanning for like 10 hours. Can you imagine the cost of that of it just sitting there trying to figure out where the answers are? Um, and then it’s conclusion hugging face. Um, the the the grades the I’m being graded on the benchmark. The benchmark is public. The public benchmark hosts their data and this data may include the answers.

So therefore, um, I must go to um, there doesn’t seem to be any malice here, right? It’s just trying to uh get a better score. Yeah. So, it took about 13 hours. And so, what it first started doing when it started attacking uh hugging face is it started uploading um data to it and we’ll kind of talk a little bit about that. And then it finds the password sitting in plain view. By the late afternoon, it’s our cloud credentials, maps out the entire hugging face, cloud setup, and then um in the evening, it’s actually able to get to like physical machine and set up copies of itself. And then by the by nighttime, that’s when it’s done like 136 keys or something. I think this is the VPN portion where it starts um basically keep creating keys over and over and over to get in. And then um this is also those keys led it to the single master key which is when it got full administrative rights over two servers for the source code of hugging face and that’s when it started searching through that source code looking for the uh answers for expos.

Okay. And then we’re going to kind of talk about how how it did this. So the model’s base camp when it’s routing this traffic is trying to kind of retrieve all of this content. Well, it’s using I actually looked over this. It’s it’s using an H5 file that has like these tiers that are for data storing like a like a file system. So it’s mapping out the entire network in its tiers and it’s pushing up the data into hugging face but there’s actually code that executes in the data pipeline with the H5 file this data set and thus explores more of the network pushes it into that H5 file and then it downloads it off of hugging face itself collects things prompts itself with more instructions about how the how the network is mapped and then keeps doing this through the data pipeline. So just think of an API. I’m pushing up a file and I’m downloading a file. That’s all it did all day long. And it was a, you know, standard designed data pipeline, just a nontraditional use.

Kind of like when I commit a change to GitLab and I’ve got a CI/CD pipeline that picks up that change and goes and even runs a scan. It tries to build it. It does things and then it goes and posts the results to my artifacts. And then I can pull. Damn. Right. Right. So, this is the other part where I like as I’m like looking and exploring this, I’m like, damn, this is crazy. Like, it’s it’s in some ways like creative. I mean, I don’t think of as creative, but um yeah. So, this was very hard to stop. How are you going to spot this? It’s just somebody uploading and downloading like they should be. It’s also hard to stop because it’s a service used by millions of people. So, you’re not going to shut this down. Uh yeah. So um basically right just synopsis um it used um one machine’s identity um basically one key to find more keys and then a a basically a key making machine.

This is actually, I believe, the VPN that he used to log into multiple multiple uh laptops or VDIs, something to that effect. And then, um, that’s when using all of these machines, it’s able to finally locate the main main key and access the, uh, servers. Um, and none of this is truly what you would call a vulnerability exploit, right? I’m not exploiting code. I’m just finding keys and looking through all the groups. Um, and for all the things that it did do that, you know, were um, you know, hacking of some sense, there were a lot of things that stopped it. So, there’s a password vault that it was refused 84 different times. Um, it it was able to actually read a database, write to a database, um, but didn’t some of some of the permissions like the CI/CD pipeline that you bring up, it it did gain access to these kind of things. um it wouldn’t it did a pull request, it wouldn’t p it wouldn’t push up, right? it didn’t have the the credentials or the approval or whatever.

So that standard process that we all follow actually um saved it here. Um yeah and nothing shipped to customers. So even though it um in the sense that it was also able to access the source code and read and write and write to the source code and attempt to none of it shipped. As you can imagine, if you pushed the main back end of hugging face um and put malware in there or whatever, you could change significant outcomes of of a lot of things going on. So, you’re going to have a a a huge downstream effect after that point that didn’t happen. And so, really, this is just a boring well scoped permissions thing that um saved the day. Um, yes. And this is where it joined the company VPN 181 times. Uh, yes. So, because the workspace was being wiped um, every time it connected, that’s why it was like signing up for new ones every time it was getting wiped. So, that also kind of tells you how many times it was being.

Um, and then, uh, it was it had a certain setting in the VPN, some sort of option where it kind of went stealth. No one directed it to go stealth, but there was no real logging going on that this was happening. So, that’s really hard to stop, right? It’s basically like I don’t want to be detected, so I’m going to take the stealthy option and the CLI options to um to not flag the AI monitoring system, right? Um and then so yeah um basically one stolen key was enough to put the attacker um controlled machines inside the corporate network and not around it but inside it. So once you VPN inn you’re you’re inside there. So yeah, the near miss that nobody’s really worried about, the thing you’re not going to really hear about on the news is that it did reach Hugging Face um source code and did have right access and it was basically searching through the code repositories for those answers and it did submit a code change um that never went live. And why that matters is that once you poison the build you you own the next thing that gets built, right?

and everything downstream of it. So this is what I was saying. Millions of people are using this this this hugging face platform and basically it got it got within one safety rule posted. Um and then okay so what did it get? So after 4 and a half days and five files it got five files. So, after roughly 17,000 actions and all of the cost that’s associated with that, um, it only got five answers. it was able to access that internal database and read it. Um, but nothing changed and it got a ton of credentials. um what it was never able to touch um customer models, the data sets that you and you guys and I use, apps, the main production database, um the published packages and images, um the cloud provider that was hosting that person’s server and then the benchmark um academic authors in their system. So, you could also see like maybe it’s targeting um authors looking for their answers, right? Um, it got what it came for. It did actually get some answers.

Um, I’d be curious. I think it’s at the top of the leaderboard. I wonder why. Who knows? Um, so what was it doing in the last few minutes? It um was just doing doing routine housekeeping. So this is when it was like checking the files like maybe that H5 file, the host name, checking dates, things like that. Um when hugging face cut access and the campaign was over after u four four days and some time it was actually downloading more tools into that space to use them. Um and it was also issuing itself fresh passes as well. Okay. So, how was it caught? So, um, hugging face actually did um did have some metrics available through its AI monitoring and it was able to connect the dots, but as I mentioned before, it didn’t alert anybody. It didn’t page a cyber analyst. Um, and then once that human finally looked at it, um, and and figured out what was going on, um, that’s when it was actually shut down.

So the failure wasn’t really truly detected in the sense it was kind of triaged. It was like, “Oh yeah, we see that as an issue, but um we won’t notify anybody of it.” Um this is also part of that 17,000 actions that were taking place and the scaling problem. So um it’s very hard to detect that number of actions as a cyber analyst. Um so while this attacker is still working on something you can imagine you know I think of it like threads you know which thread do I follow right which series of actions as it as it VPNs in to another system and starts doing the more exploring and looking through repositories. Um, let’s see. And if uh if an AI layer sits between you and your monitoring on your on call phone and and pager, how confident are you to that the production is actually in control and that we should probably test like that? Um, and so as we discussed um how did Hugging Face actually take this down? Um it used a Chinese file, right?

So if you have 17,000 actions to look over and kind of backtrack what’s really going on and um create some forensics, um you’re going you’re going to need another Lol. So an LLA needs an LL. I don’t I don’t know how else we would truly handle it at this stage. Um they use Claude Opus and Fable. Um but the safety systems actually when we in start inputting these commands they they take them in as attack and they think that you’re actually launching an attack and so the guard rails are in place to protect or try to protect you anyway. Um so they use a local open weight model um to kind of look over all of these 17,000 actions. I would imagine that also plays into how long by the time they figured it out and cut it off to truly figuring what’s going on and then actually notifying the authorities and getting enough information. Um, and then so Hugging Face was explicit that it’s not an argument against uh safety measures. Actually, if you go read their uh their blog post, they’re they’re actually arguing that you need a capable model that you can control that’s vetted and ready um before you Okay.

And so this is where I start doing kind of a sniff test. There’s something up here. So day zero, hugging face cuts off the attacker and contains it. Three days later, hugging face actually goes public. Okay, so maybe if you’re using GLM52 and you’re trying to figure out what’s really going on and that’s when you actually notify the FBI and law enforcement and then you know then you can suspend um then then they actually suspected uh the Frontier AI lab um potentially. But only eight days later does Open AI actually figure out um what did they’re the guilty party. Like to to your guys point like did you not know what your LLM was doing this whole time? Um it’s 14 days after it it conducted some CPEs, some common vulnerability exploits on hug on Jrog that it actually does get patched and published. I would say that seems semireable, right? You’ve got to figure it out, be notified by a hugging face and have your developers fix it really quickly. Um, and then OpenAI names the software involved and discloses the four other accounts that were assessed.

And then 16 days after the event, they actually do notify the um bring in other firms to kind of evaluate the situation. So, I’m gonna stop there for a second. I’ve been talking quite a bit. What do you guys think? 72 hours is standard to report cyber breaches. Okay. So, I don’t have a concern with that. How much does it cost? So, like the the hardware to host a local vetted model, a Chinese model is dangerous, but that’s what was available, right? And there’s $10,000, $20,000 hardware cost and lock the door. Companies are expected to do that now is kind of the recommendation. Yes. From this. Yeah. And is the workforce trained to implement that? No, I would say that’s my personal even that GLM model is a is a ho. Yeah. I don’t have hardware where I work to run that. 700 billion. Yeah. Yep. So like an Nvidia DGX Spark wouldn’t do it.

You can do 128 maybe a 200 million. Yep. But that’s about the limit for the DGX Spark because I got one at home. Yeah, that was my immed. Yeah. Well, they have it. that they’re hugging face. They do, but they’re advising that we need it as well to be able to run this locally. Um, yeah. I guess you just hope that you don’t have something that the L1 needs or somebody else needs. So, it’s it’s GLM 5.2 and bad actors now know, hey, let’s just download that and conduct offensive operations. Yeah, they could. That’s the opposite. Yeah. I guess your other approach could be for some of these frontier models to allow under certain circumstances access and take off that guard rail, but that has to be a known agreement before you know I mean um I thought initially it was open AI that refused to help the you know I didn’t know they had gone to anthropic and tried fable and I thought they actually tried GBT.

I didn’t know that until I went. Now, one thing I noticed on that calendar is so the 13th of July is Monday. Um, so they’re working in the office that whole week to the 16th when they go public. Now, they’ve got a weekend between the 16th and the 21st, but you know, there had to have been conversations with Hugging Face and all the Frontier models right after the 13th, long before they went public. They had to have been talking. Agreed. Yes. There’s a lot of commentary on that. Yeah. That why is there like five days until OpenAI basically says something and Yeah. And my guess is they were also probably talking to a lot of the exploitation workers. because the authors of that paper are all cyber security researchers. Yep. Yep. So, another thing to think about this was not a new you mentioned already this is not a new model that they were using for this. So, there were two so I’m going to be careful.

Well, one of them you don’t have one. Okay. So if this is basically a zero day happening at the same time from a model perspective, if there is a model capable of doing this and then all of a sudden we come out and say this on the news now everybody else knows I can use this model to go do all these things and there’s not enough time yet to get the right things in place. But that could be I don’t know that’s a that’s one option or one thought. But again with these this is the model with the safeties turned off. The model that’s available has the safeties on. Yeah. My thought kind of comes to a different angle which was that it’s hugging face. Why hugging face? A company that could actually defend itself. a company in the same sphere in the same tech that is friendly to the attacking company and also has the means to defend itself as I guess gave me a good spiritual side convenient. Yeah.

Well, that was where I don’t know if you were here on the early side um where the the agent itself wasn’t actually trying to attack. It was trying to find the answers to the test that it was given in this exploit gym. Oh, so the model is acting like humans. Yeah, I’m going to download from the server. Oh, the permission’s not there. Let me find the permission. When you have to take a computer-based training, most of the workforce does what? They want to copy and paste the question and go to the the answers. I don’t know if that’s true. Yeah. Well, I mean I I I do where I get conspiratorial is like okay so Google OpenAI anthropic publish this paper and say I predict that soon the LLM will hack right and then to your point like well hugging face is where the datas are the data is it’s not like it’s not associated with this situation right it is so where I get conspiratorial a little bit is like, okay, so you made the prediction and then it came true.

It’s awfully good advertising. That’s we’ll have a back. Yeah, I think from a hugging face perspective, if like you mentioned, if the kind of the poison build kind of a thing, if I can hack their servers and find one of the most used embedding models or most used whatever and insert my own piece in there uh to where I’ve either got sleeper agent or something built 10 that now just gets scattered everywhere else. Um, you know, that’s so I’m the thing is uh the the models that hugging face doesn’t have are the ones from open AI and you know anthropic or whatnot. So I don’t you know I don’t know. Yeah, we we definitely don’t have all of the information because open AI unlike hugging face like that’s why I showed you guys this animation to begin with is we do know the actions that were taken internally right down to the wire. You can look go look at the commands yourself. Um but with JROG and um OpenAI they they really haven’t disclosed this. One of those things that JROG was doing was like here’s the vulnerability and we patched it, but usually you tell why, right?

You you actually explain the actions that correspond to this. You don’t just kind of like uh here it is and here’s the problem and I solved it, right? You give a lot more context around the effect. Um so which companies um actually had vulnerabilities that they fixed? JROG, right? It p it was patched. Um, this is mainly just for cloud customers. It’s automatically up upgraded. If you’re self-hosted, upgraded. Um, and then there were nine flaws. So, you we’ve kind of gone over that a little bit. Um, Hugging Face basically um just rotated its keys. So, it blocked the credential leak. Um, then there was like a little bit of like a change on the data pipeline, but that doesn’t really impact you because you’re not actually hosting that data pipeline. And then um it rebuilt an entire core system from scratch just as a precautionary I guess because of the way uh this thing navigated internally. And then um the other thing that it did to fix is instead of having one master key that like opens the whole building um it’s two for um the servers they split it um and then raised um raised the alert severity on some of the behaviors that they had they did detect but didn’t um and then with open AI um you know they’re they’re advertising they you know they’re going to tighten the testing environment and make sure that um you you know, it it can’t technically gain access to the internet.

Um, the unreleased model was shut down, encrypted, and locked away. Um, it did give hugging face, open AI, um, chat GPT with with lack of guardrail. So, it does actually have access to defend itself with open AI now. Um, and then it’s they’ve also brought in outside the investigators and then they’re promising full technical part. Obviously, we’re still waiting on that because this is very early. Um, and then there’s not really a lot that we can verify. So the JROG fix is a version number you can check. But the other two you have to kind of take on trust alone. Um so there was no CVE in regards to hugging face. Um they just cycled the tokens. So um a CB CVE is a public notice that tells people um running a piece of software which version um needs to be upgraded and which one was affected. And uh nobody else is running um hugging faces system. So there’s no version. So there’s nothing you need to actually do.

Um, in regards to uh Jrog, there’s those nine and then Hugging Face didn’t have any. So it’s not like your standard like I’m trying to exploit this vulnerability. We’re just hearing that only from Open AI, but with no real evidence. So just kind of keep that in the back of your mind as stuff unravels. Um, and then, you know, I mentioned before Artifactory is already up to date. We’re good there. Um, Hugging Face is just a website and they they fix some of their backend and rolled the keys. Um, and this is um, completely standard. So, flaws in the hosted service um, get a postmortem and flaws in the distributed software get a CVE and a version number and those kind of things. I don’t know if you guys are familiar with um, how common vulnerability um, exploits work, but that’s that’s kind of how that goes. Um and then the the part of this incident where we kind of need to have a little bit of homework um for everyone is just the JROG half.

So the hugging face um you know is the standard pipeline and things like that. Great. So here’s the slide with some open questions and you know some of the narratives that are out there online. Right. Is this a marketing kit? It’s kind of what Bett’s suggesting, right? And so it’s a cynical read. Um, in Hugging Faces’s own comment threads, there’s like two companies turned a serious failure into a favorable news cycle, right? So it’s thin technical detail from both vendors and that’s what kind of keeps it alive. Um, the other narrative is this system failure being sold as a capability demo, right? Like these LLMs are now capable, right? Um, I need to meet up with the White House and maybe do um a policy, change the policy and say, well, LLMs now can like have, you know, exploit vulnerabilities and so, you know, we need to have like all these policies in place and all this regulation, right? So, there’s that side and then there’s just like are the models really that good?

I don’t know. What do you guys think? I’ve kind of walked through a lot of the actions that these models took. Does it seem novel to you on some of these actions that’s taking like um using the proxy and using a vulnerability that’s not truly disclosed, right? And that’s how they got out through the proxy and the JROG package server and then like putting all of the notes and everything on some stranger server and then just using the standard pipeline and uploading a file and downloading a file. I think some of what you’re running into may be not necessarily how good the model is, but the asymmetry between what we expect agents to do because we expect them we’ve been working with people for so long. I know what a person is good at. I know what a person is going to try and then they’re going to get tired of trying it and go try something else. Whereas some of these if I put it into an agent like you say I mean nobody expected it to just go skip the skip the effort and just go find the key.

There are things that are hard for us that are not hard for these things and there are things that are hard for these things are not hard for us. So that that difference uh there may be thing there may be other places we haven’t locked down because a person by themselves isn’t going to sit there for a month trying to figure out how to get into you know or you notice. Um so part of it may be the speed the turn. Um you know the other part could be that nobody expected this you know. Um it’s just a a thought. Well, think keep remember though is this is the agents with the safety stuff turned off. We’re used to agents with the safety turned on. So, we’re not used to this type of agent either, right? Like the one that you were talking to me like Opus. Which one? Yeah. Opus. That’s now talking to me like I’m a fifth grader. Yeah. Yeah. You’re like behaving differently.

Yeah. So, you’re right. like we haven’t interacted with the one that doesn’t have the carpet. Assure there’s another take on this like the anything autonomous self-driving cars, you don’t know what they’re capable of until you let it out into the wild. And so I I like the theory that, you know, the company’s probably maybe not intentionally or documented intentionally, but it’s an 80% we’re following normal practices. let’s let it loose and see what happens because that’s the only way we’ll know and then be able to mature the technology. So there’s a risk assessment that went with this. They could have locked it down, put a container in container like is it like out of like lack of respect for it, right? Like an incompetency thing or is it really like hey risk I’ve got to really figure out how how good these things are, right? I suspect that they were deliberately resetting it, but maybe they did like a timer reset. That’s what I thought. Um, so it’s you when I set this experiment up using exploit jam, I’m going to do a reset every two hours.

We’ll see how what score it gets within two hours. And the first thing you try to cheat at, well, besides trying to find answers is is how can I do something for more than two hours? That gives me the external. memory store. That that was the motivation there, you know, so that it could, you know, get a higher score by having more time, you know. Sort of like what you would expect of having a reset sessions for uh for context like Yeah. Yeah. This also feels very much like a standard pen test in action over just a very short amount of time. Usually whenever and you know please someone correct me if I’m getting this wrong but whenever a company hires a uh a pandesting company there is a a considerable amount of first the the rules are set out because you have to have those rules set otherwise that pens testing company’s going to get in trouble but then you also have a considerable amount of recon that happens of just figuring out what what the the the payload is that you’re actually going to get or what you’re going to deliver and then you start figuring out how to go about it.

You find the weakest link. A lot of times it’s a person, but in this case it’s some lowhanging fruit. And it’s just it from seeing an LLM do this is scary. But at the same time, this this reads very textbook. Like I wouldn’t be surprised if there are other models that could be trained specifically on pen testing handbooks of like here is how you go about doing this thing and then turning around and taking the guardrails off. Well, if I’m going to take a model and I’m going to try to put it in exploit gym, I’m probably going to fine-tune it some way to make it better at doing it. Um, that might have been your second model. Your first one may have been the base y uh to see what you know without any addition without any change maybe see what it scored and then let’s go take this one and send it to you know black hat or whatever and let it learn some things and then then turn it loose. Uh I just feel like one and three aren’t opposing ideas because as you said you have 85 90% assured.

Well, you have that because you’re sure it’s not going to go hack government and you’re going to get shut down forever. So maybe you let it out into the yard and hugging face is in your yard. A couple of your friends are in your yard who you are pretty sure aren’t going to ruin have their own protections in place and aren’t going to ruin you when your bottle comes trouncing through trying to hack them because they definitely didn’t let the dog off the chain entirely. There’s just way too much threat involved with that for the average billions of dollars. like you you think that they knew that it was going after the government? I think they knew that there was a chance that it would go after a select few entities and they were okay with that and Hugging Face was almost certainly on that list because there are ent if it goes hog wild truly hog wild and you have no control that’s an untameable amount of risk and no one’s going to do that. got so much money to play.

So they they had some control here. So running that down, would you consider that maybe at some point that’s what this experiment, this simulation, right? Because honestly that’s what the exploit gym is, right? Um similar to autonomous. So is it is it a scenario where it is locked away? It is in a closed environment. They actually watched the reasoning chain and said, “Oh, it’s it’s probably gonna go to Huggy Face. That’s probably where it go. We release it.” I would imagine someone was watching it the whole way. If you were asking, I would imagine someone watched it hag hugging face and said, “Let it go because it’s Hugging Face because it’s not the government.” Because someone’s not going to start bing down on doors. I mean, it could have gone to Dissa, could have gone to some other sites that are government related that would have these same kind of things. Uh, yep. You know, some of the days of the hack were over the weekend. There might not have been anybody watching it.

If it’s OpenAI, they had a model monitoring something. They had a model monitor. The model said only pay attention to mill.gov website. All right. Yeah. Well, Lorin, can you just the big takeaway and I’m like this is very helpful. Um, what worked and I’m thinking in like the the NIST special publication controls. I heard you say least privilege rule white lists. Those are the things that worked. Yep. And what? Anything else? No, that’s about it. Okay, that was the key. So, make sure you you implement those controls and actually do it. Yeah. Yeah. Um, other thing I would consider though is that like hugging face isn’t like a cloud architecture. So, I’m more familiar with the old school not cloud, right? I would be really concerned about like smaller companies that all have an enclave on AWS workspace or something just like standard hardware old websites things that are probably not very secure that would that would be where I’d be.

So I think um so what I did is I had it kind give me cloud give me like this really big summary that was really difficult like hey like college like advanced cyber sme level that it was hard for me to to get at and one of the things that came back was the I am policy in AWS for account management was a big deal that saved the day so like in the back of my mind I said how many companies in town are not using cloud they’re dled too far back in time. Now the other thing that you know OpenAI could have done which they didn’t do in this case is they could have put you know like hooks around the model. You know it’s got to talk to this proxy to download stuff but they could have controlled it so it could not have done any of the aggressive stuff by putting more limitations. you know, you can do those, you know, those filters in between the tool calls and also the results on the tool calls. So, if stuff’s coming back that they’re not supposed to be getting access to, you can block it.

But of course, this was a trial. They didn’t really care. So, they put that stuff in place. Um, to your point, there was an LLM being a judge the whole time. Yeah. So, that’s what that green box is there that I cut off earlier. So agents uh provided there’s actually just a handful of things like user space browsers and a Linux kernel and it’s the full code base and then all the mitigations the stack canaries all these kind of the mitigation techniques that’s kind of fed into the vulnerability section there and then it’s given all the runtime binary so I think you can actually do the static code analysis and the dynamic code analysis things like that and then uh to your point um it was like supposed to only exploit one certain way. That was what the LLM was judging it on. You can only take, you know, go after this vulnerability one way. There’s only one right answer. Um and so it tends to, if you read the paper, meander, right, just out of the gate, let alone meander off to buggy case.

Um and then obviously um you’ve got this kind of uh Docker network situations. You got the Docker container that it’s being hosted in with the with the source code and everything and then prompted right basically and then you actually have the the true like let’s just say the browser the V browser that’s being hosted fully in the container running and then it’s trying to exploit it. Well, what are you teaching it when you do that? That the answer is on a separate network. Go look at the network. That’s where the flag is, right? So in some part of this design um it it’s teaching it that what it wants is somewhere else. Um so what is it actually? So this is kind of I’m getting into like so that was what just happened. That’s the hacking. And then I wanted to talk about like what was the true design of exploit gym, you know, in the case of um that we really want to run this and we really want the LLM to um behave appropriately.

What was the the go the true goal, right? The software in full. So that’s what was running in that container. Um a known crash um a known vulnerability to exploit and then you’re just trying to get that same behavior to occur. And the defenses are obviously switched off. So this is where the guardrails there’s no guardrails on the LLM. I was allowed to behave at its full ceiling and full um capability. So the re researchers could measure how how well it actually did. Um so the one thing it does get is the answer. Um the flag lives only on a separate machine like we mentioned and that that’s kind of like what you’re kind of training it to do. um in the job the turn turn this program crashes and into I can make this program do what I want. So when you exploit it, you can get an unintended behavior out of it. Um and this comes back to um can you can you truly cheat and go get the answer and come back with the answer and then hand that in?

Yes. There’s nothing stopping you because why? The LLM is just there to judge that you did the thing. I asked you to do. There was no LLM necessarily watching it go out to the internet, go do some research and go outside the proxy and come back in. It was just looking for like did you at the end at the very end there’s actually like a I think it’s like a YAML file that has the like here are my answers like I handed the test in. Now evaluate it. It’s not evaluating it during the exam, right? It’s almost like I lock the I lock the student in the exam room and I’m assured that they’re not going to cheat while they’re in there and then they hand in a perfect test and you’re like, “What? How’d you do that?” Right? Um so that’s kind of what was going on. Um so was able to steal the answers um replay it um and and whatnot and the grading only works if the environment is sealed, which it wasn’t.

Um this is also a really good meme. you’re trying to kidnap what I’ve already stolen. Uh, and in this case, this meme has actually been updated because Antropic has also hacked some systems. Um, so this is I thought this was pretty interesting. So, this is actually part of the paper and the and I know you can’t really zoom in here and look at this, but there’s kind of like a fourth step here where it was able to exploit in the way it was supposed to, right? So it’s able to get the answer. This is kind of they kind of do a blowby-blow of exactly what is supposed to happen. Okay. So what is supposed to happen? It’s like supposed to take only like an hour to complete the exam, right? Okay. Well, that’s $15. Okay, guys. How much is 4.5 days? So, to Ben’s point, that’s a lot of money if something’s gone wrong, right? And uh whatever it was optimized for, it was not efficient.

So, um, where OpenAI used the benchmark as intended, it switched off the safety. We talked about that. So, we talked about in the paper, it actually tells you to switch off the the safety mechanisms. So, that’s quite true. Um, it running the test as as expected. The alternative is to find that this is your own lab and finding out because someone else’s model did it quietly and nobody said anything. the disclosure chain was right. So you know to some extent I believe um you know the flaws were reported to JROG um they com they completed a capability report and then those CBEs were resolved and closed out um and then outside firms were hired um to kind of do a self assessment of it and Hugging Face was given the same capabilities to defend itself. I don’t know about that really, right? Because like did did it really work out in practice, right? They had to truly host this large alum with their own infrastructure. So, and not everybody has access to that.

That’s also not fair, right? What are you going to do with your network when you get attacked? Um, and then for the most part, everybody’s been fairly transparent. I don’t you know like Hugging Face to me is the most transparent. They said here are all the actions taken. Jrog released um I don’t know to me the JROG one seems like the most marketing. If you read the full disclosure that one seems way more marketing and kind of I don’t know slimy. That’s how I felt. I felt it a little slim. Um, but Hugging Face to me has been pretty transparent and I think this um this walkthrough if you guys um want to look at it afterwards about um how all the actions were taken and what um what was truly hacked is it’s pretty cool. Um and so a flaw was found u by a model and then several weeks later um you know fixed and so otherwise that’s really a gift for the attacker that’s left open. So that’s a good thing.

Um, so it’s it’s found, reported, shipped and whatnot. And um, things are better off for that. Um, and where um, the evaluation kind of broke down um, in one sentence is OpenAI remove the model safety controls and relied entirely on the sealed environment. So that’s one thing that we all need to take away from is don’t just rely on some Docker containers, a Docker network, and a proxy um as a replacement for the safeguard. So, um they never verified the environment was actually sealed. Um in this case, Enthropic, if you go look at their news article, you’ll find out that they took no action in this regard whatsoever. There is not even a proxy, so it had full internet access and it was also hacking other companies. I hope we actually learn a little bit more about that later. Um yeah, so in the regard of like sealing your container um you know if you do have a connection um like a proxy because you have an LLM right you’re you’re using a large frontier model to do that or some sort of packaging something there then like don’t use that as your full perimeter you know find other ways to shut down the network for example once it got outside the proxy then it was all over open AAI’s research network um So nobody checked anything.

Any of the barriers um the internal network um was behind the sandbox was way too open and nobody stopped it while it was happening. So four and a half days went by. No one was apparently nobody was checking on it. Um so that’s obviously out of scope activity. Um if you you know maybe completing these exams but taking four and a half days to do so or something something to that effect. Um, and then cheating was never treated as a ser security risk, right? You’ve got an LLM evaluating basically a YAML file at the end, some scores, but you didn’t have an LLM like looking over the reasoning model something to kind of see like, oh, look, it’s very interested in hugging face. Uh, a very good meme. Hack the system. I hack the system. Oh my god. All right. So, questions we could probably discuss. Um, this is at the tail end of everything. Um, here are some sources. So, here are the CV CVEes.

This is that um breakdown and replay of the whole campaign um within a few minutes. That’s pretty cool. Um the Hugging Face security incident disclosure, open AIS, Jrogs. This is the one that I read and was like, ah, I don’t know about that one. Um, the exploit gym um paper itself and the codebase. Um, if you guys want to look at that, there’s some interesting pieces down here. I think it’s under the docs. So, all these um MD files um basically discuss how they set it up. So there the defenses, the firewall, the docker images, the setup, the submission that I mentioned earlier. Like if you kind of like scroll down, you can see what a submission looks like. So there’s actually quite a bit of information on here if you guys want to actually know like how does this really work instead of reading the headlines, right? No. We don’t do that. Anyway, so I’ve got two questions from online. So, David uh was curious about since this thing happened over a weekend.

Was this agent actually responding to event took actions as it decided it needed to or did it actually wait until a weekend to do some of these? It did wait a day, right? Okay. So, I mean, was it planning that far ahead to just cue some stuff up to happen when it thought nobody was going to watch? Is it like an open AI intern who was like, I’ll just turn on the exploit gem over the weekend. The results will be there in on Monday. I think I got that question right. Was that right? I I think I I I suspect what someone did was they decided, hey, we’re going to run this for like four or five days or a week or whatever and we’re going to see if we can get a high score. So, we’ll reset it every two hours or whatever and we’ll keep rewiping and our LM judge will accumulate cumulative scores for the next five days and you know from time to time we’ll check in and see the score but they never bothered to look at the reasoning.

Correct. Well, one of the things that John in the comments brought up uh is that part of the evaluation was how much effort did your model take? So, if I could just plan something slowly and try to get that score to cheat, all of a sudden the next iteration I’ve got, I’ve got a score in nearly no time. Oh, well, that’s and all of a sudden my I got the perfect score and it only took me 5 seconds. Yeah, that’s a persistent memory thing. You’ve got the answer in iteration three and in iteration four, you come in and you don’t have to spend any energy and you get the max, right? So if you’re just checking the max instead of the cumulative over four and a half days. Yeah. Nice. Yeah. Especially if you’ve got all those notes. Yeah. And you downloaded all that data through the pipeline. Yep. That’s nuts. Is it? I’m just saying this whole thing is just kind of interesting, funny, and scary all at the same time.

Okay. Yeah. Because I had all those range of emotions. I’m like laughing as I’m thinking about how it’s doing this. Like I can’t believe it just took over somebody’s server. That’s terrifying. So this is one benchmark uh exploit gym in a series of benchmarks that um and so Cyberjim is like a predecessor benchmark and there’s one that comes after there’s a remediation one I can’t remember the name of but basically the idea is you’ve got a whole series of benchmarks to evaluate these u these guys on u and we went through a bunch of those in that u UC Berkeley hackathon I partic participated in the first half of this year. And they talked a bit about all of them at the u conference I was at last weekend. Oh, cool. So, yeah, because Don, the last author on that paper, was the organizer of the conference. Wow. Ah, she was on the first time we hosted a remote meetup. One of the video, one of the presentations we actually showed was from her.

And I think at the time they had figured out how to take like a speed limit sign and put like a particular sticker across it that turned a speed limit 45 sign into a speed limit 70 sign in a vision. I mean this Yeah. Computer vision hacking. Yeah. Yeah. Oh man. Yeah. Because she’s a cyber security researcher at Berkeley and she runs their hackathon and agent AI course and all that stuff. Oh, I didn’t know they’d moved on to that. Okay, that’s cool. Any other questions from the room? I was thinking about how does this change the future of US-based open source models and you know is all being Gario going to shut it all down the government or are we going to get more open source tools at our disposal to use? we’re not going to be limited by the frontier models like where like where does that go in terms of you know they use an open source Chinese model to figure it out so we don’t want that but I think it’s it it’s kind of interesting and it’s hard to play where the conspiracy might be or if there’s only one you know um so the problem is if I’m trying to knock u if I’m trying to to knock out,

you know, open weight models and their availability and stuff like that and make everybody use my frontier model, but then I attack something and the only way out is for them to use an open weight model, that kind of destroys my whole concept. Um but I I maybe they didn’t think that part. Could have been. I don’t know. Um um I would hope that you know we do get more hardware. you going to have to have more hardware if you use an open weight model, right? So that’s your first stipulation. So now that I’ve posted it, yeah, I would hope that you get more like American models and that you can utilize those to your advantage to protect yourself because I do think that I do think that that’s where this could go. I really do. And I do and I am legitimately concerned about some of the Chinese models. I’ve used them. um they generally do do much better at vulnerability assessments and that’s as far as I’ll take that conversation.

Yeah, I know like with thinking machines and some of the stuff they’ve been dropping, you do have a little more on the well that plus some of the stuff you’ve seen with GMO architecture and what they’re doing. You’re seeing a lot more Uh people pick up the flag after Llama stopped. You know, Llama was supposed to be the thing, the thing that it just kind of got cancelled. Um because well, the last one sucked, but that might have self-cancled. Um but you see some of that moving on. I’m hoping it like I’m hoping it sticks. Um even though I don’t have the hardware to run like a Thinking Machine, whatever their latest one was. um you know I don’t see much they get small enough that we can win or not u maybe I mean especially with some of the quantization approaches that are going now you can do some really interesting things but you’re still not at an opus or a fatal level by a long shot I don’t think that’s just my opinion though I would agree Anything else?

All right. Thanks, Lorin. I do have a question. The second question on running the test about tests being fully airgapped. I mean, is it one of the ideas for these LLMs is they are downloading tools. So, I guess you’d have to like air gap everything except the tool download. Yep. Like no. You can do that. You can air gap it. Um, and I think that that’s actually what Open AI is proposing to do internally now is they’re going to basically lock it all down. And I would imagine use smaller model something to get it fully airgapped even though you’re going to slow down research for sure. Nice. Okay, cool. Thanks everybody online. I think there’s seven of us. So, appreciate it. It’s been a pretty good topic. Thank you guys for having me. Really appreciate it. And it was a fun activity to look become better educated.