Anthropic's Mythos Model Capabilities

This is the report from Anthropic regarding its new “Mythos” model, which is used as part of Project Glasswing. Basically it’s a frontier model that, according them, has considerable improvements in the realm of cybersecurity. You can read the entire system card to learn about its training and benchmarks.

The industry sucks, no question. But it is also true these things are finding real vulns. That is as true for attackers as defenders.

4 Likes

I’ve been thinking about this model and their findings over the past day, and this does feel a little different to me.

I have kept away from relying on AI for anything load-bearing, but perhaps this (or other similar models) will push me into the realm of actively employing some AI measures.

Coming from the AppSec side of things, my biggest worry with shoving AI into everything as a decision maker is consistency. Does the AI think these same 10 vulnerabilities are the top priority every single time if you run this same query 1,000 attempts?

Perhaps that thinking was a bit naive, and I really do need to re-scope any potential impact to work in AppSec, and how the industry responds going forward

2 Likes

Lots of valid perspectives here. I recently wrote a blog post about trying Claude Code as a skeptic that kind of escaped containment.

Relevant quote:

The CVEs [the models are] finding are real. They can perform some kinds of static analysis well—and with some agentic pipelines, dynamic analysis is possible as well. They’re not doing anything novel, but the speed and thoroughness possible can improve an application’s security. The trick is deciding what to give to the model, what to lock to deterministic automation, and what to keep in the hands of human experts.

Now that I know this, is it reasonable or responsible to omit such a step in a development pipeline? I’m not so sure.

This is the part where the “It doesn’t work!” argument against the models falls apart. It turns out, for this very specific use case, it works pretty well. That doesn’t remove the negative externalities or the shortcomings elsewhere. But damned if the thing can’t find real bugs and vulns.

Someone on Masto said something like “If the trains run on time, is fascism worth it?” which is pointlessly reductive imo. But if we begin from a first principle that the use of these models is unethical regardless of benefit, the benefits don’t matter regardless of scale.

But that’s a rather academic conversation. The models exist. The bad guys sure as hell are going to use them to find vulnerabilities faster. Should developers and defenders reject the same tool on ethical grounds? If so, we’re accepting a significant asymmetry of capability between the adversary and us.

4 Likes

I worked in model development. It was data science, but I’m pretty familiar with the business around designing/working with models.

The developers who reject these tools aren’t saving themselves any headaches. AI is here to stay and will continue to be a threat/benefit/eco-disaster depending on who you ask. The erosion already happened and at this point I believe the only real way forward is to consider these tools when developing. Not to sound blasé, but toss it on the flaming pile of requirements.

Maybe I’m wrong, and if so I’m happy to learn, but “inevitably” of AI was being business driven. It was forced into products, meetings, and every conversation. It exists and pretending it doesn’t isn’t helpful.

2 Likes

The erosion already happened and at this point I believe the only real way forward is to consider these tools when developing. Not to sound blasé, but toss it on the flaming pile of requirements.

This is a good point. Looking at the state of JS development with NPM, and all of the externalities there, and being a part of the Development world as that was really just getting going provides a similar perspective, in retrospect.

I was cranky about all the third-party dependencies, and a decade later they’re still the compelling part of developing in a modern language ecosystem. Definitely going to have some further thinks on this

1 Like

Honestly? It depends on what the temperature is set to and which service’s LLM you’re talking about. It’s not meaningful to generalize without additional details.

1 Like

Unless there can be found another way of nullifying this asymmetry.

1 Like

It’s helpful in specific ways, under specific circumstances, but it is being sold as helpful everywhere all at once. I think this is a great read on this:

Also, this is another thing that’s been living rent-free in my head:

Specifically:

Someone still has to reread, compare, test, contextualize, and sometimes rewrite. And if no one seriously takes on that work, the cost does not disappear. It reappears later in the form of errors, urgent fixes, loss of trust, and eventually litigation. What is presented as a productivity gain is often just an accounting displacement. We save at the beginning on production, only to spend later on control, or on the consequences of not having exercised it.

4 Likes

+1 for the Great Leap Forward article. Excellent context and well-written.

2 Likes

Absolutely spot on, and I was reading that article earlier and it seems very spot on. We keep finding “problems” for the “solutions” companies are designing and insisting we need.

3 Likes

Thought this was a considered saucer for the cup.

2 Likes

This is sort of an eye-opening post as a HARDCORE anti-”AI” person myself.

I don’t want to make myself even more unemployable (been laid off since October…) by ignoring something for very good reasons, but ruining my future potential.

1 Like

I hope you find a role somewhere. I’m looking for what to do next as well.

Not that I have advice, but the way I see it, this is just the new world order. The world exists in cycles and we all grow/etc. New tech replaces older tech when it’s truly better… or when enough has been sacrificed to it. The smart people are rarely truly in charge. That’s an absolute around technology that can’t be ignored.

Facebook isn’t a good alternative to anyone, but it still has features Grandma prefers and she doesn’t care about learning new stuff. So Facebook still controls market share…so businesses swallow their pride and market towards them. Swap out Facebook for whichever tool you’d prefer.

AI had some big money being pushed into it. It was successfully able to permeate the world so fully that business was promoting it. YouTube video of the ad. | Common performing a lecture about AI on prime time television. Trying their best to explain it to as many people as possible. Not the technology, mind you, just the buzzwords.

This feels like it was all cultivated to guarantee that the tech jargon would be understood enough for everyone to have an opinion. I’m not prone to conspiracy and I don’t think it was some guys dressed in cowls and trying to control the world. I think it’s a bunch of money hungry business types that understood how pliable certain markets could become. Then they plied them.

So what can we do

At this point, I think it’s best to understand the systems and the people behind them. Similar to the disruptions of the past, the rivers still work they just have to find new workflows. So AI surges, AI hype surges, and then hopefully it rubber bands to something normalish? But I doubt it just dies. I don’t think it can die, not while we have hungry power wells that can feed off of it.

So we toss them into the requirements pile along with the other standards and practices. We develop checks in process that look for and terminate bad AI stuff. Data groups focus harder on scrubbing the horrible data sets that are becoming worse, or we figure out new data sets that have shielding against AI. Like a filter against a weird cyberpunk algae in the water.

Personally? I’m becoming a tad suspicious of every dependency…which is probably good?

Even in my scribbling and silly projects, I’m becoming more cautious about which packages I bring in. In software usage, I’m “checking the label” more. So… probably good?

I don’t know…

2 Likes

My take is still that at some point the bubble will pop and the costs of running these humongous models will become untenable once the VC silly money stops subsidizing it.

At that point, we’ll see what actually happens, and what actually has value and what was just an artifact of the Gartner Hype Cycle’s peak of inflated expectations.

Meanwhile, I identified my core competency which is helping organizations pay down technical debt. I doubt I will go hungry once chickens come home to roost.

2 Likes

Relatedly, I linked this today and I think it’s an important read for all companies going hard on these things (so…all companies).

The bubble will pop, and it’s the toxic debt that will do the deed. I’m very concerned with how much worse it could be than 2008.

2 Likes

As usual, Joe Slowik has a great take.

1 Like

Also very relevant:

We took the specific vulnerabilities Anthropic showcases in their announcement, isolated the relevant code, and ran them through small, cheap, open-weights models. Those models recovered much of the same analysis. Eight out of eight models detected Mythos’s flagship FreeBSD exploit, including one with only 3.6 billion active parameters costing $0.11 per million tokens. A 5.1B-active open model recovered the core chain of the 27-year-old OpenBSD bug.

And on a basic security reasoning task, small open models outperformed most frontier models from every major lab. The capability rankings reshuffled completely across tasks.