This post has been bugging me all day. And much of her content for a while now. She’s not wrong, but also she doesn’t actually provide any actionable solutions in her publications. Just “do the math,” which, what?
It’s also the case that you can do full adversarial analysis of a model with access to the model. If you are tasked, as I am, with testing on trained commercial models and applications that present them, what are you supposed to do?
Full post recreated here:
The Problem With AI Red Teaming:
It’s Fake.
Real red teamers operate in a variety of challenging and changing circumstances, with interdisciplinary techniques.
We operate in (often) dangerous situations. Law enforcement may or may not be aware of our presence.
As a red teamer, I’ve had to scale fences, jump flights of stairs, and more. I’ve been threatened with arrest. I’ve been shot at.
Why did I do it? Because the mission mattered.
TLDR: We’ll come break your computer on-site.
AI “red teaming” is none of that.
Just like “prompt engineering” is cosplay for people who want to pretend to be engineers, “AI red teaming” is for people who want to cosplay as hackers.
How do I know?
If they were real hackers, then they would’ve read the manual.
You know, RTFM? But they didn’t.
There’s a reason these people only started “attacking AI” after the creation of GenAI, and it isn’t because of the economy or the use cases.
It’s because you could feel like you were “hacking AI” just by using natural language–no math required.
Or so they thought, because again–they didn’t do the reading.
AI attacks have always been math. AI existed long before GenAI and so did the security vectors.
As an example, my paper from 2022 predates ChatGPT, while still calling out the exact malicious applications we’ve seen it used for.
How did I do this? It wasn’t psychic powers; it’s because I learned the math.
And the math of these systems doesn’t change. It didn’t change with GenAI, and it won’t change with any other statistical “reasoning” system.
All effective AI attacks in the wild–you know, REAL hackers–are math-based. Again, this hasn’t changed with GenAI.
And if these “red teamers” were real hackers, they’d know that. Sorrynotsorry.
A lot of these “red teamers” are just shooting natural language prompts and calling it an test. GTFO.
The REAL kicker: Before I started talking about this publicly, I approached MULTIPLE of these “AI red teams” & their CEOs to tell them their methodology was critically flawed.
Know what they told me?
“We don’t care, as long as people pay for it.”
THAT’s what you’re buying when you hire these people.
Fake service, from fake hackers. But the money lost is very real.
And my paper explains why.