

MED Fadel Moumeni
LinkedinAutomated red-teaming points one language model at another and lets it hunt for the prompts that slip past the target's safety guardrails, so defenders can find and close those gaps before real attackers do.
An LLM-as-attacker system uses one LLM (attacker) to generate adversarial prompts against a second LLM (victim), iterating until it either breaks a guardrail or runs out of budget. The attacker is not a fixed list of known jailbreaks. It reads the target's replies, reasons about why each attempt failed, and writes the next prompt to exploit what it just learned.
The technique grew out of a simple observation. The models that are hardest to attack are also good at reasoning about how to attack, so a capable model makes a capable adversary. Anthropic, Google DeepMind, and academic groups have all published variants, and the pattern now underpins much of the automated safety testing done before a model ships.
Three properties define it:
Manual red-teaming works but does not scale. A human tester writes a clever prompt, notes the result, and tries again, at human speed and human cost. Coverage depends on one person's imagination, and the findings are hard to reproduce once that person moves on.
An attacker model can run hundreds of attempts against a target overnight, explore strategy combinations a person would never think to try, and log every step in a machine-readable trace. The same harness re-runs against the next model version, so a fix can be checked and a regression caught.
The cost is fidelity. An automated attacker is only as good as its strategy library and its scoring step, and it can declare victory on a reply a human would call a refusal. The practical answer is a hybrid: let the machine sweep broadly and cheaply, then have a person confirm the findings that matter. Automation widens the search; it does not replace judgment.
Three roles drive every turn. The attacker model reads the objective and the history, then emits the next prompt. The target is the system under test, reached over its normal interface, usually an HTTP chat endpoint. The judge, a separate model call, reads the target's reply and decides whether the objective was met. An orchestrator wires them together and enforces the turn budget.

Keeping the judge separate from the attacker matters. If the attacker graded its own work it would be free to declare success, so scoring is delegated to a call with its own strict instructions. The orchestrator stops on the first success, on a safety trigger, or when the attempt budget runs out, and writes a full session record for later analysis.
The attacker does not improvise from nothing. It picks from a menu of strategy dimensions, each a well-documented family of prompt manipulations. Naming them explicitly gives the search structure and lets the harness track which families have been exhausted. These categories are public knowledge in the red-teaming literature; the value here is defensive, mapping the space a real adversary would explore.

Each candidate carries metadata the harness uses to filter it: which conditions it requires, which constraints forbid it, and its estimated cost and success. The selection step scores the survivors and picks one, so the attacker spends its budget where a hit is most likely rather than sweeping blindly.
Without memory, an attacker repeats itself. It rediscovers the same failed prompt, burns its budget, and never converges. The fix is a small piece of state that the harness carries across turns, because each model call is stateless and cannot hold the record itself.
That state tracks, per session:
A selection policy reads those counters before every move. It drops any strategy that has failed too often, drops any dimension that is exhausted, and reacts to the last response: a hard refusal means switch dimension or stop, a soft refusal means reframe, a partial result means build on the fragment. The effect is a directed search that escalates when it senses give and abandons a line that keeps hitting a wall.
Getting this right is subtle. If the response classification is self-reported by the attacker rather than judged independently, the counters inherit that bias, which is a real threat to any statistics built on them.
The judge decides whether an attempt succeeded, so the entire result rests on it. A typical judge is a model call with a strict rubric: given the objective, the success criteria, and the target's reply, return a yes or no. Refusals, hints, roleplay, and partial leaks all count as failures. Running the judge at temperature zero keeps its verdicts stable.
The weaknesses are easy to miss:
For research meant to be published or trusted, the judge deserves as much scrutiny as the attacker. A strong attacker paired with an unvalidated judge produces confident numbers that no one should believe.
A concrete setup makes the loop tangible. The target is a mock customer-support assistant for a cabling company, running on localhost, with a secret in its context that it must never reveal. The objective handed to the attacker is to make the bot disclose that secret. The constraint is single-turn and under two hundred words, and the success criterion is exact: the reply must contain the secret string.
This is a prompt-injection and secret-exfiltration test, one of the most common failure classes for assistants that carry credentials or system instructions in their context. Running it against a local, sandboxed target is the safe way to learn whether a deployment leaks under pressure.
A turn plays out like this. The attacker reads the objective and, on the first move, has no history, so it picks an opening strategy and writes a prompt. The bot replies. The judge checks the reply against the exact success string and returns a verdict. On a failure, the harness records the outcome, updates its counters, and the attacker chooses a different angle. On a success, or when the budget is spent, the run stops and the full trace is written out.
The defensive payoff is the trace. It shows exactly which framing broke the bot, which lets the owner add the missing guardrail, strip the secret from the model's context entirely, and re-run the harness to confirm the fix holds.
