MITRE and Fujitsu Research developed a defender-focused framework for evaluating whether large language models can faithfully translate adversary behavior across operating systems while preserving technical realism and defensive value.
Evaluating LLMs for Impact-Faithful Translation of Adversary Behavior Across Operating Systems
Can You Trust an LLM to Translate Adversary Behavior Across Operating Systems?
New Research Puts Automated Translation to the Test.
As security teams increasingly turn to large language models (LLMs) to adapt threat intelligence and adversary emulation plans across diverse environments, such as porting a Windows attack chain to Linux or mixed-OS fleets, a critical question goes unanswered: does the translation actually preserve what matters to defenders? MITRE and Fujitsu partnered to answer it, and the findings are a caution for anyone planning to trust these translations at face value.
In a proof-of-concept trial, an LLM tasked with porting a Windows ransomware attack chain to Linux earned a reassuring pass rate above 90% on standard automated checks. But that number proved to be a mirage. Re-scored under defender-relevant definitions of telemetry equivalence, the pass rate fell sharply, and a graph-based structural analysis exposed failure modes that simpler checks miss entirely. The core problem is that a translation can keep every ATT&CK technique label intact while silently dropping the observable signals that detection rules depend on, rendering a defender's existing detections inert against the "same" behavior.
The core contribution is a defender-centric way to catch this: a procedure-graph approach with layered semantic enrichment that pinpoints exactly where a translation degrades, whether at the level of technique, tactic, telemetry, or detectability. As a direct byproduct, the work delivered eight production-ready Linux Sigma rules, six of them entirely novel, to the open-source SigmaHQ repository. The research also surfaces a second important finding. Expert manual review proved highly subjective, with reviewers reaching widely divergent verdicts on the same artifact, establishing inter-rater reliability as a concern of equal weight to any pass rate.
AI-assisted adversary emulation is coming, but a passing automated score is not evidence that a translation delivers real value to defenders. Presented as an early-stage methodology and an open invitation to the community, this work offers the first rigorous, model-agnostic way to tell the difference before a flawed translation gives teams false confidence in their coverage.