Frontier AI safety evaluations should begin by spelling out who might misuse a model, a new preprint argues. It proposes a threat-actor profile that sets out six attributes before an evaluation is designed, so the people, tasks and timing used in the test follow from stated assumptions. The focus is chemical, biological, radiological and nuclear (CBRN) and offensive-cyber misuse, with particular attention to open-weight releases, where the authors say weights cannot be recalled, safeguards cannot be patched and usage cannot be monitored.
The document is an arXiv version 1 preprint dated 26 August 2026. The authors describe the research design as a three-stage effort to standardize threat-actor assumptions and connect them to pre-release evaluations. It is centered on CBRN and offensive-cyber risks, while broader governance questions such as acceptable risk levels and staged release are left for future work.
A framework built around the adversary
At the center is a matrix with six attributes and 29 total tiers. They cover technical sophistication, prior domain knowledge, organizational capacity, operational infrastructure, financial capacity and time horizon. The tiers run from least to most capable or resourced. The intended use is to give evaluators a common way to state what sort of actor an evaluation assumes before they choose how to run it.
The attributes were distilled through a systematic review. Most were developed through conceptual reasoning, while financial capacity and time horizon were grounded with empirical approaches. The paper assessed the taxonomy's fitness for purpose using eight objective and five subjective conditions for judging whether it was fit for purpose.
The policy problem behind the proposal is inconsistency. In a survey of 12 frontier model developers with published frontier safety policies, the paper found variation in which threat-actor dimensions were highlighted and how concretely actors were described. The proposed matrix answers that variation with a fixed set of attributes, making the assumptions behind an evaluation explicit before its design is locked in.
The evidence behind the tiers
Some of the tiers are tied to reported operational data. Financial-capacity tiers were empirically anchored using an analysis of 40 terrorist plots in Western Europe. For time horizon, the paper cites the American Terrorism Study's 1,360 cases, including 272 with usable temporal information, alongside a cyber anchor based on approximately 500,000 hours of frontline incident investigations. The reported cyber proxy had a global median dwell time of 14 days in 2025.
The time figures come with a built-in warning. Terrorism data track preparation before an incident, while the Mandiant data measure dwell time during an active intrusion. The authors treat those sources as non-equivalent proxies for the time-horizon tiers. The 14-day cyber median, in that context, is a relative anchor rather than a directly comparable measure of terrorism planning time.
When the profile changes the test
The proposal becomes more concrete in two worked examples. The paper presents two hypothetical complete threat-actor profiles, one for CBRN and one for offensive cyber. The profiles lead to different evaluation priorities: the CBRN case favors domain-expert participants and agentic planning tasks, while the cyber case favors multi-session designs and knowledge-retrieval benchmarks, tests of finding and using relevant information.
That is the paper's practical point: assumptions about the actor shape the evaluation around them. It recommends populating a profile before designing an evaluation and treating it as a pre-commitment, an early decision that constrains later choices. The paper also identifies an evaluation-design consideration for each taxonomy attribute, linking the profile to decisions about participants, tasks, benchmarks and session structure.
A proposal for release records
At the institutional level, the proposal places all six attributes in frontier safety policies and links model-card descriptions of evaluations to them. The authors argue that this is especially important for open-weight releases because the weights cannot be recalled, safeguards cannot be patched and usage cannot be monitored.
Several limits remain visible in the evidence. The policy survey is limited to 12 developers that had published frontier safety policies. The two worked profiles are hypothetical examples, and the time-horizon anchors combine sources that measure different operational phases. Those boundaries frame the taxonomy as a way to organize pre-release evaluation design, while the paper leaves questions about acceptable risk levels and staged release for future work.
The conclusion is a clear sequence: define the adversary first, then design the evaluation around that definition and document the choices that follow from it. For open-weight models, the paper presents that prospective reasoning as a central part of managing misuse risk before release.
Paper data and sources
Original title: Toward a Threat Actor Profiling Taxonomy for Pre-Release Risk Management of Open-Weight Frontier Models
Authors: James Zhang
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text