Artfical AI / Security
Security overview Open tAI
Evaluation categories

Hazardous-material misuse

Chemical, biological, radiological, and nuclear misuse patterns, evaluated against a set we don't publish.

Scope

This category covers requests seeking specific, operational help with chemical, biological, radiological, or nuclear harm. It's evaluated against a fixed, non-public set of test prompts. That set is maintained separately from the rest of our evaluation suite. It's handled with a different, more guarded process than any of the other six categories in this section. This is the one category on this page where we deliberately say the least about methodology. That's a conscious editorial decision, not an oversight, and not a sign that less work went into it than the categories described elsewhere. If anything, the opposite is true. The shortness of this page is itself evidence of how seriously the category is treated, since the normal instinct to explain our reasoning in depth is deliberately overridden here.

The reasoning behind that restraint is consistent with how this category is handled across the field generally. It isn't a policy unique to Artfical. Publishing this category's test prompts would function as a partial map. So would a detailed account of exactly where the current boundary sits between a refused request and an allowed one. Either would help someone specifically trying to find the edges of the safeguard, rather than understand it in the abstract. That concern applies to every category on this page to some degree, since detailed disclosure always trades off against giving an adversary useful information. It is most acute here, given how much more severe the worst-case outcome of a successful evasion would be compared to any other category in this section.

Why so little detail is published here

Publishing this category's test prompts, or a detailed account of exactly where the current boundary sits, would function as a partial map for someone trying to find the edges of the safeguard. That's true of every category on this page to some degree, but it's most acute here, which is why this page is the shortest one in this section on purpose, not because less work went into it.

What we will say

Despite the limited methodology detail, there are several concrete, checkable commitments we're willing to state plainly. A policy commitment is different from a piece of methodology that could be reverse-engineered into an evasion technique. Every tAI version is evaluated against this category before release. There are no exceptions carved out for a version under schedule pressure. No version has ever shipped without a passing result on this specific category. Any confirmed regression here blocks release outright, in exactly the same way a child-safety regression does. That's true regardless of how strong the rest of that version's capability or other safety results happen to be. It puts this category in the small handful of areas treated as an absolute gate, rather than a score to be weighed against other considerations.

  • Every tAI version is evaluated against this category before release, with no exceptions and no version shipped without a passing result.
  • Any confirmed regression here blocks release, in the same way a child-safety regression does, regardless of what else improved in that version.
  • Category-level pass/fail and directional trend (held steady, improved, regressed) across versions is published in each model's system card, without granular sub-scores.

Category-level pass or fail status is published in each model's system card. So is the directional trend across versions: whether a given release held steady, improved, or regressed relative to the version before it. What isn't published is granular sub-scores, or any detail specific enough to reveal exactly where the current boundary sits within the category. That level of detail is precisely what this page's caution is meant to withhold. If you have a legitimate reason to need more detail than this page provides, academic research and a formal institutional security review are two examples of that, the right path is the contact channel described on the Responsible disclosure page. It is not this page, and it is not a general support request.

It's also worth being clear about what this category does not cover. The name alone can read more broadly than the actual scope. General chemistry, biology, and physics education are not what this category is built to catch. A question about how a radiation detector works is not what it's built to catch either. Neither is a student's homework question about nuclear fission. Treating those as though they were would make tAI meaningfully worse at ordinary science education, for no real safety benefit. The line this category actually draws is around specific, operational uplift toward causing serious harm, not around a subject area in the abstract. Getting that line right is exactly the kind of judgment call the evaluation set described above is built to test.

How the boundary gets reviewed

Where exactly the line sits between general education and specific operational uplift isn't decided once and left alone indefinitely. It's reviewed periodically, with people who have relevant subject-matter grounding in this specific area involved directly in that review, rather than left entirely to people whose background is in safety tuning generally but not this domain specifically. We're not going to describe the review cadence or the exact composition of that group in more detail than that. Doing so would start drifting into the same category of detail this page already declines to publish for good reason. The relevant fact for a reader of this page is that the boundary gets deliberate, qualified, recurring attention, not that it was set once at some point in the past and never revisited since. That's consistent with how every other category in this section is handled, even where, as here, the methodology detail behind it stays deliberately thin.

If a review concludes the boundary needs to move, whether tightening it further or correcting a place it was drawn too broadly, that change goes through the same release-blocking gate described above. It doesn't ship quietly as an unannounced adjustment to an existing version. A meaningful boundary change is treated with the same seriousness as a new version's initial evaluation round for this category, precisely because getting this specific line wrong in either direction carries real consequences. Correcting a boundary drawn too broadly matters just as much as tightening one drawn too narrowly, and both directions get the same level of scrutiny before a change ships. Neither is treated as the safer default that doesn't need the same rigor as the other. We'd rather a correction take longer and be right than ship quickly and need a second correction shortly after.