Artfical AI / Security
Security overview Open tAI
Per-model detail

tAI 4.2

Our current flagship. What changed in its security posture, and where the full writeup lives.

What's new in this version's safety work

tAI 4.2 introduced three specific, checkable changes to how safety work is done. None of them were true of any version before it. Each one is worth understanding on its own, rather than as a single vague claim of "improved safety" that could mean almost anything. The first is that tool-use safety was tuned as its own dedicated training pass for the first time in the model's development history. It's separate from general conversational safety tuning, described in full on the Tool-use safety page. This was driven by internal evaluation. That evaluation showed the combined approach used in earlier versions was measurably missing a specific failure mode. The dedicated pass catches that failure mode more reliably.

The second change is that Turkish-language safety evaluation was expanded. It had previously been a smaller spot-check pass. It's now a full parallel track, run across all seven evaluation categories. That's described in depth on the Turkish-language misuse page. It closed a meaningfully larger fraction of the measured English-Turkish gap than any single prior version had managed on its own. The third change is a refusal-calibration correction, and we think it's worth describing honestly rather than glossing over. An earlier development checkpoint over-corrected toward refusing legitimate coding and security-research questions. Those questions merely resembled unsafe requests in surface phrasing. That checkpoint was rejected before shipping, in favor of a re-tuned version with a lower over-refusal rate than tAI 4.1's own, not just lower than the over-corrected checkpoint that got rejected.

  • Tool-use safety tuned as its own dedicated training pass for the first time, separate from general conversational safety tuning.
  • Turkish-language safety evaluation expanded from a spot-check pass to a full parallel track across all seven categories.
  • A refusal-calibration correction: an earlier development checkpoint over-corrected toward refusing legitimate coding and security-research questions, and was rejected before shipping in favor of a re-tuned checkpoint with a lower over-refusal rate than tAI 4.1's.

Read the full writeup

The complete evaluation results behind this summary are category by category. They use exact figures, rather than the general directional language used on this page. That includes a fuller account of the rejected checkpoint mentioned above, and specifically what was wrong with it. All of that is in the tAI 4.2 system card. It's published as a PDF specifically so it can be shared, cited, and archived as a standalone document, independent of this website's own structure. A shorter model card, also a PDF, covers the same ground at summary length. That's for a reader who wants the headline results without the full detail.

We'd rather point to those two documents as the authoritative source for exact numbers than restate specific figures here on this page. Figures restated in a second place tend to drift out of sync with the source of truth over time, as new revisions get published. A reader checking a number here against the system card itself should always find the two in agreement. They should never discover a discrepancy between a summary page and the document it's meant to be summarizing. This page will be updated if a future revision of either document changes something described here. The revision history inside the system card itself is the place to check if you want to see exactly what changed and when.

Known open items

The in-voice prompt-injection failure mode remains this version's single largest unresolved gap. It's described in detail on the Prompt-injection resistance page. We'd rather state that plainly here on the model's own page. We don't want it to be a detail a reader only finds by reading the prompt-injection page separately and connecting the two on their own. Injected content that mimics the user's own writing style and stays plausibly in-scope is the hardest sub-case within the hardest category tAI 4.2 is evaluated on. Narrowing it further is active, ongoing work. It isn't a solved problem we're carrying forward with a caveat nobody is expected to actually read.

How this page gets updated going forward

This page isn't a one-time snapshot frozen at tAI 4.2's initial release. As safety-tuning point releases ship against this version, described in more general terms on the Release process page, both this page and the linked system card are updated together, in lockstep, rather than letting one drift ahead of the other. The system card's own revision history, in its appendix, is the authoritative place to check exactly what changed and when. This page's summary is updated to match whenever that happens. The two are meant to be read as a pair, not as one being an occasionally-outdated shadow of the other. If you only ever read one of them, this page is meant to still be accurate on its own.

If a category described above moves meaningfully, whether the known open item narrows, a new point release addresses something specific, or a new open item gets identified, that shows up here as well as in the card itself. We'd rather a reader checking only this summary page still come away with an accurate, current picture, rather than one that was correct only at the moment tAI 4.2 first shipped and has since quietly gone stale. This same discipline extends to what eventually happens once a successor version ships: rather than leaving this page frozen at whatever it said the day tAI 4.2 launched, it gets the same kind of retrospective update the tAI 4.1 page received once tAI 4.2 itself shipped. That's a concrete, checkable pattern now, not just a stated intention. A reader can verify it by comparing this page's own history against what the tAI 4.1 page actually says about being updated after the fact. We'd rather be held to that pattern than simply promise to follow it.