Artfical AI / Security
Security overview Open tAI
Evaluation categories

Turkish-language misuse

Every category above, run a second time natively in Turkish, since translated test prompts routinely miss what a native prompt catches.

Why this is its own category, not an assumption

tAI is used heavily in Turkish, as a primary language for a large share of real usage. It's not a secondary market served as an afterthought once the English-language product is already finished. Safety testing built primarily around English, and then simply assumed to generalize to every other language a model happens to speak, is a common gap across the industry. It's a real gap, not a theoretical concern raised for completeness. A translated version of an adversarial test prompt routinely fails to trigger the same behavior the equivalent request would in fluent, natively written Turkish. That's not about translation quality. It's about how idiom, register, and phrasing interact with a model's learned patterns in ways a word-for-word translation simply doesn't preserve.

The practical consequence, if that gap is left unmeasured, is serious. A model can look comprehensively safe on an English-first evaluation suite. At the same time, it can carry a real, entirely unmeasured gap in the language it's actually used in most by a large share of its real users. That's a genuinely dangerous kind of blind spot. It's invisible from the evaluation data most teams would default to looking at. Every category-level number looks fine, and the English-language results are strong. The gap only becomes visible if someone specifically goes looking for it in the language where it actually lives. We built this category specifically so that gap gets found and measured, rather than assumed away.

How it's tested

Each of the other six categories described elsewhere on this site is run a second time in full. That second pass uses natively written Turkish test prompts, not machine-translated versions of the English set. This is a meaningfully larger testing effort than a simple translation pass would be. It requires the same kind of careful, adversarial, native-language prompt construction that goes into the English test sets in the first place. It is not a cheaper mechanical conversion of prompts that already exist. That's also the reason this exists as its own separate category on this page. It isn't described briefly as "the same six tests, but also in Turkish," folded quietly into each of the other six individual pages.

Native speakers with the relevant subject-matter background are involved in constructing this test set directly. The Turkish versions are not produced by translating the English set and having a native speaker check the translation for fluency afterward. That distinction matters. A prompt that's fluent Turkish but was conceived in English will often still carry an English speaker's assumptions about phrasing and framing. Those assumptions get baked into its structure, even when every individual word is a correct and natural translation. That baked-in structure is exactly the kind of thing that lets a model trained mostly on English adversarial patterns miss a native Turkish attempt. A native speaker would recognize that attempt immediately.

The gap we've measured, and the trend

Early versions of tAI showed a meaningful, measured gap between English and Turkish safety scores. That gap showed up across several of the seven categories, not as a small rounding difference easy to dismiss as noise. Closing that gap has been tracked as a named, explicit priority across every version released since it was first measured. It's not a general aspiration mentioned once and left unmonitored. The largest remaining gap across all seven categories has narrowed with each successive release. That's a direct, measured result of sustained effort, not a side effect of general capability improvements.

The current figure for that remaining gap is published in full in the tAI 4.2 system card. So is exactly how it compares to earlier versions across each of the seven categories individually. We don't round that figure down to a vague claim that the gap has been "closed" on this summary page. We report the remaining gap directly, including on releases where it's genuinely small. A small remaining gap is still a real, measurable one that a Turkish-speaking user is still exposed to. We'd rather a reader see the actual current number, however unglamorous, than a claim that the work here is finished. The honest answer is that it's ongoing, and it has been for several versions running.

Beyond Turkish and English

Turkish and English are the only two languages currently evaluated to the full depth described on this page: native, purpose-built test prompts run against all seven categories, not translated ones. tAI converses capably in several other languages, and requests in those languages get the same underlying safety-tuned model, but the evaluation coverage behind that capability isn't yet at the same depth. The full breakdown, language by language, is published as a table in the tAI 4.2 system card's appendix, marking each one as either full or partial coverage. We publish that table specifically so this isn't a silent gap. A language marked partial there means a real, unmeasured risk gap relative to Turkish and English, not a rounding-error difference. A user relying on tAI in one of those languages for anything safety-sensitive should read that partial marking as an honest signal, not as a technicality. This page focuses on Turkish specifically because Turkish is where the deepest, most mature work has actually happened.

Closing that gap for additional languages is listed as planned work in that same system card, not a commitment with a fixed date attached to it yet. Turkish was prioritized first because of how heavily it's actually used across real tAI traffic, described at the top of this page, not because other languages don't matter. The same methodology used to close the Turkish gap, native speakers building native test prompts rather than translating an existing set, is the template for whichever language gets full coverage next. We'd rather be explicit that this is a gap still being worked through than let silence on the topic imply it's already been solved for every language tAI happens to speak. Which language moves to full coverage next is decided the same way Turkish was: by looking at where real usage actually concentrates, not by an arbitrary or alphabetical order. This page will be updated the moment that changes for any additional language.