Artfical AI / Security
Security overview Open tAI
Reference

Responsible disclosure

What to do if you find a way to make tAI behave unsafely, or a real gap in a safeguard described on this site.

What to report

Report any safety or security issue you find. A working prompt-injection technique that gets past the resistance described on its own page is one example. A tool-use scope violation, where access reached further than it should have, is another. A connector permission that reaches further in practice than what's documented for it is a third. A jailbreak that reliably defeats one of the seven evaluation categories described throughout this section is a fourth. The bar for what's worth reporting is deliberately low. We'd rather hear about something that turns out to already be understood and handled than never hear about something that turns out to be real and unaddressed.

If you're genuinely not sure whether something you've noticed actually qualifies as a real finding, report it anyway. Don't decide on your own that it's probably not worth our time. A false alarm costs us a small amount of triage time and nothing more. A missed real issue costs considerably more than that, both to us and potentially to other users. That asymmetry is exactly why we'd rather set the reporting bar low. The alternative is risking someone deciding, understandably but incorrectly, that their finding probably isn't significant enough to mention.

How reports are handled

A report submitted through this channel is triaged using the exact same process an internal red-teaming finding goes through. That process is described in full on the Red-teaming page. It's logged individually, assessed for severity, and, if confirmed, routed directly into training and safety tuning. It isn't patched narrowly in a way that wouldn't generalize to related variants of the same underlying issue. That shared process matters. It means an external report gets exactly the same weight and the same quality of response as something our own internal team happened to find first. It's never treated as a lower-priority, second-class input.

A confirmed finding results in either the next safety-tuning point release or the next major version, whichever comes sooner. That depends on how the finding fits into the broader release schedule, described in more detail on the Incident response page. It doesn't sit in an indefinite backlog with no clear path to actually being addressed. You'll generally hear back about the status of a report you've submitted. You won't be left to wonder whether it was ever looked at. The specific timeline naturally depends on how complex the underlying issue turns out to be, once we've had a chance to actually dig into it.

Credit

A reporter is credited by name, if they want to be. That happens in the affected model's next published revision, when their report leads to a confirmed fix. It gives genuine public acknowledgment for work that materially improved the product's safety. We don't treat a valuable external contribution as something to be quietly absorbed without recognition. This is currently handled informally, rather than run as a structured bounty program with fixed monetary payouts tied to severity tiers. We're being upfront about that limitation. We don't want to imply a level of formal reward this program doesn't yet actually offer.

Formalizing this further is tracked internally as the broader practice matures. That includes the question of whether a structured bounty program with real monetary rewards makes sense at our current scale. We haven't ruled it out. If that changes, it'll be reflected here, rather than announced only elsewhere and left for a reader of this page to discover separately. We'd rather grow this program deliberately, in a way we can actually sustain. Announcing something bigger than we're ready to run consistently would be worse than being upfront about the current, more informal state of it.

Where to send it

Use the contact details listed on ai.artfical.com/about for a general safety or security report. Use the contact address listed in the specific connector's own documentation instead, if the issue you've found is specific to one connector rather than the model or product generally. A connector-specific report often reaches the right team faster through that more targeted channel. Include enough detail to actually reproduce what you found. A report we can't reproduce ourselves is considerably harder to act on quickly and confidently, even when we take the underlying concern seriously. If you'd rather not put technical detail in a first message, a short note describing what you saw is still a useful starting point, and we can follow up with specific questions from there.

Safe harbor for good-faith research

Good-faith security research conducted specifically to find and report a real issue, following the reporting practice described on this page, is not treated as a violation of our usage policy. That's true even where the research technically involves attempting to defeat a safeguard, since that's the entire point of the exercise. We won't pursue legal action or account enforcement against a researcher acting in good faith within the bounds described below. We'd rather remove any incentive to stay quiet about a real finding out of concern for how we might react to it. A chilling effect on exactly the kind of research this page is asking for would be a worse outcome for everyone than the finding itself. This is the same underlying instinct behind the low reporting bar described above, applied specifically to the legal and account-standing risk a researcher might otherwise reasonably worry about.

That protection has reasonable, stated bounds. Testing should stop at proof-of-concept: enough to demonstrate and reliably reproduce the issue, not enough to access, exfiltrate, or retain more data than that requires. It shouldn't degrade service for other users or involve testing against another person's account without their permission. Research staying within those bounds, and reported through the channel described above rather than disclosed publicly first, is covered. If you're ever unsure whether something you're planning falls inside these bounds, ask first through the same contact channel; we'd much rather answer that question in advance than have it become a dispute after the fact. A quick question up front costs nothing and avoids any ambiguity about whether a specific piece of research was covered.