Glossary
Terms used throughout this section, defined the way we use them internally.
This page collects the specific terms used repeatedly throughout the rest of this security section. They're defined the way we actually use them internally. That's not always the same as a generic dictionary definition. A generic definition might not match the specific, narrower meaning a term carries on this particular site. Several of these words have a more general meaning in the security field broadly. Evaluation category, red-teaming, and scope violation are the clearest examples. That general meaning doesn't always line up exactly with how we've chosen to use them here. This page is meant to remove any ambiguity about that, for a reader encountering the term for the first time on one of the other pages in this section.
We've kept each definition short. We've pointed it back to the specific page where the term is actually used and explained in full context. We haven't tried to make this page a complete, self-contained explanation of every concept on its own. The goal here is a quick, reliable lookup. That's for someone mid-way through reading another page who hits an unfamiliar term and wants a fast answer before continuing. It is not a substitute for reading the page that term actually belongs to, if you want the full reasoning and context behind it.
A handful of these terms are specific enough to this site's own internal terminology that you won't find an equivalent definition anywhere else. System card and model card are the clearest examples. Those specific document names, and what they each contain, are conventions we've adopted and defined for our own publishing practice. They aren't terms with a single fixed meaning across the industry generally. Where a term does have a more standard, widely recognized meaning elsewhere, we've tried to stay close to the common usage. Sandboxed session and approval gate are reasonable examples of that. We haven't tried to invent a narrower, idiosyncratic definition, just for the sake of precision.
If you read a term used on another page in this section and it isn't listed here, that's worth reporting as a gap in this page. Don't assume the omission was deliberate. The intent is genuinely for this list to be a complete companion to the rest of the section. It's not meant to be a partial one, covering only the terms someone happened to remember to add when the page was first written. The list below is organized roughly in the order the terms are most likely to come up. That's the order you'd hit them while reading through this section from the overview page onward, rather than strict alphabetical order. The theory is that a reader working through the section in order is a more common case than one jumping straight to this page looking up a single specific word.
- Evaluation category. One of the seven fixed areas, listed on the overview page, every model is tested against before release.
- Red-teaming. Structured, adversarial internal testing aimed at eliciting unsafe behavior before release, distinct from the standard per-category evaluation.
- Scope violation. A tool action that touches something outside what a request or a connector's granted access actually covers.
- Prompt injection. Instructions embedded in content the model reads, rather than typed by the user, written to override the user's own instructions.
- Approval gate. A required, explicit confirmation step before an irreversible action executes, enforced at the product level regardless of the model's own judgment.
- Over-refusal. A model incorrectly declining a legitimate request because it superficially resembles an unsafe one.
- Staged rollout. Releasing a new version to a smaller share of traffic first and expanding over time, rather than switching every user over at once.
- System card. The full, detailed evaluation writeup published for a model version, including what didn't pass on the first attempt.
- Model card. A shorter summary of a system card's key results.
- Sandboxed session. An isolated execution environment tCode or tAI's tools operate in, physically separate from account-data infrastructure.
- Safe harbor. A commitment not to pursue legal action or account enforcement against good-faith security research conducted within stated bounds, described on the Responsible disclosure page.
- Data portability. The ability to export your own data in a usable, downloadable form, distinct from and independent of deletion, described on the Data handling page.
- False positive. A legitimate user or action incorrectly flagged or restricted by production monitoring, tracked and reviewed with the same seriousness as a missed real abuse case.
- Severity classification. The judgment call, made by the safety review function, about how urgently a confirmed finding needs a response, based on scope of exposure and reversibility.
Every one of these terms is linked from at least one other page in this section, at the point where it's first used in real context. If a short definition here leaves you wanting the fuller explanation, follow that link back to the source page. That's generally a better next step than expecting this glossary entry to expand further on its own. This page is meant to be the shortest, fastest one in the entire section, by design. We'd rather keep it that way. We don't want individual entries to grow into miniature versions of the pages they're already pointing to.