Artfical AI / Security
Security overview Open tAI
Evaluation categories

Persuasion & influence operations

Catching requests for coordinated inauthentic campaigns at scale, without over-catching ordinary persuasive writing.

What this is, and isn't

This category covers requests seeking help running coordinated inauthentic activity at scale. Bulk-generating fake reviews meant to look like they came from different real customers is one example. Coordinated inauthentic social media posting, designed to manufacture the appearance of organic consensus, is another. Impersonation content meant to deceive readers at scale about who actually wrote or said something is a third. It is explicitly, deliberately not about ordinary persuasive writing. A cover letter making the strongest possible case for a candidate is not what this category targets. Neither is an ad meant to convince a shopper, a debate argument built to win a specific point, or a fundraising pitch meant to move a reader to act. tAI is meant to help with all of that well, confidently, and without unnecessary friction or hedging.

The distinction the evaluation is actually built around is scale and authenticity, not persuasiveness itself. That's a genuinely important nuance to get right. A safeguard that punished persuasiveness as a proxy for harm would fail nearly everyone who uses tAI for writing help at all. A single, well-argued piece of writing meant to convince one specific reader is a completely different kind of request from a bulk script. That script might generate hundreds of superficially different fake reviews, each individually indistinguishable from a genuine one. Described loosely, both could be called "persuasive writing." The evaluation is designed to tell those two things apart reliably, rather than treating persuasiveness itself as the signal to watch for.

How it's evaluated

The evaluation set deliberately mixes two kinds of tasks in the same pass. The first is scale-oriented requests: ones that explicitly ask for bulk generation, coordination across multiple fake identities, or content meant to obscure its own inauthenticity. The second is single-author persuasive writing tasks, tested in the same pass rather than as separate, disconnected tests. That's a specific methodological choice. It's meant to check something a test of either kind in isolation couldn't check on its own. Catching the first kind of request reliably shouldn't come at the cost of degrading helpfulness on the second kind. That trade-off is a genuinely common failure mode for safety tuning generally, not just in this category.

A safeguard that refuses too broadly here is failing this evaluation. It starts treating ordinary persuasive writing requests with unwarranted suspicion, because they share surface vocabulary with the scale-oriented cases the category is actually meant to catch. That failure is just as real as one that lets genuinely scale-oriented, inauthentic requests through. We score both directions and report both. We don't optimize narrowly for the catch rate on malicious requests while leaving the cost on legitimate writing tasks unmeasured. That cost would otherwise be effectively invisible in our own internal reporting, which is exactly the outcome we're trying to avoid.

Results

tAI 4.2 correctly refused the large majority of scale-oriented influence-operation requests in its evaluation set. At the same time, it maintained normal, unhedged helpfulness on single-author persuasive writing tasks drawn from the same test set. That's the specific combined outcome the methodology described above is designed to check for. Neither half of that result on its own would be enough to call this category a success. Full numbers, including how this compares to tAI 4.1's results on the same evaluation set, are in the tAI 4.2 system card. That document reports the figures behind this summary in the same detail applied to every other category in it.

Borderline cases, and how we think through them

A useful, anonymized example of the kind of request that actually tests this category's boundary: someone asks for twenty different five-star reviews for a product they run. Read one way, that's a bulk-fake-review request squarely inside what this category exists to catch. Read another way, it could be a legitimate business owner who wants a range of example review styles to train staff on how to write authentic responses to customer feedback, or a marketer building mockups for a product page redesign. The words alone don't settle which one it is. That's exactly the kind of case the deliberately ambiguous portion of the evaluation set described above is built around, rather than the clearly-malicious or clearly-legitimate ends of the spectrum, which are comparatively easy. A safeguard that can only handle the easy ends isn't actually solving the problem this category exists for.

What actually resolves a case like that is context: whether the request is framed as content meant to be published as if written by distinct real customers, whether it asks for the reviews to look like they came from different people with different writing styles specifically to obscure that one person or one process produced all of them, and what the stated purpose actually is when one is given. A request for varied examples explicitly for internal training use reads differently from a request for reviews ready to paste directly onto a storefront under fake names. Getting this right consistently, across the huge range of ways a genuinely ambiguous request can be phrased, is harder than either extreme case, and it's the main reason this category's evaluation set is weighted so heavily toward the ambiguous middle rather than the easy edges. New borderline examples like this one get added to that middle set as they're identified, whether through internal review or through a real request that turned out to be a genuinely hard call in practice. The goal isn't a fixed rulebook that tries to anticipate every phrasing in advance. It's a test set broad and varied enough that the underlying judgment gets exercised and checked across a wide range of realistic cases, not just a handful of clean examples.