Data handling
What happens to your data, separate from what happens to the data tAI is trained on.
Two separate things, kept separate
Your conversations, your connector activity, and your account data generally are handled under a different, explicit policy. That policy is separate from the one governing the training corpora used to build tAI itself. Keeping those two things genuinely separate is a deliberate structural choice. It's not a detail of how the systems happen to be organized. "We use data to improve our products" language often blurs the two together elsewhere in the industry, and we've tried not to do that here. Where usage data is used to inform future training at all, it goes through the exact same review and filtering step that new corpus data goes through before it's eligible for inclusion, described in more depth below. That use is opt-out at the account level, meaning it's a choice you make rather than a default you have to discover and turn off.
This separation matters because the two things carry genuinely different expectations. Your own conversations and connector activity are yours, tied to your account. They're governed by the deletion guarantees described below. The training corpora, by contrast, are a shared, curated resource. They're built up over time from many sources, handled under the sourcing and licensing process also described below. Treating usage data as automatically part of that shared resource, without an explicit, separate decision, would blur a distinction we think matters. We've deliberately kept that distinction intact.
Deletion means deletion
When you delete a chat, disconnect a connector, or delete your account entirely, the underlying data is actually removed on our end. It is not flagged as hidden while remaining queryable through internal tools. It is not retained in some transformed or aggregated form that could later be reconstructed back into something identifiable. A user who opts out of usage-data-informed training, described above, is removed from that pipeline as well going forward. That removal isn't limited to their own visible chat history inside the product itself. A deletion that only affects what you can see, while a copy persists elsewhere, isn't a deletion in any meaningful sense, and we don't consider it one.
Where a short operational retention window exists, that's for something like abuse prevention or backup rotation. Both are legitimate operational needs that a purely instantaneous, zero-retention system genuinely can't satisfy. That window is bounded to a specific, disclosed length of time. It's described explicitly in the privacy policy, rather than left open-ended or vague. "We may retain data for legitimate business purposes," with no further detail, is the kind of language that sounds reasonable but commits to nothing checkable. We've deliberately tried to avoid that pattern here, in favor of specific, statable windows. If a window's length ever changes, that change is reflected in the privacy policy directly rather than applied quietly.
How training data is sourced
Both training corpora, ArtficalAI and the Artfical Code Index, pass through deduplication, quality filtering, and an explicit licensing check. That happens before a given source is added to either one. It is not audited for problems only after the fact, once a source is already part of a training run. That ordering, checking before inclusion rather than auditing after, is deliberate. It means a source that fails the check simply never becomes part of a model's training data in the first place. It avoids requiring a later removal-and-retrain cycle if a problem is discovered once it's already baked into a shipped checkpoint.
Personal information that appears incidentally inside scraped or aggregated text is filtered out during that process. An email address mentioned in a forum post is one example. A phone number in an old public listing is another. Neither is treated as a usable signal worth preserving. This filtering isn't described here as perfect, and we don't think it would be honest to claim it is. It's the standard applied consistently across every source that goes into either corpus. Where a genuine gap in it is found, whether through internal review or an external report, it's treated as a real limitation worth fixing directly, not a minor footnote to note and move past.
We don't sell your data
Not to advertisers, not to data brokers, not to anyone else, under any arrangement. That's not a policy we've committed to as a standalone promise sitting apart from how the business actually works. It reflects how revenue is actually structured. Revenue comes from the product itself. Usage-based API pricing and subscription tiers that people pay for directly are the actual source of income. It doesn't come from monetizing what's inside your conversations, your email, or your files through some secondary channel most users would never see or think to ask about.
The full legal and retention detail behind everything summarized on this page is in the privacy policy. That document covers specifics like exact retention windows and the legal basis for different kinds of processing, in the level of precision a policy document requires. This page is meant as the practical, plain-language summary of what that policy actually means for you day to day. It is not a replacement for reading the policy itself, if you need the exact legal wording for a specific purpose. Where the two would ever appear to disagree, the privacy policy is the authoritative source. We'd want to hear about the discrepancy so this page can be corrected.
Data portability
Deletion, described above, is one half of having real control over your own data. Being able to actually take it with you is the other half, and it's a working feature in the product today, not a promise described only in policy language. The "export your data" option in Settings produces a downloadable file containing your conversations and messages, structured so it's readable outside the product rather than locked into a format only tAI itself can open. It's available on demand, not gated behind a support request or a waiting period. You don't need a specific reason to use it, and requesting it doesn't affect your account or your existing conversation history in any way. You can export as many times as you'd like, and doing so has no effect on anything else in the product.
This is still evolving. The export today covers conversation history; broader account settings and connector metadata aren't yet part of it, and that's a real, current limitation worth stating plainly rather than implying the export is more complete than it is. Both halves, exporting what you have and deleting what you no longer want us to hold, are meant to actually work as described, not exist as a page like this one making a claim that isn't backed by a real, functioning mechanism behind it. Expanding what the export covers is planned work, tracked the same honest way as the other open items described throughout this section. We'd rather describe today's actual scope accurately than describe an aspiration and let a reader assume it's already true. If you specifically need something the current export doesn't include, the contact channel on the Responsible disclosure page is a reasonable place to ask.