Collecting ChatGPT Conversations for eDiscovery
OpenAI's own CEO has said it plainly: conversations with ChatGPT carry no legal confidentiality, and the company is required to retain user chats, including deleted ones, when a court orders it. As Tyson & Mendes details, in OpenAI, Inc., Copyright Infringement Litigation, a federal judge compelled OpenAI to produce 20 million de-identified ChatGPT conversation logs to plaintiffs, rejecting privacy objections because the logs were relevant to the case regardless of whether they touched the disputed content directly. Most legal operations teams have not updated their collection playbook to match.
ChatGPT collections refer to the process of identifying, preserving, and exporting ChatGPT prompts, responses, and related conversation data as electronically stored information (ESI), so this content can be reviewed and produced in litigation, investigations, or regulatory matters. Because ChatGPT retention and export options differ by account tier, from free consumer accounts to Enterprise and API deployments, a ChatGPT collection strategy has to account for where the data lives before it can account for how to get it out.
ChatGPT Conversations Are Discoverable, and Courts Have Confirmed It
The relevance standard for ChatGPT data is the same standard that applies to email or Slack messages: is it tied to a claim or defense and is production proportional to the needs of the case. In the OpenAI copyright litigation, the court found that even logs unconnected to the plaintiffs' specific works could be relevant, since they could help establish or rebut a fair use defense, a reasoning that broadens what counts as discoverable well beyond a narrow keyword match. That single ruling reframes ChatGPT interaction data as a mainstream discovery category rather than an edge case.
There is one meaningful nuance legal teams should track. As Kang Haggerty notes in The Legal Intelligencer, when an attorney uses ChatGPT to draft analysis or explore legal theories, some courts have treated those prompts as reflecting counsel's mental impressions, extending work-product protection to them in the same way it would apply to handwritten notes. That distinction does not extend to prompts and outputs generated by non-legal employees in the ordinary course of business, which remain subject to standard discovery rules.
Why ChatGPT Data Is Different from Traditional ESI
Email and chat platforms were built with retention and legal hold in mind. ChatGPT was not, and that gap shows up in two places.
Retention Depends Entirely on Account Tier
A free, Plus, Pro, or Team ChatGPT account behaves very differently from an Enterprise or API deployment. Deleted conversations on consumer tiers are generally purged within 30 days absent from a specific legal hold, while Enterprise and API customers with a zero data retention agreement may never have prompts or outputs retained at all. That means the same employee, using a different login, can produce two entirely different discovery postures for the same organization. Before a matter opens, legal and IT teams need a clear map of which ChatGPT tier every custodian is using, not just which tier the organization has licensed.
Shadow AI Use Creates Its Own Blind Spot
Enterprise ChatGPT deployments come with admin controls and audit logs. Personal accounts used for work do not, and employees pasting internal documents into a free ChatGPT session leave almost no trace an IT team can see without looking for it. Onna's connectors close part of this gap by pulling structured activity data from the collaboration platforms where ChatGPT use shows up, such as shared drives, chat tools, and email, so legal teams are not limited to whatever a custodian volunteers in an interview. Pairing that with data activity monitoring as an early warning system means unusual file exports or mass downloads tied to AI tool use can surface before a matter reaches the collection stage, not after.
What a Defensible ChatGPT Collection Requires
A defensible collection needs more than the AI's final output. It needs the prompt that produced it, the timestamp, the account and session identifiers, and enough structure to show which parts of a document were human-authored and which were AI-generated. Onna's guidance on what counts as evidence when AI wrote it walks through why the prompt itself is often the most substantively important part of the record, since it can reveal what an employee knew, was researching, or intended at the time.
Preservation has to happen before collection can succeed. Many ChatGPT interactions live in a session that auto-deletes on a fixed schedule unless a hold explicitly pauses it. Onna's guidance on how to preserve AI-generated content across collaboration platforms sets out the components a defensible preservation record needs: the full output text as it existed at the time, the originating prompt, and the metadata layer that establishes provenance and authenticity.
Building a Repeatable ChatGPT Collection Workflow
Organizations that treat ChatGPT collection as a one-off scramble every time a matter opens are recreating the same manual process each time. A short list of practices closes most of that gap:
- Name ChatGPT and other AI tools explicitly in hold notices. A generic instruction to preserve relevant electronic communications rarely prompts a custodian to export an AI chat history.
- Map account tier by custodian, not just by organizational license, since free and Enterprise accounts carry different retention defaults and export paths.
- Extend custodian interviews to cover AI tool use, asking which tools, how often, and whether outputs were saved to a shared system.
- Monitor for AI-related data activity, including unusual exports or deletions tied to custodians on hold, so gaps surface before production deadlines.
- Centralize ChatGPT data alongside other collaboration data, rather than managing it as a separate, manually reconciled export.
The Outcome Legal Teams Should Be Planning For
ChatGPT conversations are not a future discovery problem. They are already being compelled, litigated over privilege, and produced at a scale of millions of logs in active cases. Organizations that build a repeatable collection workflow now, one that names AI tools in holds, tracks account tiers by custodian, and treats AI interaction data as a first-class ESI category, will not be reconstructing that process for the first time under the pressure of a discovery deadline.
To see how a structured ChatGPT collection workflow works in practice, book a demo with Onna or contact the Onna team to talk through what an AI-ready information governance program should look like for your organization.
Subscribe to our newsletter
Get Complete Visibility into Your Unstructured Data, Today
Complete initial setup and first collection in one business day. No lengthy implementations. No IT backlog. Just full visibility into your collaboration data when you need it most.

