What Counts as an AI-Generated Document When a Human Edited the Output
Most governance policies still treat AI-generated content as a binary category: either a document came from an AI tool or it did not. In practice, almost no document works that way anymore. An employee asks an AI assistant to draft a summary, rewrites two paragraphs, deletes a section, and sends it. Is that document AI-generated, human-authored, or something in between?
AI-generated content evidence refers to any document, message, or record produced with AI assistance, including drafts substantially rewritten by a human, that may be subject to collection, preservation, or production obligations because of how it was created, regardless of how much human editing occurred afterward. The amount of human editing does not remove the record from scope. It changes how the record needs to be classified, collected, and explained.
Why This Question Doesn't Have a Clean Answer
The closest existing framework for thinking about human contribution to AI output comes from copyright law, not eDiscovery, but it is instructive. The U.S. Copyright Office's January 2025 report on AI and copyrightability concluded that prompting alone does not establish meaningful human authorship, while substantial human modification, selection, and arrangement of AI output can. That threshold question, how much human involvement changes the character of a document, is exactly the question legal and compliance teams now face when deciding how to classify a document for collection purposes.
In practice, AI-assisted content spans a spectrum rather than two categories: a fully AI-generated draft that is produced and sent unedited, an AI draft that receives light copyediting, an AI draft that is substantially rewritten, and a human draft that an AI tool polishes or reformats. Each point on that spectrum raises a different question about what needs to be preserved.
Why the Distinction Matters for eDiscovery and Compliance
Whether a document counts as AI-generated affects more than a label. Onna's guide on what counts as digital evidence when AI wrote it explains that prompts, drafts, and outputs are each independently evidentiary, meaning a heavily edited final document does not erase the relevance of the AI draft that preceded it. If a custodian's intent, knowledge, or reasoning is at issue, the earlier AI draft may matter more than the polished final version.
The authentication risk compounds this. The 2025 EDI Leadership Summit's panel on AI-generated data in eDiscovery was convened specifically to address unresolved questions about authentication and admissibility, alongside authorship and privilege, as AI-generated content becomes embedded in everyday business operations. A document that started as AI output and was later edited by a human carries that same authentication burden, whether or not the final version reads as fully human-written.
How to Classify Human-Edited AI Content for Collection Purposes
Preserve the Full Edit History, Not Just the Final Document
A final document alone cannot answer how much of it originated with AI. Onna's guidance on preserving AI-generated content in collaboration platforms covers why capturing the original AI output, alongside the edited final version, is what actually allows a later reviewer to assess the degree of human involvement.
Treat the Prompt and Draft as Part of the Record
The prompt that produced the first draft, and the draft itself, are part of the document's history even if neither survives in the final version. Onna's approach to identifying AI-generated content in data outlines the detection signals, including revision histories that jump from a short prompt to a complete draft, that help reconstruct this history during collection rather than relying on a document's final appearance alone.
Document the Workflow, Not Just the Output
Classification decisions need to be repeatable and defensible, not made ad hoc by whichever reviewer encounters the document first. Onna's guidance on legal holds for AI-generated content explains why hold and collection processes need to name AI tools explicitly and capture session data, so the question of how much a human edited an AI draft can be answered from the record itself rather than reconstructed after the fact.
Building an AI-Generated Content Collection Platform Approach
Organizations that handle this well tend to build their collection process around a few consistent practices:
- Capture drafts and final versions together, so the edit history is available rather than lost the moment a document is finalized.
- Preserve prompts alongside outputs, since the prompt often reveals context the final document does not.
- Avoid binary AI-generated tags. A classification system that only allows "AI" or "not AI" will misclassify the majority of real-world documents.
- Apply consistent criteria across matters, so classification decisions do not vary based on which reviewer or team handles a given collection.
Getting Ahead of a Question Courts Are Still Defining
Courts have not settled on a bright-line rule for how much human editing changes a document's status, and they may not for some time. Organizations do not need to wait for that clarity to build a collection process that captures the full history of AI-assisted content rather than just its final form.
If your organization needs a more precise way to classify and collect human-edited AI content, Onna's team can walk through the right approach for your environment. You can also see how this works in practice by requesting a demo.
Subscribe to our newsletter
Get Complete Visibility into Your Unstructured Data, Today
Complete initial setup and first collection in one business day. No lengthy implementations. No IT backlog. Just full visibility into your collaboration data when you need it most.

