A TAR technique where the system learns from attorney-coded seed documents to predict relevance across the full document set; court acceptance depends on validation methodology.
Last reviewed: 2026/05/19
Move from this definition to role-based legal AI shortlists and the selection criteria that matter for each type of legal team.
Am Law 200 and global firm workflows: accuracy at scale, security compliance, and matter-level auditability.
Last reviewed: 2026/05/19. Definitions are written by the LawyerAI Editorial team. Commercial relationships are disclosed and do not determine editorial scores or conclusions. See our Sponsorship & Affiliate Disclosure.
Predictive coding is a specific technology-assisted review methodology in which a supervised machine learning model is trained on a seed set of documents that attorneys have coded as relevant or non-relevant, then applies those relevance predictions to the remaining unreviewed document population. The model learns the characteristics of relevant documents from attorney-coded examples and ranks or classifies the full document set accordingly. Predictive coding is a form of TAR (Technology-Assisted Review) distinguished by its supervised learning approach — requiring a carefully selected and consistently coded seed set. Court acceptance of predictive coding as a discovery methodology depends on the transparency and rigor of the validation process.
Predictive coding emerged as a practical response to document volumes in large-scale litigation that made traditional linear review economically and logistically impossible. A case with five million documents cannot be reviewed linearly at any reasonable cost; predictive coding enables meaningful review of the relevant document population at a fraction of the cost.
The seed set selection and coding quality are the most critical inputs. A poorly selected or inconsistently coded seed set produces a poorly calibrated model that misclassifies documents. Attorneys responsible for predictive coding workflows must invest in seed set quality, not merely volume.
Validation methodology is the dimension on which courts most frequently scrutinize predictive coding implementations. The supervising attorney must understand recall and precision metrics, be able to explain the validation sampling process, and be prepared to defend the choice of completeness threshold if challenged.
Relativity supports multiple TAR implementations including traditional predictive coding with seed set review and its Active Learning module, which implements continuous active learning — a more efficient variant. DISCO offers predictive coding integrated with its review platform, with validation reporting built into the workflow.
CasePoint provides predictive coding with automated reporting on training progress, precision, and recall, supporting the documentation needed to defend the methodology in court or in meet-and-confer discussions with opposing counsel.