AI that processes multiple input types — text, images, tables, scanned PDFs — in a unified model; legal applications include scanned document review, exhibit analysis, and financial disclosure extraction.
Last reviewed: 2026/05/19
A subset of machine learning using multi-layered neural networks that powers contract clause extraction, semantic search, and LLMs; modern legal AI tools are predominantly deep learning systems.
Tech / ModelAlgorithms that learn patterns from labeled legal data — relevance decisions, risk labels, outcome records — to make predictions on new documents or cases; TAR is the most established application.
Enterprise AI for portfolio-level contract analysis and institutional memory.
Thomson Reuters' GPT-backed legal research and drafting with Westlaw integration (relaunched as CoCounsel Legal, 2025).
Move from this definition to role-based legal AI shortlists and the selection criteria that matter for each type of legal team.
Am Law 200 and global firm workflows: accuracy at scale, security compliance, and matter-level auditability.
Legal department workflows: contract lifecycle, regulatory tracking, outside counsel management, and risk.
Last reviewed: 2026/05/19. Definitions are written by the LawyerAI Editorial team. Commercial relationships are disclosed and do not determine editorial scores or conclusions. See our Sponsorship & Affiliate Disclosure.
Multimodal AI refers to artificial intelligence systems capable of processing and reasoning across multiple input modalities — including text, images, structured tables, handwritten content, scanned documents, charts, and audio — within a unified model, rather than requiring separate specialized tools for each input type. In legal applications, multimodal capabilities enable processing of scanned contracts and court filings, extraction of data from financial exhibits and spreadsheet-format disclosures, analysis of diagram evidence, and review of mixed-format document sets containing both text-native and scanned materials. Most current legal AI tools are primarily text-based; multimodal capabilities are expanding rapidly as foundation models like GPT-4V and Gemini integrate vision capabilities.
Legal practice involves a wide range of document formats beyond text-native PDFs. Scanned legacy contracts, handwritten notes, deposition exhibit binders with mixed document types, financial statements with tables and charts, technical drawings in patent matters, and photographs as evidence are all common inputs that text-only AI tools cannot process.
Multimodal AI expands the set of documents that AI tools can analyze. A document review exercise that includes 20% scanned documents — a common situation in older litigation matters — would benefit from multimodal AI that can process scanned documents directly, rather than requiring OCR pre-processing that may introduce errors.
In transactional work, the ability to extract data from financial exhibits — tables, pro forma financial statements, cap tables — without manual data entry is a time-saving capability. In IP matters, analyzing patent drawings alongside claim text in a unified model enables more sophisticated prior art and infringement analysis.
Harvey integrates multimodal capabilities from underlying foundation models, enabling document analysis that spans text-native and image-based content in the same workflow. Luminance applies multimodal processing to contracts and documents that include tables and charts, extracting structured data from non-text-native formats.
CoCounsel has expanded its document processing capabilities to handle mixed-format document sets more comprehensively as underlying model capabilities have advanced. Most tools continue to perform better on text-native than image-based content; the gap is narrowing.