Harvey and EngramLab have released an open-source synthetic dataset containing more than 100 million tokens, essentially building a fake law firm’s entire body of work so that real AI systems can learn how actual firms operate. The dataset covers over 250 synthetic client matters across 46 clients, with roughly 10,000 individual files representing the kind of institutional knowledge that typically lives inside the heads of partners who’ve been practicing for decades.
What the dataset actually contains
EngramLab’s contribution centers on its memory-layer technology, which the company claims can compress organizational context enough to cut token usage by up to 100x. In practical terms, that means an AI system could process the equivalent of a partner’s career worth of institutional knowledge without burning through compute budgets.
Harvey, for its part, has been building toward this kind of release. Earlier in 2026, the company open-sourced the Legal Agent Benchmark, known as LAB, which established standardized ways to measure how well AI agents perform on legal tasks. The synthetic dataset is a natural companion piece, giving researchers and developers actual training material to work with alongside those benchmarks.
The companies behind the release
Harvey has become one of the most prominent players in legal AI, currently valued at $11 billion. The company has positioned itself as a platform for law firms looking to integrate AI into their workflows without surrendering control of their proprietary knowledge.
EngramLab raised $98 million in funding on June 23, 2026, at a valuation of approximately $600 million. Its core technology focuses on learned memory layers for AI systems, reducing the computational overhead needed to maintain context over long interactions.
The partnership between the two companies predates this dataset release. Their collaboration has centered on encoding law-firm processes through AI, with the Legal Agent Benchmark serving as an early public output of that work.
Why open source matters here
Law firms are famously protective of their data, for good reason. Attorney-client privilege, work-product doctrine, and competitive secrecy all create enormous barriers to assembling the kind of training data that AI systems need. Synthetic data offers a workaround: all the structural complexity of real legal work, none of the confidentiality concerns.
What this means for the legal AI landscape
EngramLab’s 100x compression claim, if it holds up in practice, could prove particularly significant. A single M&A transaction might produce tens of thousands of pages of documents. Any technology that can maintain meaningful context across that kind of volume without proportional cost increases addresses one of the fundamental bottlenecks in deploying AI at scale within law firms.
For Harvey, the open-source strategy reinforces its position as an infrastructure layer for legal AI rather than just another point solution. By providing both the benchmark (LAB) and now the training data, the company is shaping the standards that the entire sector will build against.
Disclosure: This article was edited by Editorial Team. For more information on how we create and review content, see our Editorial Policy.

1 hour ago
17









English (US) ·