On August 20, 2026, US legal AI company Harvey unveiled its first post-trained open-weight model, Harvey Tenet, built on the Kimi K3 foundation model from Moonshot AI, with Fireworks Research collaborating on legal-domain post-training. Over the past six months, Harvey's research has focused on two key areas: leveraging open-weight models to develop cutting-edge legal intelligence, and creating systems that enable law firms to train and own their own specialized models. Tenet represents the first public outcome of this research effort.
Tenet uses Kimi K3 as its base, with training data incorporating synthetic datasets, public legal materials, and human expert input, with an emphasis on improving performance in long-horizon, agentic legal tasks.
134 B300 GPUs and Two Months of Training
Harvey and Fireworks Research started from Kimi K3 and conducted asynchronous reinforcement learning within real-world legal work environments. The training setup mirrors the structure of the Legal Agent Benchmark (LAB). Each task includes partner-level work instructions, client matters with associated materials, and expert-defined scoring criteria. The model operates within a sandbox environment to search materials, review documents, and produce final legal work products.
Tenet employs Group-Sequence Policy Optimization for reinforcement learning, using Rank-64 LoRA to adapt the full Kimi K3 network, including Attention, MLP, and routed expert weights. Harvey's training dataset comprises roughly 1,750 legal agent task environments, with each training epoch executing 150 optimization steps, accumulating over 10,000 independent rollouts. The entire training process ran on 134 NVIDIA B300 GPUs over approximately two months. Harvey has explicitly confirmed that no client data was used during the post-training phase.
Test Results
Harvey's published results show that legal-domain post-training significantly boosted Kimi K3's performance on complex agentic legal tasks. In held-out LAB test tasks, Tenet's All-pass Rate improved by 9 percentage points over the original Kimi K3, and by 2 percentage points on LAB Contracts. In terms of successfully completed tasks, Tenet completed nearly twice as many tasks as the base Kimi K3 on LAB, and roughly 20% more on LAB Contracts.
Harvey's disclosed metrics show Tenet achieving an All-pass Rate of 19.7% on the Legal Agent Bench, compared to Kimi K3's 10.8%; on LAB Contracts, Tenet scored 11.3% versus 9.3% for Kimi K3. In relative terms, Tenet's All-pass Rate improved approximately 82% on LAB and 22% on LAB Contracts compared to Kimi K3. Tenet currently achieves state-of-the-art results on LAB Contracts and ranks second on the comprehensive LAB benchmark.
Harvey also tested Tenet on third-party legal agent benchmarks outside its training data. On Mercor APEX Agents Corporate Law, Tenet scored 74.0, surpassing Kimi K3's 58.8; on Crosby Redline Bench, Tenet recorded 55.5, also higher than Kimi K3's 49.3. Harvey notes that neither APEX Agents nor Redline Bench appeared in Tenet's training data, demonstrating that post-training capabilities transfer effectively across different legal agent benchmarks and runtime environments. On benchmarks emphasizing legal knowledge and reasoning—including APEX-v1, Scale Professional Reasoning Benchmark, LegalBench, CUAD, and MAUD—Tenet largely preserved Kimi K3's original legal knowledge levels.
Higher Performance, Stable Costs
Another key finding Harvey announced relates to cost efficiency. Open-weight models inherently offer lower token prices, but total model cost also depends on the number of tokens consumed to complete a task. During post-training, Harvey used reward shaping to encourage the model to reduce unnecessary tool calls and reasoning steps, prioritizing execution paths that consume fewer tokens when task performance remains equal. Results show that Tenet's All-pass Rate on held-out LAB tasks rose from 10.8% to 19.7% compared to Kimi K3, while per-task cost remained essentially stable.
Harvey Expands Investment in Open-Weight Models
Harvey views Tenet and its extensibility as a milestone in the company's research into training environment design, post-training, and open-weight models. Going forward, the company plans to expand LAB across more jurisdictions, practice areas, and workflows, while scaling up training compute and gradually moving current research findings into production environments. Tenet remains in the Research Preview stage.
More notably, Harvey's underlying model strategy is shifting. Rather than training a large foundational model from scratch, Harvey uses Kimi K3 as its base, then builds its own legal model using legal data, specialized task environments, expert scoring systems, and reinforcement learning. For Kimi K3, this also represents a significant deployment in a high-value professional services context.
Beyond serving as the base for Tenet, Kimi K3 has also performed strongly on independent third-party legal agent evaluations. In the Harvey LAB-AA benchmark independently run by Artificial Analysis, Kimi K3 achieved an All-pass Rate of 26.7%, well ahead of Claude Fable 5's 14.2% and Grok 4.5's 13.3%. This metric tracks the proportion of tasks where all scoring criteria are fully met, with no partial credit given. The Harvey LAB-AA benchmark is based on 120 private legal tasks provided by Harvey, spanning 24 legal practice areas, and was run independently by Artificial Analysis using its own agent harness.
Harvey provides AI services to large law firms and corporate legal departments, with more than 2,400 institutional clients and over 200,000 professional users across more than 70 countries. Over 75 Am Law 100 firms currently use Harvey. In March of this year, Harvey raised $200 million at a valuation of $11 billion (approximately 73.9 billion RMB), with current annualized recurring revenue around $350 million, and is seeking a new valuation of roughly $15.5 billion in its latest fundraising round.
Harvey's choice of Kimi K3 for its first publicly released legal post-trained open-weight model also signals that open models are entering high-value professional applications historically dominated by closed-source frontier models.
Note: If any of the above content contains errors or infringes upon the rights of your company, institution, organization, or individual, please contact us with an explanation, and we will cooperate fully and remove it without condition.