估值达156亿美元的法律科技新创公司Harvey,由于旗下AI代理使用量暴增导致运算成本飙升,毛利率从年初的50%暴跌至6月的-50%,促使其转向拥抱开放权重模型以降低对OpenAI及Anthropic等AI巨头的依赖。Harvey随后采用中国月之暗面(Moonshot AI)的Kimi K3开发出自有客制化模型,以极低成本达到媲美顶级专有模型的表现,并成功让毛利率转正;此趋势同时获得红杉资本与General Catalyst等顶级创投支持,成为新创对抗高昂运算费用与平台依赖的重要手段。
除了Harvey之外,Abridge、Decagon、Ramp、Rogo及Cursor等众多软体与垂直领域新创也纷纷投入建构客制化模型或分流查询,借此降低对外部高价API的依赖并强化自主控制权。然而,这种转向开源与开放权重的趋势同时引发了关于资安、数据隐私与地缘政治的疑虑,美方监管单位亦对开源模型及中国AI开发者是否蒸馏美国前沿模型展开严密审查,促使部分新创在采用时必须向客户审慎沟通以化解顾虑。
自行建构与微调模型同样面临高昂的前置技术门槛,包含争夺年薪高达数百万美元的稀缺AI专业人才、获取高品质专有训练数据,以及自行维护推论运算基础设施的庞大开销。因此,并非所有企业都能在此模式中获益,部分低流量或资源有限的新创仍发现直接采用闭源商业API更具成本效益,多数业者目前也未完全脱离顶级闭源模型,而是将最繁复的核心任务保留给如Claude Opus等先进系统处理。
Valued at $15.6 billion, legal AI startup Harvey experienced soaring compute costs following a surge in agent usage, which caused gross margins to plummet from 50% to -50% and forced a reevaluation of its reliance on proprietary giants like OpenAI and Anthropic. In response, Harvey launched a custom model powered by Moonshot AI's Kimi K3, matching top-tier performance at a fraction of the expense and restoring positive margins—a cost-saving strategy backed by prominent venture firms like Sequoia Capital and General Catalyst.
A growing cohort of startups, including Abridge, Decagon, Ramp, Rogo, and Cursor, are similarly developing bespoke models and leveraging open-weight architectures to cut operational expenditures and gain greater technological sovereignty. However, this pivot to open-weight solutions—particularly those originating from China—has introduced complex challenges regarding data privacy, potential regulatory restrictions, and heightened client scrutiny regarding distillation practices and cybersecurity risks.
Despite the economic incentives, post-training and hosting proprietary models impose steep barriers, including fierce competition for multimillion-dollar engineering talent, the necessity of acquiring vast proprietary datasets, and expensive infrastructure overhead. Consequently, building custom models is not universally viable for every enterprise, leading many emerging startups to maintain a hybrid strategy that relies on open models for bulk efficiency while reserving premier closed systems for the most demanding tasks.