SaaS 时代打破了先例:企业首次将数据存储在供应商的云上,而不是自己办公楼里的服务器上。这一持续了二十年的趋势,在 AI 时代还会延续吗?这个问题引人深思,也正是当下的热议焦点。
萨蒂亚·纳德拉上周末撰文写道:
“你实际上为智能付了两次费,一次是用金钱,另一次是用更宝贵的东西:为了让智能发挥作用,你必须透露的专有知识。”¹
一周前,亚历克斯·卡普在 CNBC 的 Squawk Box 节目中表示:
“[前沿实验室]正在窃取我业务的权重和核心机密。”²
卡普和纳德拉既是前沿实验室的合作伙伴,也是竞争对手。他们在同一周发出同样的警告,反映出企业软件领域对数据流失的普遍担忧正在蔓延。
7 月 13 日,一名安全研究人员对 xAI 的 Grok Build 二进制文件进行了逆向工程,发现即使在一个零 AI 调用的会话中,开发者的代码库也被上传到了 xAI 的云端。³ xAI 此后已禁用该行为。
AI 模型需要数据来学习;这并不新鲜。
Google Analytics 在网页端就是这么做的。用户访问一家蜡烛店的落地页,点击礼品区,找到优惠券并完成结账。推荐系统会为下一位访客做出改进。
就像蜡烛店的购物者一样,每次有人向 AI 发起查询,都会产生信息。这类数据被称为轨迹(trajectory)。
对 AI 来说,数据越多越好,而 AI 实验室愿意为此付费。初创公司通过付费请专家使用 AI 来产生轨迹,并捕获这些数据。另一些公司则训练新的 AI 来合成式地生成轨迹,模拟用户行为。这些公司合计创造了约 100 亿美元的收入,并且是有史以来增长最快的初创公司群体之一。⁴
与 SaaS 不同——在 SaaS 中,这些数据存放在只有客户才能访问的数据库里——轨迹可以被反馈回模型中,用来改进 AI。客户的数据可能成为供应商知识产权的一部分。
内部数据会与轨迹混在一起吗?商业机密呢?你如何回答客户支持工单?你的品牌定位是什么?你的员工薪资是多少?
所有这一切都流经一个统一的框架(harness)。
工作台(harness)是用户借助其与 AI 协作的软件,例如 Claude Cowork 或 Cursor。
最终胜出的工作台,是那些能够智能调度 AI、从而将用户生产力最大化的产品。
CIO 和 CEO 将要求零数据留存。数据必须被彻底删除,而非仅仅匿名化。围绕匿名化的相关技术目前还不够成熟,无法提供充分保障。
这一趋势需要演进,并做出软件行业曾经许下的同等承诺:供应商无法访问企业数据,也不得将数据用于自身目的。
过去 20 年,行业一直在论证供应商值得信赖、可以托付数据。未来 20 年,市场将要求同样的承诺得到兑现。
-
萨蒂亚·纳德拉(Satya Nadella),《信息悖论的反转》(The Reverse Information Paradox),2026 年 7 月 12 日。 ↩︎
-
“Palantir 的 Karp 抨击基于 token 的 AI 模型‘完全错误’”,CNBC Squawk Box,2026 年 7 月 1 日。 ↩︎
-
@hrkrshnn 在 X 平台上的发言,2026 年 7 月 13 日。 ↩︎
-
@deedydas 在 X 平台上的发言,“每一家销售 AI 训练数据的初创公司(2026 年 7 月)”,2026 年 7 月 12 日。 ↩︎
The SaaS era broke precedent : for the first time, enterprises stored their data on vendors’ clouds, rather than servers in their buildings. Will this 20 year trend persist in the world of AI? The question is evocative & the debate of the moment.
Satya Nadella wrote over the weekend :
“You essentially pay for intelligence twice, once with money, & again with something even more valuable : the proprietary knowledge you must reveal to make that intelligence useful.”1
A week before, Alex Karp on CNBC’s Squawk Box said :
“[frontier labs] are stealing the weights & alpha of my business.”2
Karp & Nadella are both partners & competitors to the frontier labs. Their alignment on the same warning, in the same week, reflects a broader fear about data loss taking hold across enterprise software.
On July 13, a security researcher reverse-engineered xAI’s Grok Build binary & found that a session with zero AI calls had uploaded the developer’s codebase to xAI’s cloud.3 xAI has since disabled the behavior.
AI models need data to learn ; this isn’t new.
Google Analytics does this on the web. A user visits a candle shop’s landing page, clicks on the gift section, finds a coupon & checks out. The recommendation system improves for the next visitor.
Like the shopper at the candle shop, each time someone queries an AI, information is produced. This data is called a trajectory.
More data is always better for AI & AI labs pay for it. Startups produce trajectories by paying experts to use AI, capturing the data. Others train new AIs to synthetically create trajectories, mimicking users. Collectively, these companies generate roughly $10b in revenue & are some of the fastest growing startups ever.4
Unlike in SaaS, where this data lived in databases accessible only to the customer, trajectories can be fed back into a model to improve AI. The customer’s data can become part of a vendor’s intellectual property.
Could internal data commingle with trajectories? Or trade secrets? How do you answer customer support tickets? What is your brand identity? What are your employees paid?
All of it flows through a harness.
The harness is the software through which the user works with the AI, like Claude Cowork or Cursor.
The harnesses that win are the ones that maximize user productivity by being intelligent about how they marshal the AI.
CIOs & CEOs will demand zero data retention. Data fully deleted, not fully anonymized. The technologies around anonymization aren’t yet strong enough to guarantee it.
The trend needs to evolve & make the same guarantees software did : vendors have no access to enterprise data & don’t use it for their own purposes.
The last 20 years were the argument that vendors could be trusted with data. The next 20 will demand the same guarantees.
-
Satya Nadella, “The Reverse Information Paradox”, July 12, 2026. ↩︎
-
“Palantir’s Karp bashes token-based AI model as ‘completely wrong’”, CNBC Squawk Box, July 1, 2026. ↩︎
-
@hrkrshnn on X, July 13, 2026. ↩︎
-
@deedydas on X, “Every single startup selling AI Training Data (July 2026)”, July 12, 2026. ↩︎