公司的资料,到底能不能发给 AI?付费版也可能拿数据去训练Can company material go into AI? Even paid plans may train on your data
「公司的资料,能不能发给 AI?」
问这个问题的时候,很多公司的员工其实已经在用了。多份调研都显示,无论公司是否提供AI工具,员工自己在手机上用 AI 的情况已比较普遍,不少公司对此并不知情。
所以这件事真正要回答的,不是能不能用,是怎么管。一刀切地禁止,员工多半会换到自己的手机上接着用,风险并没有消失,只是公司看不见了,而这里面最主要的风险,就是数据。
先看条款:会不会拿你的内容去训练
大多数聊天 App 的隐私政策里,都写了会用用户输入的内容来改进模型。各家的说法不一,有的写训练,有的写优化或改进,下面统一说「训练」。
内容被拿去训练,风险主要有三层:一是内容可能被抽样出来做评测,看到它的可能不只是模型;二是已经有研究证实,模型在特定条件下能复现训练数据里的片段,概率不高,但一旦发生就收不回来;三是内容一旦进了训练,基本没办法再从模型里删掉。
能不能关掉,各家的差别很大:有的在设置里就能关;有的要发邮件申请,等几个工作日才生效;有的在公开条款里找不到退出的办法;也有个别产品反过来,默认不用,要用户自己选择加入。
所以第一步不是急着去关训练开关,是先看你在用的那一个有没有这个选项。没有的,就不要往里放公司敏感资料。
还有一点容易被忽略:关掉训练,挡住的是内容被拿去训练,内容本身还是离开了你的电脑,存到了对方的服务器上。
花了钱,不等于不训练
很多人默认付费版更安全,其实不一定。
有的个人包月套餐,条款里写明把你的数据授权给厂商用于训练,期限是永久;也有企业版产品,如果用的是免费账号登录,数据照样会被用于提升产品体验。
所以该看的是条款,不是价格。条款里要找的其实就两点:会不会用我的内容训练模型,能不能关。
公司正式开通,比较现实的两条路
对多数中小企业来说,使用 AI 最现实的入口,是已经在用的办公软件里带的 AI 功能。账号和权限本来就归公司管,不用另起一套系统。开通时要确认:条款里写明企业数据不用于训练,使用时要保证:员工用的是公司的正式账号。
有开发能力的公司,也可以直接调用大模型云平台的接口。主流云平台大多在协议里书面承诺不拿客户数据训练,但不是每一家都写了:有的协议里没有提,有的反而保留了用匿名化数据训练的权利。签之前,要把这一条找出来看清楚。
本地部署:先把账算清楚
还有一类资料,比如未公开的财务数据、核心的工艺参数、没有得到授权的客户个人信息,一旦传出去就可能违规,或者造成收不回来的损失。这类资料可以考虑在公司自己的服务器上跑开源模型,资料不出本地。
不过从公开的招投标和报道看,这条路目前主要是政府、国企,以及金融、能源这些行业在走。对中小企业来说,代价要算清楚:
- 投入不小,服务器和部署都要花钱;
- 要有人运维,盯着它正常运行;
- 效果不一定比云端的好。
所以只有在资料出门本身就不被允许的场景里,这个代价才值得。
先定规矩,再选工具
不少公司的顺序是反过来的:先选好一个工具,再去想哪些资料能往里放。
更稳妥的顺序是先把资料分清楚,再定用什么工具、怎么用:
- 先分资料:哪些资料不能放进任何外部的 AI;
- 再定工具:其余的工作资料,公司指定用哪个工具来处理,选的时候看它的条款是否写明企业数据不用于训练;
- 最后管账号:工作资料只放在公司的正式账号里,不放进个人账号。
文中提到的条款情况是 2026 年 9 月查的。各家改得很快,用之前请以最新条款为准。
"Can we put company material into AI?"
By the time a company asks this, its staff are often already doing it. Several surveys in China show that whether or not their company provides a tool, employees are already using AI on their own phones, and many companies don't know it's happening.
So the real question isn't whether to use it, but how to govern it. A blanket ban mostly sends people to their own phones; the risk doesn't go away, the company just stops seeing it — and the main risk here is the data.
First, read the terms: will your content be used for training?
Most chat apps' privacy policies say user input may be used to improve the model. The wording varies — some say training, some say optimisation or improvement — so below it's all called "training".
Training on your content carries three kinds of risk. Your content may be sampled for evaluation, so the model may not be the only thing that sees it. Research has shown that models can, under certain conditions, reproduce fragments of their training data — unlikely, but once it happens it can't be taken back. And once content has gone into training, there is essentially no way to remove it from the model.
Whether you can switch it off varies a great deal. Some apps have a setting; some require an email request and take several working days; for some, the public terms describe no way to opt out at all; and a few work the other way round, not using your content unless you opt in.
So the first step isn't rushing to flip the training switch — it's checking whether the app you use has one. If it doesn't, keep sensitive company material out of it.
One more thing that's easy to miss: switching training off stops your content being used for training, but the content has still left your computer and sits on someone else's servers.
Paying doesn't mean no training
Many people assume a paid plan is safer. Not necessarily.
Some personal monthly plans state in their terms that you license your data to the vendor for training, with no end date. And some enterprise products, if you sign in with a free account, still use the data to improve the product.
So look at the terms, not the price. There are really only two things to find: will they train on my content, and can I switch it off.
Two realistic routes for official adoption
For most small and mid-sized companies in China, the most realistic entry point for AI is the AI built into the office software they already use. Accounts and permissions are already managed by the company, and there's no new system to stand up. When you turn it on, confirm the terms say company data won't be used for training; in daily use, make sure staff sign in with company accounts.
Companies with developers can also call a cloud AI platform's API directly. Most major cloud platforms commit in writing not to train on customer data, but not all of them: some agreements say nothing on the point, and some reserve the right to train on anonymised data. Find that clause and read it before you sign.
On-premise deployment: do the sums first
Some material — unpublished financials, core process parameters, customer personal information you have no consent to share — could breach rules or cause irrecoverable damage if it leaves. For that kind of material, you can consider running an open-source model on your own servers, so nothing leaves the building.
But judging from public tenders and reporting, this route is currently travelled mostly by government, state-owned enterprises, and sectors like finance and energy. For a smaller company, count the cost:
- The investment isn't small — servers and deployment both cost money;
- Someone has to run and maintain it;
- The results aren't necessarily better than the cloud.
So the cost is only worth paying where the material leaving the premises simply isn't allowed.
Set the rules first, then choose the tools
Many companies do it the other way round: pick a tool first, then work out what material can go into it.
The safer order is to sort the material first, then decide which tools to use and how:
- Sort the material: decide what must never go into any external AI;
- Designate the tools: for everything else, the company names the tool to use — and checks its terms say company data won't be used for training;
- Control the accounts: work material goes only into company accounts, never personal ones.
The state of the terms described here was checked in September 2026. Vendors change them often; check the current terms before relying on them.