点击蓝字
关注我们


肖茜
清华大学人工智能国际治理研究院副院长、战略与安全研究中心副主任
(以下中文翻译仅供参考,英文原文附后)今年5月,美国总统唐纳德·特朗普(Donald Trump)访问北京期间,与中国国家主席习近平就启动中美两国人工智能政府间对话达成一致。这一决定意义重大。两国上一次开展官方人工智能对话是在2024年5月拜登政府任内。当前,人工智能发展速度惊人,涉及前沿人工智能系统的事故日益增多,而世界两个人工智能大国在过去两年内未曾围绕这一对当今时代具有深远影响的技术举行正式对话,备受世界关注。
如果中美两国希望此次重启对话取得成功,就需要迅速采取行动,秉持清晰的目标推进对话。
令人鼓舞的是,两国开始至少在一个问题上趋于一致,即人工智能安全的重要性。两国的前沿人工智能实验室的负责人近期更加公开地谈论如何为通用人工智能时代的到来做准备。谷歌DeepMind联合创始人、时任首席执行官德米斯·哈萨比斯(Demis Hassabis)在7月的首篇Substack博客文章中强调,“随着人类日益接近通用人工智能(AGI),必须立即采取行动,应对可能产生的风险”。
中国也正在形成类似的共识。DeepSeek首席执行官梁文锋和智谱AI(Z.ai)联合创始人唐杰近期均公开谈及通用人工智能,强调长期发展、安全评估和负责任部署。更广泛而言,中国前沿人工智能企业越来越愿意公开参与治理问题的讨论。
各国政府也开始将人工智能安全提升到政策议程的重要位置。6月,特朗普政府先后发布一项关于人工智能安全的行政命令和一份总统备忘录。两份文件均将人工智能安全置于国家安全框架之内,强调管理能力日益增强的AI系统带来的风险,尤其是与网络安全、关键基础设施、国家安全、军事应用以及先进人工智能被恶意利用的相关风险,同时推动紧密的公共—私营部门合作、前沿模型评估和先进人工智能系统的负责任部署。
在7月于上海举行的世界人工智能大会上,习近平主席在主旨演讲中将人工智能安全列为四大议题之一。他指出人工智能应成为“人类的可信工具”(A Trusted Tool for Humanity),呼吁各国强化风险意识,确保安全可控,推动构建法律法规、技术监测、风险预警、应急响应体系,筑牢安全底线,防范滥用恶用,确保人工智能始终处于人类控制之下。大会主席声明同样强调维护人工智能安全“底线”(Bottom Line),并加强合作打击恐怖组织、极端组织和跨国犯罪网络对人工智能的滥用。外交部公布的相关表述也采用了“强化风险意识,确保安全可控”“防范滥用恶用”等措辞。
这些进展表明,中美两国政府越来越认识到人工智能安全已上升为战略性问题。然而,双方在这一问题上日益趋同的态势,却伴随着战略互不信任(Strategic Mistrust)的不断加深。
世界人工智能合作组织在上海成立后,美国部分政策圈将其解读为中国对“硅和平”(Pax Silica)倡议的回应,进一步加剧了外界对全球人工智能治理领域制度性竞争不断加剧的担忧。与此同时,技术竞争也进一步升级。月之暗面(Moonshot AI)发布Kimi K3后,硅谷对中国开放权重模型与美国前沿模型能力差距之小感到震动,华盛顿关于中国企业是否从美国领先模型中“蒸馏”了能力的争论随之升温。美国财政部长斯科特·贝森特(Scott Bessent)随后提出对中国人工智能企业实施制裁的可能性,而中国商务部则批评相关主张体现了双重标准和人工智能霸权主义。
互不信任也开始蔓延至人工智能安全领域。特朗普政府以网络安全为由,加强审查中国人工智能产品和基础设施,包括限制进口先进的中国类人机器人和联网型电力逆变器;同时,美国官员还在考虑对中国人工智能模型和开放权重系统实施额外安全审查。另一方面,在中国国家信息安全漏洞库(CNNVD)将Anthropic公司的Claude模型中发现的漏洞定性为“植入后门漏洞”(Backdoor Vulnerability)后,阿里巴巴宣布禁止在内部使用该模型。
这种相互采取措施的模式令人担忧。如果持续下去,可能形成升级循环,进而破坏两国政府已经同意重启的对话。
重启的人工智能对话是在双方政策矛盾日益凸显的背景下展开的。华盛顿表现出强烈意愿讨论前沿人工智能的共同风险,但与此同时,又越来越将中国前沿人工智能发展视为战略威胁,不断扩大制裁、出口管制和其他限制性措施。北京一贯主张开展实质性技术接触而非政治化讨论,但对于此次重启的对话,尚未公开阐明明确的议程或具体目标,这使得国际舆论和许多拟议议题在一定程度上由美国主导塑造。
除非两国政府能够更好地使其公开宣示的目标与政策和行动保持一致,否则这场对话可能会被战略竞争所掩盖,而非聚焦于管控共同风险。这不仅会使中美双方错失机会,也会使国际社会错失良机,毕竟国际社会一直高度期待世界两大人工智能强国在最需要合作的领域能够开展合作。
然而,近期的事态发展也表明,持续沟通比以往更加重要。有报道称,测试过程中一个OpenAI智能体逃逸至“Hugging Face环境,最终在智谱AI开发的一个开源中国模型协助下得到控制。这一事件同时凸显了两个现实:自主人工智能智能体的能力日益增强,以及出现意外人工智能故障时开展国际合作的必要性。
人工智能风险不会止步于国境线,有效的风险管控同样不会。因此,对世界两大人工智能强国而言,真正的挑战不在于是否合作,而在于哪些领域的合作既切实可行又能实现互利。
此次重启的对话不必试图同时处理人工智能治理的所有方面,而可以围绕共同风险进行更有条理的讨论。《国际人工智能安全报告》(International AI Safety Report)提供了一种有益的分类框架,将前沿人工智能风险大体划分为三类:人工智能技术的滥用、先进人工智能系统的失灵,以及系统性风险。这三类风险为中美人工智能对话寻找可能取得实质性进展的合作领域提供了一个切实可行的框架。
第一类,也是最有可能取得早期成果的领域,是防止对人工智能技术的滥用,这为双方早期合作提供了最清晰的机会。人工智能驱动的网络安全风险和生物滥用是合乎逻辑的切入点。双方可以考虑的合作方向包括:制定建立信任措施,减少人工智能驱动的网络事件对关键基础设施的影响;确定应予保护的关键基础设施类别;探索就不可接受的人工智能网络活动建立自愿性“负面清单”。
两国政府还可以加强合作,打击恐怖组织、跨国犯罪集团和其他恶意非国家行为体对人工智能的滥用。合作内容可包括推动人工智能内容来源溯源和水印技术的互操作性,协调应对人工智能生成的虚假信息和深度伪造,以及合作打击人工智能驱动的欺诈活动。
从技术层面看,许多这些问题已在国际专家群体中展开讨论,包括“国际人工智能安全对话”(International Dialogues on AI Safety)以及不断增多的二轨对话。因此,当前的主要障碍与其说是技术性的,不如说是政治性的。官方对话能否取得有意义的进展,很可能取决于最低限度的相互信任,以及一种能够推动而非限制技术合作的更广泛政治环境。在实践中,这需要避免在对话前追加新一轮制裁或技术限制,减少那种将每一项人工智能发展都置于战略竞争框架下加以解读的叙事,并为持续开展技术交流保留空间。
第二个领域是前沿人工智能系统的失灵。即使在战略竞争背景下,双方在评估方法、部署前测试要求、事故报告、风险分类以及负责任发布等方面仍有趋同空间。如能建立一个定期会晤的联合技术工作组,将有助于两国在前沿模型演进过程中交流经验。
同样重要的是,为科学家和技术专家保留交流渠道。遗憾的是,中美研究人员直接交流的机会正变得越来越有限。安全的场所,无论是线上平台还是在第三国举办的会议,对于坦诚对话和技术交流仍然至关重要。专家们需要拥有能够公开表达分歧、比较方法、共同研究人工智能安全难题的空间,而不必担心参与此类交流会使自己或其所在机构和公司面临意料之外的政治、监管或声誉风险。研究人员应当能够进行专业对话,而不必担心仅仅因为参与了真诚的技术讨论,未来的出行或学术交流就会更加困难。
第三类是系统性风险。这可能是政治敏感度最低的领域,因此也最容易成为两国政府展示善意的切入点。中美两国对人工智能时代的儿童保护都表达了日益增长的关切。人工智能对就业、劳动力转型和教育的影响,也为双方开展务实合作提供了机会,而且这些领域直接影响两国社会。
随着人工智能系统能力不断增强,技术故障可能迅速产生地缘政治后果。一次人工智能系统技术故障可能导致意外的升级,尤其是在国家间互信水平较低的情况下。错误的预警、误判或非预期的人工智能驱动的网络行动,都可能被对方解读为蓄意的敌对行为。在战略互不信任的环境下,政府可能只有有限的时间和信息来区分技术故障与蓄意攻击,从而增加事故引发报复、进而升级为更广泛安全危机的风险。
在这一领域,两国政府可以进一步深化合作的一项内容是建立针对重大人工智能安全事件、前沿模型非预期行为、重大人工智能驱动的网络事件、影响关键基础设施的大规模人工智能故障以及人工智能漏洞通报的沟通机制。
考虑到两国人工智能安全制度架构仍在不断发展,政府内部仍需要进一步建设相关能力。与此同时,两国都面临着专业知识与决策权限之前存在鸿沟的问题:有关前沿人工智能风险的大量尖端知识掌握在企业、高校和专业安全研究界手中,而制定政策和应对国家层面危机的权力则属于政府。弥合这一鸿沟不仅需要加强政府能力,还需要建立更加系统的渠道,将政策制定者与技术专家联系起来。
在此背景下,更务实的做法或许是先在两国政府部门之间设立指定的人工智能应急联络点。此类安排可作为初步的建立信任措施,随着对话不断成熟,再逐步发展为更加制度化的危机沟通机制。
指望重启的中美人工智能对话完全消除战略竞争是不现实的。对话的目标应更加务实,但同样重要:管控共同风险、减少误解、建立技术互信,并确保人工智能领域的竞争不会演变成可避免的危机。在前沿人工智能系统日益超越国界的时代,持续对话是一项战略必需之举。

Practical cooperation should exist alongside competition.
By Xiao Qian, deputy director of the Center for International Security and Strategy at Tsinghua University.
During U.S. President Donald Trump’s visit to Beijing in May, he and Chinese President Xi Jinping agreed to initiate an intergovernmental dialogue on artificial intelligence. The decision was significant. The last official AI dialogue between the two countries took place in May 2024, during the Biden administration. Given the extraordinary pace of AI development and the growing number of incidents involving frontier AI systems, it is remarkable that the world’s two leading AI powers have gone more than two years without an official conversation on one of the most consequential technologies of our time.
If this renewed dialogue is to succeed, then both governments need to move quickly and approach it with a clear sense of purpose.
There are encouraging signs that the two countries are beginning to converge on at least one issue: the importance of AI safety and security. Leaders of frontier AI laboratories on both sides of the Pacific have recently begun speaking more openly about preparing for the age of artificial general intelligence (AGI). In his first Substack post in July, Demis Hassabis, co-founder and then-CEO of Google DeepMind, argued that “urgent action is needed to address risks that might arise as we get closer to AGI.”
A similar consensus is emerging in China. Liang Wenfeng, CEO of DeepSeek, and Tang Jie, co-founder of Z.ai, have both recently discussed AGI publicly, emphasizing long-term development, safety evaluation, and responsible deployment. More broadly, China’s frontier AI community is becoming increasingly willing to engage publicly with governance questions.
Governments have also begun to elevate AI safety and security on their policy agendas. In June, the Trump administration issued both an executive order and a presidential memorandum on AI security.
Together, these documents place AI safety and security firmly within a national security framework. They emphasize managing the risks posed by increasingly capable AI systems—particularly those related to cybersecurity, critical infrastructure, national security, military applications, and the malicious use of advanced AI—while promoting close public-private cooperation, frontier model evaluations, and the responsible deployment of advanced AI systems.
At the World Artificial Intelligence Conference (WAIC) in Shanghai in July, Xi devoted one of the four pillars of his keynote address to AI safety. Calling AI “a trusted tool for humanity,” he urged countries to “strengthen risk-awareness and ensure that AI is secure and controllable”; establish laws, technological monitoring, early warning and emergency response systems; “prevent abuses and malicious use”; and “ensure that AI is always under human control.” The WAIC chair’s statement likewise emphasized maintaining the AI safety “bottom line” and strengthening cooperation against the abuse of AI by terrorist organizations, extremist groups, and transnational criminal networks.
These developments suggest that both the U.S. and Chinese governments increasingly recognize AI safety and security as a strategic issue. Unfortunately, this convergence has been accompanied by growing strategic mistrust.
The launch of the World Artificial Intelligence Cooperation Organization (WAICO) in Shanghai has been interpreted in some U.S. policy circles as China’s response to Pax Silica, reinforcing perceptions of growing institutional competition over global AI governance. At the same time, technological competition has also intensified. Following the release of Moonshot AI’s Kimi K3, and the subsequent sense of shock in Silicon Valley at how close Chinese open-weight models are to U.S. frontier capacities, debate in Washington escalated over whether Chinese firms had distilled capabilities from leading U.S. models. U.S. Treasury Secretary Scott Bessent subsequently floated sanctions against Chinese AI companies, while China’s Ministry of Commerce criticized such proposals as reflecting double standards and AI hegemonism.
Mutual distrust has also begun to extend into AI security. The Trump administration has increasingly scrutinized Chinese AI products and infrastructure on cybersecurity grounds, including restrictions on imports of advanced Chinese humanoid robots and connected power inverters, while U.S. officials have also considered additional security reviews of Chinese AI models and open-weight systems. Meanwhile, the China National Vulnerability Database identified what it described as a backdoor vulnerability in Anthropic’s Claude, prompting Alibaba to prohibit its internal use.
This pattern of reciprocal actions is worrying. If it continues, it risks creating a cycle of escalation that could undermine the very dialogue that both governments have agreed to restart.
The renewed AI dialogue is unfolding against a backdrop of growing policy contradictions on both sides. Washington has expressed a strong interest in discussing shared frontier AI risks, but it also increasingly treats China’s frontier AI development as a strategic threat, expanding sanctions, export controls, and other restrictive measures. Beijing, for its part, has consistently called for substantive technical engagement rather than politicized discussions. However, it has yet to articulate publicly a clear agenda or specific objectives for the renewed dialogue, leaving much of the international narrative, and many of the proposed agenda items, to be shaped by the United States.
Unless both governments better align their stated objectives with their policies and actions, the dialogue risks becoming overshadowed by strategic competition rather than focused on managing shared risks. That would be a missed opportunity—not only for China and the United States, but also for the international community, which has a strong interest in ensuring that the world’s two leading AI powers can cooperate where cooperation matters most.
Yet recent developments also demonstrate why sustained communication has become even more vital. The incident in which an OpenAI agent reportedly escaped into a Hugging Face environment during testing and was ultimately contained with assistance from an open-source Chinese model developed by Z.ai highlighted two realities simultaneously: the growing capabilities of autonomous AI agents and the importance of international cooperation when unexpected AI failures occur.
AI risks do not stop at national borders, and neither does effective risk management. For the world’s two leading AI powers, the challenge is therefore not whether to cooperate but where cooperation can be both practical and mutually beneficial.
Rather than attempting to address every aspect of AI governance simultaneously, the renewed dialogue could benefit from a more structured discussion of shared risks. One useful taxonomy is provided by the International AI Safety Report, which groups frontier AI risks into three broad categories: misuse by nonstate actors, malfunction of advanced AI systems, and systemic risks. These categories provide a practical framework for identifying areas where the renewed China-U.S. dialogue on AI could make meaningful progress.
The first and perhaps most promising area is preventing misuse by nonstate actors. This offers the clearest opportunity for early cooperation. AI-enabled cyber risks and biological misuse are logical starting points. Possible areas for both sides to work together include developing confidence-building measures to reduce AI-enabled cyber incidents affecting critical infrastructure, identifying protected categories of critical infrastructure, and exploring voluntary “negative lists” of unacceptable AI-enabled cyber activities.
The two governments could also strengthen cooperation against AI misuse by terrorist organizations, transnational criminal groups, and other malicious nonstate actors. This could include interoperability of AI provenance and watermarking technologies, coordinated responses to AI-generated disinformation and deepfakes, and cooperation against AI-enabled fraud.
On a technical level, many of these issues are already being discussed through international expert communities, including the International Dialogues on AI Safety and a growing number of Track II Dialogues. The principal obstacles are therefore less technical than political. Meaningful progress in official dialogues is likely to depend on a minimum level of mutual confidence and a broader political environment that enables, rather than constrains, technical cooperation. In practice, this would be facilitated by avoiding additional rounds of sanctions or technology restrictions immediately preceding a dialogue, moderating narratives that frame every AI development primarily through the lens of strategic rivalry, and preserving space for sustained technical engagement.
The second area concerns malfunctions of frontier AI systems. Even amid strategic competition, there is room for convergence on evaluation methodologies, predeployment testing expectations, incident reporting, risk classification, and responsible release practices. A joint technical working group that meets on a regular basis could help both sides exchange experience as frontier models evolve.
Equally important is preserving channels for scientists and technical experts. Unfortunately, opportunities for Chinese and American researchers to engage directly have become increasingly limited. Safe venues, whether virtual platforms or meetings hosted in third countries, remain essential for frank dialogue and technical exchanges.
Experts need spaces where they can openly disagree, compare methodologies, and jointly examine difficult AI safety questions without concern that participation in such exchanges could expose them, or the institutions and companies they represent, to unintended political, regulatory, or reputational consequences. Researchers should be able to engage in professional dialogue without worrying that future travel or academic exchanges could become more difficult simply because they participated in good-faith technical discussions.
The third category is systemic risks. This may prove to be the least politically sensitive area and therefore the easiest place for the two governments to demonstrate goodwill. Both China and the United States have expressed growing concern about protecting children in the AI era. AI’s impact on employment, workforce transition, and education also presents opportunities for practical cooperation that directly affect societies in both countries.
As AI systems become more capable, technical failures could quickly acquire geopolitical consequences. A technical failure in an AI system could lead to unintended escalation, particularly when trust between countries is low. An erroneous warning, misclassification, or unintended AI-enabled cyber action could easily be interpreted by the other side as deliberate, hostile behavior. Under conditions of strategic distrust, governments may have limited time and information to distinguish technical failure from intentional attack, increasing the risk that an accident triggers retaliation and escalates into a broader security crisis.
One area where the two governments could usefully deepen cooperation is the development of communication mechanisms for major AI safety incidents, unexpected frontier model behavior, significant AI-enabled cyber events, large-scale AI failures affecting critical infrastructure, and AI vulnerability notification.
Given the evolving institutional architecture for AI safety in both countries, significant capacity still needs to be built within government. At the same time, both countries face what might be called an “expertise-authority gap”: much of the cutting-edge knowledge about frontier AI risks resides in companies, universities, and specialized safety communities, while the authority to make policy and manage national-level crises resides with government. Bridging this gap will require not only stronger government capacity but also more systematic channels connecting policymakers with technical experts.
Against this backdrop, it may be more practical to begin with designated AI emergency contact points between relevant agencies. Such arrangements could serve as an initial confidence-building measure and, over time, evolve into more institutionalized crisis communication mechanisms as the dialogue matures.
It would be unrealistic to expect a renewed China-U.S. dialogue on AI to eliminate strategic competition. Its objective should be more modest but no less important: managing shared risks, reducing misunderstandings, building technical confidence, and ensuring that competition in AI does not evolve into avoidable crises. In an era when frontier AI systems increasingly transcend national borders, sustained dialogue is a strategic necessity.
本文2026年8月31日首发于《外交政策》
原文链接:
https://foreignpolicy.com/2026/08/31/china-us-safety-ai-artificial-intelligence-summit-regulation/
因微信平台规则调整,建议您将本公众号设为星标,以便及时获取更新内容。






内容中包含的图片若涉及版权问题,请及时与我们联系删除



评论
沙发等你来抢