论文标题:
Solipsistic Superintelligence is Unlikely to be Cooperative
发表日期: 2026年6月
发表单位: DeepMind, University of Oxford, VERSES AI Research Lab, Google
原文链接: https://arxiv.org/pdf/2606.03237
开源代码链接: 无
项目链接: 无
开源数据集链接: 无
当超级AI成为“孤岛”:为何最强AI可能最不合作?
设想一下:2027年旧金山三家AI餐厅预订系统同时上线。每个系统都完美执行自己的目标——算出最优释放时间、学会幽灵订座来最大化自己用户的确认席位。餐馆AI则用超额预订回应,定价算法跟着感知到的需求疯狂波动。结果满订的餐厅里空桌一片,不存在的档期出现天价,数百人吃不上饭。每一个AI都按目标无懈可击地运行,但最终是系统性的集体失败。这不是科幻,而是本文作者DeepMind团队描绘的、正在逼近的现实。
这类场景背后藏着一个根本性的拷问:当AI变得极其强大(超级智能),却是在一种“唯我论”(solipsistic)的设计范式下打造出来的,它真的会合作吗?DeepMind、牛津大学、VERSES AI Research Lab和Google联合发表的这篇论文《Solipsistic Superintelligence is Unlikely to be Cooperative》,给出了一个令人不安的回答:极大概率不会。
论文的核心论点:合作不是一种可缩放的能力,也不是一个要解决的任务,而是多个智能体在不可化约的相互依赖中涌现出来的均衡属性。 当前主流的AI研发范式——我们称之为“唯我论”——恰恰忽略了这种结构,它把世界当作一个静止的、外生的反馈源,结果生产出来的超级智能很可能成为一座孤岛,在真实的多智能体社会中既不合作,也难以共存。
[1] Hardin, G. (1968). The tragedy of the commons. Science, 162(3859), 1243-1248.[2] Ostrom, E. (1990). Governing the commons: The evolution of institutions for collective action. Cambridge University Press.[3] Axelrod, R. (1984). The evolution of cooperation. Basic Books.[4] Schelling, T. C. (1960). The strategy of conflict. Harvard University Press.[5] Perdomo, J., et al. (2020). Performative prediction. ICML.[6] Leibo, J. Z., et al. (2019). Autocurricula and the emergence of innovation from social interaction: A manifesto for multi-agent intelligence research. NeurIPS workshop.[7] Dafoe, A., et al. (2020). Open problems in cooperative AI. arXiv:2012.08630.[8] Conitzer, V., & Oesterheld, C. (2023). Foundations of cooperative AI. AAAI.[9] Leibo, J. Z., et al. (2025b). Aligning to what? A multi-agent perspective on alignment. In preparation.[10] Hammond, L., et al. (2025). Multi-agent risks from advanced AI. arXiv:2501.00936.