从价值对齐到智能正义:生成式人工智能价值风险的哲学审视
From Value Alignment to Intelligent Justice: A Philosophical Examination of the Value Risks of Generative Artificial Intelligence
摘要: 生成式人工智能的价值对齐方案试图通过技术手段确保AI系统与人类价值观保持一致,但方案悬置了一个根本问题,即对齐何种价值观,又以何种程序实现对齐。从批判理论视角审视,当前主流价值对齐方案并非中立性的技术完善,而是技术资本主义语境下的资本逻辑主导下,对马克思所揭示的“一般智力”进行的价值嵌入,它将特殊的、服务于特定群体利益的价值取向伪装成普遍适用的人类共同价值,由此生成的价值风险具有三重维度。认识论层面的社会认知结构性偏差;价值论层面的技术理性对人文价值的排挤;存在论层面的人的主体性风险。若要超越这一困境,就需要从价值对齐转向智能正义。因而智能正义围绕AI设计、部署与收益分配,建立能够抵抗资本与权力扭曲的价值批判和制度反思框架。所以智能正义不是对价值对齐的简单否定,而是对其深层权力关系的揭示、批判与重构。
Abstract: The value alignment scheme of generative artificial intelligence attempts to ensure that AI systems align with human values through technological means, but it leaves a fundamental question unresolved: which values to align with, and through what procedures to achieve alignment. From the perspective of critical theory, the current mainstream value alignment schemes are not neutral technological improvements but rather the embedding of values under the dominance of techno-capitalist logic into what Marx revealed as general intelligence. They disguise the particular and value-oriented interests serving specific groups as universal and common human values, thereby generating value-related risks with three dimensions: social cognitive structural deviations at the epistemological level; the exclusion of human values by technical rationality at the axiological level; and the risk to human subjectivity at the ontological level. To overcome this predicament, it is necessary to shift from value alignment to intelligent justice. Intelligent justice focuses on the design, deployment, and revenue distribution of AI, establishing a framework of value criticism and institutional reflection that can resist the distortion of capital and power. Intelligent justice is not a simple negation of value alignment but rather the revelation, criticism, and reconstruction of its deep power relations.
参考文献
|
[1]
|
海德格尔. 追问技术[M]//查常平. 人文艺术: 第1辑. 成穷, 译. 贵阳: 贵州人民出版社, 1999: 285.
|
|
[2]
|
吴静. 人工智能价值观的动态适应对齐研究——从意图-价值-情境互构范式到智能正义[J]. 华中科技大学学报(社会科学版), 2025, 39(6): 1-10.
|
|
[3]
|
赵泽林. 《1857-1858年经济学手稿》与人工智能的三重审思[J]. 南京社会科学, 2024(5): 30-36+48.
|
|
[4]
|
马克思. 资本论: 第1卷[M]. 中共中央马克思恩格斯列宁斯大林著作编译局, 译. 北京: 人民出版社, 2018: 88-99.
|
|
[5]
|
郭良婧. 人工智能时代道德他律的悖论性生产——Anthropic对齐实验的伦理学批判[J]. 电子科技大学学报(社会科学版), 2026, 28(3): 12-20.
|
|
[6]
|
马克思, 恩格斯. 马克思恩格斯文集: 第8卷[M]. 北京: 人民出版社, 2009: 184.
|
|
[7]
|
马尔库塞. 单向度的人: 发达工业社会意识形态研究[M]. 刘继, 译. 上海: 上海译文出版, 2006: 135.
|
|
[8]
|
Anthropic (2023) Claude’s Constitution. https://www.anthropic.com/news/claudes-constitution
|
|
[9]
|
Anthropic (2023) Collective Constitutional AI: Aligning a Language Model with Public Input. https://www.anthropic.com/news/collective-constitutional-ai-aligning-a-language-model-with-public-input
|
|
[10]
|
OpenAI (2024) Democratic Inputs to AI Grant Program: Lessons Learned and Implementation Plans. https://openai.com/index/democratic-inputs-to-ai-grant-program-update/
|
|
[11]
|
马克思, 恩格斯. 马克思恩格斯文集: 第2卷[M]. 中共中央马克思恩格斯列宁斯大林著作编译局, 译. 北京: 人民出版社, 2009: 53.
|