人工智能训练数据池中的ODbL协议合规性研究
Research on ODbL Protocol Compliance in AI Training Data Pool
摘要: 开源协议是开源大模型生态治理中的关键要素,对其认识不足、履行不当既可能引发知识产权侵权风险,也可能陷入投入产出失衡的“合规陷阱”,亟需合规指引。随着人工智能技术的飞速发展,高质量、大规模的训练数据成为核心竞争力。数据池作为汇聚多源数据的重要载体,其构建与使用过程中的法律合规性问题日益凸显。开放数据库许可协议(ODbL)作为一种“著佐权”式的数据许可协议,旨在确保数据的共享与后续使用的开放性,其在人工智能数据领域的应用引发了新的法律挑战。
Abstract: Open-source licenses are a critical element in the governance of the open-source large model ecosystem. Insufficient understanding or improper implementation of these licenses can not only lead to intellectual property infringement risks but may also result in a “compliance trap” where input and output are unbalanced, creating an urgent need for compliance guidance. With the rapid advancement of artificial intelligence technology, high-quality, large-scale training data has become a core competitiveness. Data pools, as important carriers for aggregating multi-source data, increasingly face prominent legal compliance issues in their construction and usage. The Open Database License (ODbL), as a “copyleft” style data license agreement, aims to ensure the openness of data sharing and subsequent use. Its application in the field of artificial intelligence data has sparked new legal challenges.
文章引用:吴若. 人工智能训练数据池中的ODbL协议合规性研究[J]. 争议解决, 2026, 12(8): 77-82. https://doi.org/10.12677/ds.2026.128239

参考文献

[1] 郭亚军, 徐苑茜, 梁艳丽, 等. 从ChatGPT到DeepSeek: 生成式人工智能迭代对图书馆的影响[J]. 图书馆论坛, 2025, 45(7): 140-149.
[2] 喻玲, 邵滨. 开源社区知识产权治理模式及变革——基于36个开源社区使用协议的考察[J]. 科学学研究, 2024, 42(9): 1938-1945.
[3] 希瑟·米克. 商业开源: 开源软件许可证实用指南[M]. 刘伟, 译. 北京: 人民邮电出版社, 2021: 4.
[4] 朱鸿军, 王涛. 开源许可协议: 人工智能作品保护的另一种制度安排[J]. 探索与争鸣, 2025(12): 136-145+211.
[5] 王利明. 生成式人工智能侵权的归责原则与过错认定[J]. 中国法律评论, 2025(4): 15-30.
[6] 黄如花, 李楠. 开放数据的许可协议类型研究[J]. 图书馆, 2016(8): 16-21.
[7] 王丹. 开源大模型嵌入档案数据空间的基本架构、安全风险与规范路径[J]. 情报科学, 2025, 43(11): 20-26+117.