中国科学院文献情报中心发起的语料创新生态联盟近日在北京启动1。该联盟发布了五类42项科技语料标准(采用三级架构)、敖仓·语料社区服务平台及英文期刊《Data Express in AI Corpus》等系列成果1。联盟首批吸纳了50家共建单位,覆盖国家实验室、科研院所和人工智能企业1。
中国工程院院士孙凝晖表示,大模型要"读懂"科学,靠的不是通用互联网数据,而是高质量科技语料1。该联盟致力于解决科技领域语料分散、专业语义复杂等问题,为人工智能发展提供基础性支撑1。敖仓·语料社区服务平台将打通正式语料与社区语料的双轨供给渠道1。联盟将重点聚焦联合共建高质量科技语料、推进智能就绪科技语料集成服务等五大任务1。
A corpus innovation ecosystem alliance initiated by the Chinese Academy of Sciences' Literature and Information Center was launched in Beijing, unveiling a suite of foundational resources designed to equip artificial intelligence systems with specialized scientific knowledge 1. The initiative released five categories of 42 scientific corpus standards employing a three-tier architecture, alongside the Aocang Corpus Community Service Platform and the English-language journal Data Express in AI Corpus 1.
The alliance brings together 50 founding member institutions spanning national laboratories, research institutes, and artificial intelligence enterprises 1. According to Sun Ninghuai, an academician of the Chinese Academy of Engineering, large language models require access to high-quality scientific corpora rather than generic internet data to effectively comprehend scientific concepts 1. The platform will establish dual supply channels linking formal corpora with community-generated linguistic resources 1, addressing persistent challenges of fragmented data sources and complex domain-specific semantics in scientific literature.
The alliance has prioritized five core objectives, including collaborative development of high-quality scientific corpora and advancement of intelligent readiness for integrated scientific corpus services 1. This structural approach aims to provide foundational technological support for artificial intelligence development by enabling systems to conduct scientific reasoning and knowledge discovery grounded in vetted, discipline-specific language resources 1.
评论
还没有评论,欢迎留下第一条。