法国初创公司Kog正在开发软件优化技术,以在标准数据中心GPU(如NVIDIA H200和AMD MI300X)上实现更快的AI推理速度。[1]公司声称能实现30倍的大语言模型推理加速,并在Laneformer 2B小模型的演示中达到每秒3,000个令牌的处理速度。[1]
Kog团队规模为11人。[1]CEO Gaël Delalleau表示,公司在5月登上Hacker News首页展示了在标准GPU上实现"极快的单请求解码"后,获得了200个商业线索。[1]为了支持新GPU,Kog需要投入数周甚至数月时间进行工程研究。[1]
Delalleau预计公司将于9月实现首个主流大模型的10倍速优化,随后开始展示客户采用成果,以推动融资进展。[1]
French startup Kog is developing software techniques designed to extract faster AI inference performance from standard datacenter GPUs, including NVIDIA H200 and AMD MI300X processors.[1] The company claims it can achieve 30 times faster large language model inference[1] and has accumulated 200 tangible business leads.[1] In a current demonstration running on the Laneformer 2B model, Kog's technology reaches 3,000 tokens per second.[1]
Kog gained attention in May when it appeared on Hacker News' front page, showcasing what it described as "extremely fast single-request decoding" on standard GPUs.[1] The startup operates with a team of 11 people, and supporting each new GPU architecture requires weeks or even months of engineering work.[1] CEO Gaël Delalleau stated that the company expects to have "implemented our first major model at 10x speed" by September, after which it plans to "start demonstrating customer traction and from there, raise our Series A" funding round.[1]