Cedana推出了一款令牌到GPU计算器,帮助AI从业者快速评估所需的硬件资源1。该工具通过输入令牌量和模型类型,可自动计算所需的GPU数量1。
以月均1万亿令牌处理为例,在Llama 3.3 70B模型上需要约367块H100 GPU1。根据计算逻辑,1万亿令牌在30天内换算为每秒约385,802个令牌1。在50%利用率下,每块H100可处理1,000令牌/秒1。此外,B200芯片相比H100在大规模混合专家模型上的性能提升可达4到21倍1。需要注意的是,该工具的规划估计在各个方向可能存在2倍以上的误差1。
Cedana has introduced a token-to-GPU calculator tool designed to help AI practitioners quickly estimate the hardware resources required for their workloads 1. The calculator accepts token volume and model type as inputs, then computes the necessary number of GPUs to handle the specified throughput 1.
Using one trillion tokens per month as a benchmark case, the tool demonstrates that running the Llama 3.3 70B model would require approximately 367 H100 GPUs over a 30-day period 1. This calculation breaks down to roughly 385,802 tokens per second when dividing one trillion tokens by 2,592,000 seconds in a month, and assumes each H100 can process 1,000 tokens per second at 50% utilization 1. The calculator also highlights performance comparisons across different GPU generations, noting that the B200 can deliver 4 to 21 times better performance than the H100 on large mixture-of-experts models 1. Users should be aware that capacity planning estimates derived from such tools can have error margins of 2x or greater in either direction 1.
评论
还没有评论,欢迎留下第一条。