Opper推出了Jevman基准测试项目,用于评估多个AI决策模型在实时游戏场景中的表现和低延迟性能1。该项目让包括OpenAI的o1、Cloudflare的Clef、GPT-4o等多个模型在Pac-Man游戏中竞技,每个模型进行100场游戏测试1。
该项目已开源发布在GitHub上,用户可以运行自己的模型加入排行榜进行排名1。平台为所有用户提供了免费额度来尝试这项功能1。除了让模型相互对战外,用户还可以作为玩家与AI控制的幽灵对手进行游戏,每局游戏的成本约为2美分1。
Opper has launched Jevman, an open-source benchmark project designed to evaluate the real-time decision-making capabilities and low-latency performance of multiple AI models in the classic arcade game Pac-Man 1. The platform tests various decision-making systems, including OpenAI's o1, Cloudflare's Clef, GPT-4o, and several others, each playing 100 games to assess their performance under interactive constraints 1.
The benchmark allows users to run their own models and compete on a shared leaderboard, with each game costing approximately two cents to play 1. Players also have the option to take on the role of Pac-Man and face off against AI-controlled ghosts 1. The project is fully open-sourced on GitHub at https://github.com/opper-ai/jevman-benchmark/, and the platform provides free credits for all users to experiment with different models 1.
评论
还没有评论,欢迎留下第一条。