Benzi是一款代码智能工具,在SWE-bench Verified基准测试中展示了其性能表现1。该工具在500个真实GitHub问题的测试中解决了78.2%,单个问题的处理成本低于10美分1。测试涵盖了24个GitHub issues和10种编程语言1,对标对象包括Claude Code和CodeGraph等代码智能方案1。
Benzi在成本效益上显示出优势1。根据对比数据,DeepSeek系列模型的成本约为Claude Sonnet的二十分之一1。在难度递增的问题处理上,Claude Code的成本增长斜率最陡,而Benzi展现出相对更缓和的成本增长曲线1。
Benzi, a code intelligence platform, has released benchmark results showing its capability to solve real-world GitHub issues at a fraction of the cost of competing solutions 1. On the SWE-bench Verified benchmark, Benzi resolved 78.2% of 500 authentic GitHub problems at a cost below 10 cents per issue 1. The evaluation encompassed 24 GitHub issues across 10 programming languages, positioning Benzi against established code intelligence tools including Claude Code and CodeGraph 1.
The performance comparison reveals notable cost advantages across multiple dimensions. DeepSeek series models, which power Benzi's infrastructure, operate at approximately one-twentieth the cost of Claude Sonnet 1. When analyzing scaling behavior, Claude Code exhibits the steepest cost growth trajectory as problem difficulty increases, while Benzi demonstrates more moderate cost escalation under similar conditions 1. These metrics—measured across code lines produced, execution time, and overall expenses—highlight Benzi's efficiency in handling progressively complex coding tasks 1.
评论
还没有评论,欢迎留下第一条。