본문으로 건너뛰기
개발 뉴스로
AIHacker News··원문 약 2

Show HN: Benzi – Claude Code 및 CodeGraph를 능가하는 코드 관성/하네스

Show HN: Benzi – A Code Intillegence/Harness Beating Claude Code and CodeGraph

버그가 더 어려워질수록 모든 하네스는 더 많은 소스를 엽니다.

핵심 요약

자동 요약
  1. 1버그가 더 어려워질수록 모든 하네스는 더 많은 소스를 엽니다.
  2. 2문제는 경사입니다.
  3. 3각 포인트는 하나의 버그입니다.

원문 본문

출처 · Hacker News

Benzi · benchmarks

Three ways we've measured it.

Bug-fixing benchmark →

Apples to apples comparison: Benzi harness vs mainstream harnesses, 24 GitHub issues, 10 languages.

SWE-bench Verified →

Benzi harness on SWE-bench Verified. 78.2% of 500 real issues resolved at under 10¢ a fix.

CodeGraph comparison →

Apples to Apples comparison: Code Graph's code intelligence vs Benzi's AI-native code intelligence

See all runs →

Every task, every attempt, verbatim — nothing held back.

Learn more about Benzi: benzi.fly.dev/about

Benzi's KPI (Key Performance Indicator) — source lines read

Every harness opens more source as bugs get harder. The question is the slope. Each point is one bug; the 24 are laid out easiest to hardest, left to right.

Difficulty is Claude Code's turn count on that bug — a third-party yardstick, so no harness sets its own position on the axis. Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The figure beside each line is its slope: how many extra lines that harness opens per step of difficulty. Hover any point for the bug and its count.

source lines read · 24 bugs · lowest of the four marked bug Benzi
Sonnet Benzi
DeepSeek Claude Code
Sonnet DeepSeek Harness
DeepSeek mux64643831,372 commons-cli39235456771 addressable1802402001,603 jsoup6117060616 yaml-cpp120383194461 cJSON17187105854 dayjs64336307610 gson7815880516 CsvHelper42390335805 semver3534685631,857 hashie153172300458 money6815726611,892 rich2555147361,421 fmt6043152641,097 sqlglot227115800777 scrapy3861,5801,2592,127 marked9058072,1674,091 http-parser4491,2011,2702,635 zod8411,4761,9563,608 quartznet4881,7071,8974,197 sqlparser6201,0071,2252,334 nlohmann/json3799291,7645,201 ts-pattern1,1301,2321,1391,910 nats-server9892,1492,5832,385 all 249,12516,40720,70443,598

Lines read counts only what came back from file-read calls; grep and shell output are search, not reading. The green figure in each row is the lowest of the four.

The same 24 bugs in the same order, with wall clock in place of lines read.

Wall clock is raw here — unlike the tables above, Benzi's per-repo index build is not subtracted, so these seconds run slightly higher than the warm figures quoted elsewhere on this page. Each point is that harness's most recent solved run for that bug; unsolved and unfinished runs are left out rather than plotted as fast. Benzi on Sonnet never solved http-parser and the DeepSeek harness never ran nats-server, so those two points are absent and neither enters its fit.

And the same again with dollars on the vertical axis.

Priced at the published per-token rates, same run selection as the chart above it. The two DeepSeek series run an order of magnitude cheaper than the two Sonnet ones, so at this scale they sit close to the baseline — the per-bug figures behind them are in the DeepSeek table further down. What the axis does show is the slope: Claude Code's cost climbs with difficulty faster than any other series here.

cost per fix · USD at list price · lowest of the four marked bug Benzi
Sonnet Benzi
DeepSeek Claude Code
Sonnet DeepSeek Harness
DeepSeek mux$0.39$0.036$0.25$0.023 commons-cli$0.26$0.030$0.29$0.014 addressable$0.57$0.043$0.35$0.042 jsoup$0.23$0.025$0.30$0.051 yaml-cpp$0.25$0.068$0.37$0.024 cJSON$0.27$0.036$0.44$0.073 dayjs$0.21$0.038$0.55$0.040 gson$0.19$0.015$0.45$0.022 CsvHelper$0.25$0.038$0.58$0.042 semver$0.62$0.10$1.43$0.097 hashie$0.69$0.051$0.92$0.043 money$0.79$0.036$0.95$0.052 rich$0.33$0.11$1.30$0.053 fmt$0.54$0.14$1.12$0.071 sqlglot$0.62$0.033$1.23$0.059 scrapy$0.73$0.10$1.35$0.19 marked$1.08$0.20$3.59$0.13 http-parser—$0.36$3.44$0.23 zod$1.53$0.13$2.47$0.33 quartznet$0.95$0.23$3.99$0.31 sqlparser$0.72$0.14$3.22$0.33 nlohmann/json$1.03$0.17$3.68$0.44 ts-pattern$3.74$0.31$4.33$0.053 nats-server$1.99$0.23$2.94— all 24$17.96$2.66$39.54$2.70

Priced at published per-token rates. The green figure in each row is the lowest of the four; the two DeepSeek columns are cheaper largely because that model costs roughly twenty times less per token. Blank cells are the two runs that never produced a fix.

이 글은 Hacker News 의 원문을 정제해 보여드립니다. 저작권은 원저작자에게 있습니다.

전체 내용이 궁금하다면

Hacker News 원문에서 이어 읽기

원문 보기

비슷한 글

5유사도 추천