Task Resolve Rate: 50.8% (63/124) — big-pickle, the free stealth model on OpenCode Zen, evaluated on Scale AI's SWE Atlas Codebase QnA benchmark using the mini-swe-agent scaffold.
Run on 2026-08-11 with the official open-source harness, task data, and judge model.
Against the official SWE Atlas QnA leaderboard (updated 2026-07-28):
Within the Mini-SWE-Agent scaffold class — the apples-to-apples comparison — this run outscores every entry on the official leaderboard, and it also tops the Codex-scaffold GPT entries. Only the two Claude models running on their native Claude Code scaffold score higher. Note the caveats below before treating this as a leaderboard-equivalent number.