I'm excited for the rollout of the verifiers stack...
# success-stories
e
I'm excited for the rollout of the verifiers stack! I recently discussed the "death of SWE-Bench Verified" (due to contamination) in a segment of our AI-First Show:

https://youtu.be/ycTg7GnNuCg?t=4461โ–พ

One of my students commented with a couple of questions, one of them about agent benchmarking, and I just replied referencing the OpenHands trajectory-level verifier. Aside from ensuring coding agent reliability and increasing trust, I think this could also be an important verification component in agent benchmarking.
๐Ÿ™Œ 3
l
Awesome, thanks for the shout out!
Small correction, I think that Scale AI created SWE-Bench Pro: https://labs.scale.com/leaderboard/swe_bench_pro_public
e
Oh! I was convinced it was OpenAI - thanks for the correction ๐Ÿ˜… I guess I got confused due to OpenAI's support and evangelism there.
l
Great analysis though!
Weโ€™re thinking of replacing SWE-Bench Verified with SWE-Bench Pro for the next version of the OpenHands Index
๐Ÿ‘๐Ÿผ 1
e
Have you seen this? They just released SWE-Atlas https://scale.com/blog/swe-atlas
l
Yeah! I think we discussed it in the #C09MTM3QPLG project
โœ”๏ธ 1
Because of the Codebase QnA component, which we were trying to develop there