Back to News
Advertisement
Advertisement

⚡ Community Insights

Discussion Sentiment

0% Positive

Analyzed from 102 words in the discussion.

Trending Topics

#development#agent#quality#findings#harnesses#releases#level#abstract#reveal#despite

Discussion (1 Comments)Read Original on HackerNews

wek•about 3 hours ago
From their abstract: "Our findings reveal that despite continuous development activity and growing codebase complexity of the agent harnesses, there is no statistically significant improvement in SWE-bench benchmark score (i.e., resolve rates of bugs) across releases for a given fixed LLM version. Worse, later agent harness versions consume nearly double the computational tokens and tool calls without corresponding quality gains. We explain this paradox from two angles: at the project level, we identify the development patterns (e.g., feature additions, fix-heavy releases, scattered small changes) correlating with quality fluctuations, while at the architecture level, we localize regressions to specific high-risk architectural components. Our findings call for a new practice of quality assurance in the development of agent harnesses."