계산 또는 계산: 도구가 행을 반환하면 해당 행을 올바르게 계산하는 모델이 토큰을 사용합니다.
단계 에이전트에 대한 Kaggle 벤치마크에서는 도구가 반환하는 내용을 세는 테스트를 거의 수행하지 않습니다. 10개의 모델, 68개의 질문, 개수를 반환하는 도구 하나, 행을 반환하는 도구 하나. 카운트로 모든 모델은
Count It or Compute It: When a Tool Returns Rows, the Models That Count Them Right Spend the Tokens
A Kaggle benchmark of the step agents rarely test: counting what a tool returns. Ten models, 68 questions, one tool that returns the count and one that returns the rows. With the count, every model is
dev.to 15점, 한글로 옮겨왔어요