I guess the agentic coding benchmarks don't have many rewards for stopping and clarifying what the user wants?