"In April 2026, Scale updated the MCP-Atlas evaluation, upgrading the scoring judge and adding retry handling for transient tool errors. In addition, we also moved away from the previous 20 turn limit for the eval, and replaced it with a tool call budget of 100 max tool calls per task. We believe this config gives the models sufficient room to execute tool calls to complete the task. More importantly, it standardizes the eval harness for reasoning models and models that may prefer parallel v/s sequential tool calling per each turn."
"In April 2026, Scale updated the MCP-Atlas evaluation, upgrading the scoring judge and adding retry handling for transient tool errors. In addition, we also moved away from the previous 20 turn limit for the eval, and replaced it with a tool call budget of 100 max tool calls per task. We believe this config gives the models sufficient room to execute tool calls to complete the task. More importantly, it standardizes the eval harness for reasoning models and models that may prefer parallel v/s sequential tool calling per each turn."