Skip to content

Request for the new evaluation script #27

Description

@Sunliangtai

"In April 2026, Scale updated the MCP-Atlas evaluation, upgrading the scoring judge and adding retry handling for transient tool errors. In addition, we also moved away from the previous 20 turn limit for the eval, and replaced it with a tool call budget of 100 max tool calls per task. We believe this config gives the models sufficient room to execute tool calls to complete the task. More importantly, it standardizes the eval harness for reasoning models and models that may prefer parallel v/s sequential tool calling per each turn."

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions