Hello BrowseComp-Plus maintainers,
EvidenceMesh is assessing whether BrowseComp-Plus could satisfy our internal admission policy for an officially comparable fixed-corpus benchmark. In the bounded evidence set reviewed so far, we do not yet have sufficient evidence to close two gates: component-level rights disposition and immutable effective evaluation configuration.
This is a request for documented facts and artifact identities, not legal advice or a legal conclusion. We do not assume that the project maintainers can speak for every third-party rights holder. Where your team is not the competent authority, please identify the relevant authority if known or mark the status as unknown or not confirmed.
Please do not provide benchmark payloads, corpus content, examples, queries, answers, qrels, model weights, secrets, personal data or a newly executed score. Identifiers, digests, immutable references and a machine-readable configuration are sufficient.
1. Component-level rights disposition
For every component required to run the official fixed-corpus evaluation, could you provide or identify an authoritative record covering:
- the component name, type, provenance and exact version or snapshot;
- the relevant rights holder or authority competent to document its status;
- the applicable license, permission, terms or other documented basis, with an immutable artifact identity where available;
- whether that basis separately covers acquisition, private local storage and processing solely for evaluation, publication of aggregate metrics or scores without content or excerpts, and those uses under a strict non-redistribution condition; and
- every restriction, unresolved status or component whose responsible authority or documented permission remains unknown.
Please distinguish rights documented by the benchmark project from rights originating with third-party Web-content owners. A project-level license or availability statement should not be treated as a component-level disposition unless its scope explicitly covers that component and the intended uses above.
2. Immutable effective evaluation closure
For each published result, result row or result family intended to represent the official fixed-corpus comparison, could you provide a stable result identifier and a one-to-one mapping to:
- the evaluation-runner repository, exact commit, path and content identity;
- the judge model repository, exact revision and relevant artifact digests;
- the tokenizer repository, exact revision, tokenizer configuration and chat-template identity;
- the exact effective grader prompt, including system and user wrappers or templates, with a source identity or content digest;
- every effective generation setting and override, including decoding parameters, maximum output length, stop conditions, thinking mode, seed or seed policy, repetition policy, batch size and parallelism;
- the fully resolved runtime, including Python version, exact dependency versions or commits, effective runtime flags, and a lockfile or container-image digest where available;
- the parser source identity and effective configuration, including normalization and malformed-output handling;
- the scoring source identity and effective configuration, including metric definitions, aggregation, missing-output handling and error handling; and
- the exact invocation or machine-readable effective-run manifest binding the result identifier to all these elements.
Please distinguish source-code defaults from the effective values actually used for each reported comparison. If different official results used different configurations, a separate mapping for each configuration would avoid treating them as one evaluator identity.
This request records an EvidenceMesh evidence gap only. It does not assert that any permission exists or does not exist, and a future response would not automatically admit the benchmark or authorize acquisition, evaluation or publication.
Thank you.
Hello BrowseComp-Plus maintainers,
EvidenceMesh is assessing whether BrowseComp-Plus could satisfy our internal admission policy for an officially comparable fixed-corpus benchmark. In the bounded evidence set reviewed so far, we do not yet have sufficient evidence to close two gates: component-level rights disposition and immutable effective evaluation configuration.
This is a request for documented facts and artifact identities, not legal advice or a legal conclusion. We do not assume that the project maintainers can speak for every third-party rights holder. Where your team is not the competent authority, please identify the relevant authority if known or mark the status as unknown or not confirmed.
Please do not provide benchmark payloads, corpus content, examples, queries, answers, qrels, model weights, secrets, personal data or a newly executed score. Identifiers, digests, immutable references and a machine-readable configuration are sufficient.
1. Component-level rights disposition
For every component required to run the official fixed-corpus evaluation, could you provide or identify an authoritative record covering:
Please distinguish rights documented by the benchmark project from rights originating with third-party Web-content owners. A project-level license or availability statement should not be treated as a component-level disposition unless its scope explicitly covers that component and the intended uses above.
2. Immutable effective evaluation closure
For each published result, result row or result family intended to represent the official fixed-corpus comparison, could you provide a stable result identifier and a one-to-one mapping to:
Please distinguish source-code defaults from the effective values actually used for each reported comparison. If different official results used different configurations, a separate mapping for each configuration would avoid treating them as one evaluator identity.
This request records an EvidenceMesh evidence gap only. It does not assert that any permission exists or does not exist, and a future response would not automatically admit the benchmark or authorize acquisition, evaluation or publication.
Thank you.