Hello,
I maintain Agent RCA Bench, an open-source benchmark measuring whether the observability interface behind an LLM agent changes root cause analysis quality. Six of its schema-discovery cases are built on OpenRCA 1.0 Bank, Market, and Telecom telemetry. Before publishing, I would like to confirm our use is consistent with your terms.
What we do with the data.
Each case is downloaded from your Google Drive release at a pinned revision, ingested into a local database, queried by an LLM agent, and then deleted. We redistribute no telemetry, no labels, and no archives — the repository is downloader-backed. What we publish is derived and sanitized: source case identifiers, component and table names, time windows, sample counts, and content hashes, plus measurements of model behaviour (token counts, tool calls, whether the diagnosis was correct).
What the study asks.
Whether an all-in-one observability database, and a semantic layer on top of it, help an agent do RCA — stated as an open question, not a claim to confirm. The published result is largely negative for the layer we build: none of the eight pre-specified endpoints for the semantic layer survived multiple-comparison correction. The benchmark is Apache-2.0, the report is free with no paywall or registration, and the full protocol, scorers, and artifacts are public so the measurement can be checked or refuted.
Two things I want to be transparent about. The benchmark is sponsored and maintained by Greptime, and this is stated in the repository and in the report itself. The database under test, GreptimeDB, is Apache-2.0; every capability we measured lives in its open-source core and is present in a default build, so anyone can reproduce the run from the same public commit without a commercial agreement.
My questions:
-
Could you confirm the license that applies to the released telemetry? The repository declares MIT for the code, and I have seen CC BY-NC 4.0 cited for the data, but I could not locate an authoritative statement to point at. I want to record the correct one.
-
If CC BY-NC 4.0 applies: given the above, would you consider this use to fall within NonCommercial? If you would rather treat it otherwise, I would welcome a written permission, or I will adjust what we publish — including removing the OpenRCA-derived cases — to stay within your terms.
Either way we will attribute OpenRCA and the ICLR 2025 paper wherever derived facts appear. Happy to share the repository and the draft report if useful.
Thanks for building and releasing OpenRCA,
Dennis Zhuang
Greptime
Hello,
I maintain Agent RCA Bench, an open-source benchmark measuring whether the observability interface behind an LLM agent changes root cause analysis quality. Six of its schema-discovery cases are built on OpenRCA 1.0 Bank, Market, and Telecom telemetry. Before publishing, I would like to confirm our use is consistent with your terms.
What we do with the data.
Each case is downloaded from your Google Drive release at a pinned revision, ingested into a local database, queried by an LLM agent, and then deleted. We redistribute no telemetry, no labels, and no archives — the repository is downloader-backed. What we publish is derived and sanitized: source case identifiers, component and table names, time windows, sample counts, and content hashes, plus measurements of model behaviour (token counts, tool calls, whether the diagnosis was correct).
What the study asks.
Whether an all-in-one observability database, and a semantic layer on top of it, help an agent do RCA — stated as an open question, not a claim to confirm. The published result is largely negative for the layer we build: none of the eight pre-specified endpoints for the semantic layer survived multiple-comparison correction. The benchmark is Apache-2.0, the report is free with no paywall or registration, and the full protocol, scorers, and artifacts are public so the measurement can be checked or refuted.
Two things I want to be transparent about. The benchmark is sponsored and maintained by Greptime, and this is stated in the repository and in the report itself. The database under test, GreptimeDB, is Apache-2.0; every capability we measured lives in its open-source core and is present in a default build, so anyone can reproduce the run from the same public commit without a commercial agreement.
My questions:
Could you confirm the license that applies to the released telemetry? The repository declares MIT for the code, and I have seen CC BY-NC 4.0 cited for the data, but I could not locate an authoritative statement to point at. I want to record the correct one.
If CC BY-NC 4.0 applies: given the above, would you consider this use to fall within NonCommercial? If you would rather treat it otherwise, I would welcome a written permission, or I will adjust what we publish — including removing the OpenRCA-derived cases — to stay within your terms.
Either way we will attribute OpenRCA and the ICLR 2025 paper wherever derived facts appear. Happy to share the repository and the draft report if useful.
Thanks for building and releasing OpenRCA,
Dennis Zhuang
Greptime