6 comments

  • smithclay 2 minutes ago
    More benchmarks comparing effectiveness of various AI SREs is welcome and overdue: especially vendor-vs-vendor comparisons. One area that I think is going to be really interesting and important is the best way to emulate complex IT environments for evals.

    Some related work I recommend checking out: - https://arxiv.org/abs/2609.33023 (new last week!) - https://github.com/SREGym/SREGym - https://github.com/hyperdxio/hyperdx/tree/main/packages/hdx-... (Clickstack's version) - https://github.com/grafana/o11y-bench (Grafana's version)

  • tkkiran 9 minutes ago
    Cool, as I see your product react proactively rather than Claude being reactive so that with your tools automated investigations happens without me asking explicitly the problem and root cause as I understand? Do you guys also have mitigations?
    • emrahsamdan 4 minutes ago
      Yes depending on which connector you chose to connect. There could be a new PR on GitHub, a rollback in CircleCI or a feature flag switch on LaunchDarkly. All waits for a human to chime in by default.

      These are all becoming tablestakes, imo for any AI SRE. The real moat is to build the data layer in an efficient manner so that the investigation (and mitigation) is fast and cost-efficient. The models are making it easier for any company at the same time.

  • nikhilunni 1 hour ago
    Cool that you guys created this, but interpreting the results: why wouldn't I just use Claude instead of an AI SRE tool?

    Or if there are some features of an AI SRE tool that make it better than Claude + some MCPs, should those be captured in this same benchmark?

  • rafiyashaheen11 1 hour ago
    How much time did it take to build this? Loved it! What problem does this solve?
    • emrahsamdan 50 minutes ago
      Running the test takes around 5-6 hours per vendor depending on how much wrestle it requires (not every system is ideal for ai agents - including ourselves) but the actual work was to come up with the real world scenarios and what would be the expected RCA and remediation for this. Our own SRE team worked for 1.5 weeks for that.
  • esafak 41 minutes ago
    Emrah, it seems that EDX does not offer any edge over GCX, when paired with an agent like Claude?
  • tanmoy1139 26 minutes ago
    [flagged]