研究团队开发的Jev驱动诊断管道在自动化故障排查中展现出显著优势。该系统在21个SREGym-Lite故障场景的105次诊断中成功率达76.2%,中位诊断时间仅为14.6秒1。与GPT-5.6 Sol (medium)的77.8%成功率相比,Jev诊断速度快约7倍,成本低约200倍1。
诊断管道的表现存在明显差异。在18个故障场景中,5次诊断尝试获得了相同的诊断分数,其中nginx-thrift内存限制故障实现了5次尝试全部通过的结果1。然而,包括edge_request_filter_cpu_saturation、kafka_poison_pill、search_rate_retry_collapse、service_dns_resolution_failure和valkey_auth_disruption在内的5个故障场景中,诊断完全失败,无一通过1。研究团队识别出两个主要失败模式:Jev选择了错误的线索,以及决定性证据在可用诊断数据中的缺失1。
A research team developed a diagnostic pipeline powered by Jev for automating troubleshooting without requiring LLM agent involvement.1 Across 105 diagnostic attempts spanning 21 SREGym-Lite failure scenarios, the pipeline achieved a 76.2% success rate, with a median diagnosis time of 14.6 seconds.1 When compared to GPT-5.6 Sol (medium), which achieved a 77.8% success rate, the Jev-driven approach operated approximately 7 times faster while costing roughly 200 times less.1
The results revealed considerable consistency in diagnostic outcomes. Among the 21 failure scenarios, 18 faults produced identical diagnosis scores across all five attempts, while the nginx-thrift memory constraint fault achieved a perfect diagnostic success rate of 5 out of 5 attempts.1 However, five failure types demonstrated complete diagnostic failure with zero successful resolutions: edge_request_filter_cpu_saturation, kafka_poison_pill, search_rate_retry_collapse, service_dns_resolution_failure, and valkey_auth_disruption.1 Analysis identified two primary failure modes: instances where Jev selected incorrect diagnostic leads, and cases where decisive evidence was absent from the available data.1
评论
还没有评论,欢迎留下第一条。