GitHub于8月17日遭遇长达7小时47分钟的服务中断1。此次事故影响范围广泛,github.com、用户认证、GitHub Actions、API、Pull Request、Issue和Copilot等核心服务均受波及1。
根据官方分析,服务中断的根本原因是位于Central US数据中心的关键基础设施组件未能随流量高峰而扩展,导致容量压力触发认证失败,进而引发整体服务中断1。值得注意的是,这是GitHub在8月内发生的第二次重大事故,两次事故均由容量失败引起,而非代码或配置变更所致1。
为应对业务增长压力,GitHub正在积极推进向Azure的基础设施迁移。截至事故发生时,Azure已承载GitHub平台负载的58%和一半的Git操作,相比5月的12%有显著提升1。与此同时,GitHub新增超过300万个CPU核心和120PB高速存储以扩展容量1。这些投入反映了平台流量的快速增长——从4月的14亿次commit增长到8月的29亿次1。
GitHub experienced a significant service disruption on August 17 that lasted 7 hours and 47 minutes, affecting core platforms including github.com, authentication, GitHub Actions, APIs, Pull Requests, Issues, and Copilot 1. The outage was triggered by capacity constraints in the company's Central US data center, where critical infrastructure components failed to scale adequately in response to traffic surges 1.
The root cause stemmed from insufficient expansion of key infrastructure resources during peak demand periods, which ultimately led to authentication failures cascading across GitHub's ecosystem 1. This incident marked the second major outage GitHub experienced in August alone, with both disruptions attributed to capacity failures rather than code or configuration errors 1. The platform has experienced substantial growth, with monthly commits increasing from 1.4 billion in April to 2.9 billion by August 1.
To address these vulnerabilities, GitHub has invested heavily in infrastructure expansion, adding over 3 million CPU cores and 120 petabytes of high-speed storage 1. The company's migration to Azure has progressed significantly, with the cloud platform now hosting 58 percent of GitHub's platform load and handling half of all Git operations, up from just 12 percent as of May 1.
评论
还没有评论,欢迎留下第一条。