xAI· Infrastructure· London, UK
Site Reliability Engineer (SRE)
Classified Tasks (10)
Automate 0%Augment 90%Human-Only 10%
Augment (9)
AI assists, human decides
Develop and maintain backend services that power products such as grok.com and the API
technical
Design and implement services to efficiently process tens of thousands of queries per second
technical
Deploy, operate, and maintain Kubernetes clusters across on-premises and cloud environments
operational
Build and maintain continuous deployment pipelines using Buildkite and ArgoCD
operational
Implement and operate monitoring, alerting, and incident response using Prometheus, Grafana, and PagerDuty
operational
Define, provision, and manage infrastructure as code using Pulumi or Terraform
operational
Write and maintain systems-level code in languages such as Rust, C++, or Go for backend services
technical
Configure and manage traffic management and HTTP proxies such as nginx and Envoy
technical
Ensure production services meet scalability, reliability, and performance targets
operational
Human-Only (1)
Requires human judgment
Collaborate with team members to operate production backend services and resolve operational issues
operational
Job description
ABOUT xAI xAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. All employees are expected to have strong communication skills. They should be able to concisely and accurately share knowledge with their teammates. ABOUT THE ROLE: You will work on the team that is responsible for the backend services that power our products such as grok.com and the API. We focus on writing and maintaining highly scalable and reliable services that can efficiently process tens of thousands of queries per second. The services are hosted on a number of Kubernetes clusters (on-prem & cloud). BASIC QUALIFICATIONS: Expert knowledge of Kubernetes. Expert knowledge of continuous deployment systems such as Buildkite and ArgoCD. Expert knowledge of monitoring technologies such as Prometheus, Grafana, and PagerDuty. Expert knowledge of infrastructure as code technologies such as Pulumi or Terraform. Familiarity with a systems programming language like Rust, C++ or Go Experience with traffic management and HTTP proxies such as nginx and envoy. xAI is an equal opportunity employer. For details on data processing, view our Recruitment Privacy Notice .