About the role
SpaceX was founded under the belief that a future where humanity is out exploring the stars is fundamentally more exciting than one where we are not. Today SpaceX is actively developing the technologies to make this possible, with the ultimate goal of enabling human life on Mars.
SR. IT SYSTEMS ENGINEER, DEVOPS
SpaceX is looking for a Senior IT Systems Engineer with a focus on DevOps to join our growing team. You will own the architecture, design, implementation, and evolution of our CI/CD platforms, infrastructure automation, and reliability engineering practices at scale, including on-premise VM and Kubernetes infrastructure. This employee will be a senior member of the Information Technology IT Applications team. The ideal candidate will be a technical leader who is flexible and flourishes in a fast-paced, high-stakes environment. He or she should be a self-starter, self-motivator, and possess the ingenuity and ownership to excel at this position while mentoring others.
RESPONSIBILITIES
Architect, design, implement, and manage enterprise-scale CI/CD pipelines and platforms running on both on-premise VM infrastructure and Kubernetes clusters.
Lead automation of deployment, configuration, monitoring, and operations to drive reliability, scalability, and efficiency across hundreds of on-premise servers, VMs, and containerized environments.
Define and enforce DevOps, SRE, and release engineering best practices, standards, and security controls across on-premise VM and Kubernetes infrastructure.
Partner closely with software engineering, development, and operations teams to streamline code integration, testing, delivery, and feedback loops.
Own system performance, capacity planning, troubleshooting, and optimization of on-premise VM clusters and Kubernetes environments; drive root-cause analysis and long-term improvements.
Design, maintain, and enhance observability (logging, metrics, tracing, alerting) for on-premise infrastructure using tools such as Grafana, Prometheus, Influx, Splunk, or equivalents.
Lead high-availability, disaster recovery, and business continuity strategies for on-premise VM and Kubernetes platforms, including planning, executing, a…