All vacancies
Site Reliability Engineer, Infrastructure Platforms - UK (Intermediate to Senior Staff)
GitLab
Remote · United KingdomSalary not disclosedfull-timeGitLab Careers
An overview of this role Site Reliability Engineers keep GitLab's user-facing services and production systems running reliably at scale. They combine software engineering with operational excellence, applying sound engineering principles, automation, and continuous improvement to build, operate, and evolve our production infrastructure.
Responsibilities
- Keep user-facing services and production systems reliable, scalable, and efficient
- Build automation and tooling that reduces toil and replaces manual work with repeatable, infrastructure-as-code-driven workflows
- Operate and troubleshoot production systems on Kubernetes, including deployments, rollouts, and scaling
- Write and maintain infrastructure as code, and ship changes safely through CI/CD and GitOps
- Participate in on-call, triage alerts, follow and improve runbooks, and escalate appropriately
- Contribute to the observability stack, using metrics, logs, and SLOs to detect symptoms early rather than just outages
- Take part in incident response and post-incident reviews, turning learnings into changes in automation and process
- Document runbooks, architecture decisions, and reviews so your findings become repeatable practices
Languages
- Work format
- Remote
- Seniority
- Senior
- Posted
- 8 Sept 2026
- Last verified
- In the last 5 days
