Remote · ArmeniaSalary not disclosedfull-timeVerified recentlyhh.ru
We are looking for a Senior DevOps observability (SRE) engineer who will own the development of Release Management and CI/CD, and the operation of the production infrastructure behind our key products.
Responsibilities
Maintaining and developing the observability strategy, standardising approaches and advising the development teams
Keeping applications running reliably in Kubernetes in production, including incident diagnosis and analysis of distributed systems
Developing and maintaining the observability platform (metrics/logs/traces), including the VM stack, CloudWatch, ELK and related tooling
Providing monitoring, alerting and dashboards that reflect the services' architecture and production requirements
Supporting the development teams on service operation, incidents and interaction between systems
Owning the full production incident lifecycle: on-call response (PagerDuty), coordinating mitigation and service recovery, RCA/post-mortems, and implementing the reliability improvements that come out of them
Advancing incident management and reliability engineering practices that reduce repeat incidents and improve the stability of production systems