Remote · worldwide, UTC+4 acceptedSalary not disclosedfull-timeVerified 1w agoHimalayas
You’ll be working in a team that drives the company's AI development to develop algorithms for a variety of different use cases. This includes e.g.
Responsibilities
Lead complex infrastructure and delivery initiatives across teams and environments.
Help evolve the platform toward self-service environments, clearer ownership of operations by product teams, and an SRE-style operating model.
Define and improve standards for cloud infrastructure, deployment, observability, and operational readiness.
Design and operate scalable, secure, and cost-efficient cloud platforms.
Work with Infrastructure-as-Code, Kubernetes, containers, CI/CD, and automation frameworks.
Improve monitoring, alerting, tracing, SLOs, disaster recovery, and incident-management practices.
Lead root-cause analysis and drive sustainable post-incident improvements.
Build security and compliance into infrastructure and delivery processes, including least-privilege access, identity federation and OIDC, short-lived credentials, secrets management and rotation, auditability, certificate management, supply-chain security, and policy-as-code.
Implement automated testing, linting, validation, and rollback processes for infrastructure and delivery pipelines.
Use AI agents responsibly to improve infrastructure design, automation, incident analysis, and operational workflows while applying strong engineering judgement and quality guardrails.