job description
Are you a seasoned DevOps/Site Reliability Engineer passionate about building scalable, resilient AI infrastructure? Join our dynamic team in Bali as a Principal Software Engineer and lead the charge in optimizing AI efficiency through cutting-edge operational frameworks.
In this role, you will architect and maintain high-availability systems, automate deployment pipelines, and ensure seamless performance across distributed environments. Your expertise will directly impact the reliability and scalability of our AI-driven solutions, empowering innovation at scale.
Based in Bali’s vibrant tech hubs (Canggu, Ubud, Denpasar, Jimbaran, Nusa Dua, Kuta, or Badung), you’ll collaborate with cross-functional teams to drive operational excellence while enjoying a flexible, results-driven work culture.
Responsibility
- Design, implement, and maintain scalable infrastructure for AI workloads.
- Lead incident response, root cause analysis, and system reliability improvements.
- Automate CI/CD pipelines and deployment strategies for zero-downtime releases.
- Optimize cloud resources (AWS/GCP/Azure) for cost-efficiency and performance.
- Mentor junior engineers and foster a culture of DevOps best practices.
- Monitor system health, capacity planning, and proactive scalability measures.
- Collaborate with AI/ML teams to integrate reliability into model deployment.
Qualifications
- 8+ years in DevOps/SRE roles, with expertise in cloud platforms (AWS/GCP/Azure).
- Proficiency in IaC tools (Terraform, Ansible, Pulumi) and containerization (Docker, Kubernetes).
- Strong scripting skills (Python, Bash, Go) and experience with monitoring tools (Prometheus, Grafana).
- Proven track record in improving system reliability and reducing MTTR.
- Experience with AI/ML infrastructure is a plus.
- Excellent problem-solving and communication skills.