job description
Join Search Index Pte Ltd, a leading technology company, as our Senior Infrastructure Engineer (Monitoring & Operations) and play a pivotal role in shaping the reliability, performance, and cost-efficiency of our enterprise IT operations. Based in Bali, Indonesia (with remote flexibility in Canggu, Ubud, Denpasar, Jimbaran, Nusa Dua, Kuta, or Badung), youâll lead critical initiatives in enterprise monitoring, service assurance, and cloud cost optimization to ensure seamless, high-performing IT infrastructure.
In this role, youâll collaborate with cross-functional teams to design, implement, and optimize monitoring solutions that proactively detect and resolve issues before they impact business operations. Youâll also drive cloud cost optimization strategies, ensuring our infrastructure remains scalable, secure, and cost-effective. If youâre passionate about observability, automation, and cloud-native technologies, this is your opportunity to make a tangible impact in a dynamic, fast-growing environment.
Search Index Pte Ltd values innovation, collaboration, and work-life balance. As part of our team, youâll enjoy competitive compensation, professional growth opportunities, and the flexibility to work from one of Baliâs most vibrant tech hubs. Apply now and help us build the future of reliable, high-performance IT infrastructure!
Responsibility
- Design, implement, and maintain enterprise-grade monitoring and observability solutions (e.g., Prometheus, Grafana, ELK Stack, Datadog) to ensure system reliability and performance.
- Lead service assurance initiatives, including incident response, root cause analysis (RCA), and proactive issue resolution to minimize downtime.
- Optimize cloud infrastructure costs (AWS, GCP, or Azure) by identifying inefficiencies, rightsizing resources, and implementing FinOps best practices.
- Develop and automate monitoring alerts, dashboards, and reporting to provide real-time insights into system health and performance.
- Collaborate with DevOps, SRE, and engineering teams to improve infrastructure resilience, scalability, and security.
- Implement and manage logging, tracing, and APM (Application Performance Monitoring) tools to enhance observability across distributed systems.
- Document monitoring policies, procedures, and runbooks to ensure knowledge sharing and operational continuity.
- Stay updated on emerging trends in infrastructure monitoring, cloud optimization, and observability tools to drive continuous improvement.
Qualifications
- Bachelorâs degree in Computer Science, Information Technology, or a related field (or equivalent experience).
- 5+ years of experience in infrastructure engineering, monitoring, or DevOps/SRE roles, with a focus on enterprise environments.
- Hands-on experience with monitoring tools such as Prometheus, Grafana, Nagios, Zabbix, Datadog, or New Relic.
- Proficiency in cloud platforms (AWS, GCP, or Azure) and cloud cost optimization strategies (FinOps).
- Strong scripting skills (Python, Bash, or Go) for automation and tooling development.
- Experience with logging and tracing tools (ELK Stack, Jaeger, OpenTelemetry).
- Familiarity with infrastructure-as-code (IaC) tools like Terraform, Ansible, or Pulumi.
- Excellent problem-solving skills and ability to work in a fast-paced, collaborative environment.
- Certifications such as AWS Certified DevOps Engineer, Google Professional Cloud DevOps Engineer, or Certified Kubernetes Administrator (CKA) are a plus.