Beranda Job Details
N
Information & Communication Technology 🏢 Full Time ⭐️ Terverifikasi

Site Reliability Engineer (SRE) - Remote from Bali, Indonesia

NCS Philippines
Canggu, Ubud, Denpasar, Jimbaran, Nusa Dua, Kuta, Badung
Salary Estimate
Rp 40.000.000 – Rp 70.000.000
Newest
Live Update
5 Agustus 2026
Deadline
5 Agu 2027

job description

Join NCS Philippines as a Site Reliability Engineer (SRE) and play a pivotal role in ensuring the reliability, scalability, and performance of our critical systems. This hybrid/remote role allows you to work from the vibrant hubs of Bali, Indonesia, including Canggu, Ubud, Denpasar, and more, while collaborating with a global team to deliver seamless digital experiences.

As an SRE, you will bridge the gap between development and operations, leveraging automation, monitoring, and incident response to maintain system stability. Your expertise in cloud infrastructure, CI/CD pipelines, and observability tools will drive efficiency and innovation across our platforms.

We offer a competitive salary, flexible work arrangements, and the opportunity to grow in a dynamic, tech-driven environment. If you are passionate about building resilient systems and thrive in a collaborative culture, we’d love to hear from you!

Responsibility

  • Design, implement, and maintain scalable and highly available systems using cloud platforms (AWS, GCP, or Azure).
  • Develop and optimize CI/CD pipelines to automate deployments, testing, and rollbacks.
  • Monitor system performance, identify bottlenecks, and proactively resolve issues to minimize downtime.
  • Collaborate with development teams to improve application reliability through best practices in coding, testing, and infrastructure.
  • Implement and manage observability tools (Prometheus, Grafana, ELK Stack) for real-time monitoring and alerting.
  • Automate operational tasks using scripting (Python, Bash) and infrastructure-as-code (Terraform, Ansible).
  • Participate in on-call rotations to respond to incidents, conduct post-mortems, and drive preventive measures.
  • Evaluate and adopt new technologies to enhance system resilience and efficiency.

Qualifications

  • Bachelor’s degree in Computer Science, Engineering, or a related field (or equivalent experience).
  • 3+ years of experience in Site Reliability Engineering, DevOps, or Cloud Operations.
  • Proficiency in Linux/Unix systems and containerization technologies (Docker, Kubernetes).
  • Hands-on experience with cloud platforms (AWS, GCP, Azure) and infrastructure-as-code tools.
  • Strong knowledge of monitoring, logging, and alerting systems (e.g., Prometheus, Nagios, Datadog).
  • Experience with scripting languages (Python, Go, Bash) and automation frameworks.
  • Familiarity with networking concepts (TCP/IP, DNS, load balancing) and security best practices.
  • Excellent problem-solving skills and ability to work in a fast-paced, collaborative environment.

Required Skills

Site Reliability Engineering DevOps Cloud Computing AWS GCP Azure Kubernetes Docker CI/CD Terraform Ansible Python Prometheus Grafana Linux Observability Incident Response Automation

Ready to Take This Challenge?

Make sure your resume is ready. Submit your application now before the deadline..

Apply Now

Lowongan Terkait

Rekomendasi pekerjaan serupa untuk Anda

Lihat Semua