job description
Join TalentVibe Business Consultancy as a Site Reliability Engineer (SRE) and play a pivotal role in ensuring the stability, scalability, and performance of our cutting-edge digital infrastructure. Based in the vibrant and dynamic regions of Bali (Canggu, Ubud, Denpasar, and more), this is a unique opportunity to work remotely while contributing to high-impact projects in the Information & Communication Technology sector.
As an SRE, you will bridge the gap between development and operations, implementing automation, monitoring, and incident response strategies to maintain seamless system reliability. Your expertise will directly enhance user experience, minimize downtime, and drive continuous improvement in our cloud-native environments.
This contract role offers a competitive salary of $5,500 – $6,000 per month, along with the flexibility to work from Bali’s most sought-after locations. If you are passionate about building resilient systems and thrive in a collaborative, fast-paced environment, we want to hear from you!
Responsibility
- Design, implement, and maintain scalable and highly available systems using cloud-native technologies (e.g., Kubernetes, Docker, AWS/GCP).
- Develop and optimize automation scripts (Python, Bash, or Go) to streamline deployment, monitoring, and incident response.
- Monitor system performance, identify bottlenecks, and proactively resolve issues to ensure 99.9% uptime.
- Collaborate with development teams to integrate SRE best practices into CI/CD pipelines.
- Implement robust logging, metrics, and alerting systems (e.g., Prometheus, Grafana, ELK Stack) to enable data-driven decision-making.
- Conduct post-mortems for incidents, identify root causes, and drive preventive measures to enhance system reliability.
- Manage infrastructure as code (IaC) using tools like Terraform or Ansible to ensure consistency and reproducibility.
- Stay ahead of industry trends and advocate for the adoption of new technologies to improve efficiency and security.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related field (or equivalent work experience).
- 3+ years of experience in Site Reliability Engineering, DevOps, or Cloud Operations roles.
- Proficiency in Linux/Unix systems and scripting languages (Python, Bash, or Go).
- Hands-on experience with containerization (Docker, Kubernetes) and cloud platforms (AWS, GCP, or Azure).
- Strong knowledge of monitoring tools (Prometheus, Grafana, Nagios) and logging systems (ELK, Splunk).
- Experience with Infrastructure as Code (IaC) tools like Terraform, Ansible, or Puppet.
- Familiarity with CI/CD pipelines (Jenkins, GitLab CI, GitHub Actions) and version control systems (Git).
- Excellent problem-solving skills and the ability to work under pressure in a fast-paced environment.