job description
Join Gamer2Gamer as a Site Reliability Engineer (SRE) and become a key player in our technical team, ensuring the reliability, scalability, and performance of our gaming platforms. In this role, you will bridge the gap between development and operations, implementing automation, monitoring, and incident response strategies to deliver seamless user experiences. Based in the vibrant tech hubs of Bali, you'll collaborate with cross-functional teams to optimize infrastructure, reduce downtime, and enhance system resilience.
This is a unique opportunity to work in a dynamic, fast-paced environment where innovation meets gaming. If you're passionate about building robust systems and thrive in a culture of continuous improvement, we'd love to hear from you!
Responsibility
- Design, implement, and maintain scalable and reliable infrastructure to support high-traffic gaming applications.
- Develop and refine monitoring, alerting, and logging systems to proactively identify and resolve issues.
- Automate deployment, scaling, and management processes using CI/CD pipelines and infrastructure-as-code tools.
- Collaborate with development teams to improve system architecture, performance, and security.
- Lead incident response efforts, conduct post-mortems, and implement preventive measures to minimize future disruptions.
- Optimize cloud resources (AWS, GCP, or Azure) to balance cost, performance, and reliability.
- Mentor team members on SRE best practices and foster a culture of reliability and ownership.
- Stay ahead of industry trends and emerging technologies to drive innovation in system reliability.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related field, or equivalent practical experience.
- Proven experience as a Site Reliability Engineer, DevOps Engineer, or similar role in a high-availability environment.
- Strong proficiency in scripting languages (Python, Bash, Go) and automation tools (Ansible, Terraform, Puppet).
- Hands-on experience with cloud platforms (AWS, GCP, Azure) and containerization technologies (Docker, Kubernetes).
- Deep understanding of networking, security, and system architecture principles.
- Experience with monitoring tools (Prometheus, Grafana, Datadog) and logging systems (ELK Stack, Splunk).
- Excellent problem-solving skills and the ability to troubleshoot complex system issues under pressure.
- Strong communication skills and a collaborative mindset to work effectively with cross-functional teams.