job description
Are you an experienced SRE passionate about building resilient, high-scale infrastructure? Hays Recruitment is partnering with a market-leading technology organization in Kuala Lumpur to find a Senior Site Reliability Engineer to join their core Platforms Team. In this critical role, you will be the architect of reliability, ensuring our systems are not only robust and performant but also highly scalable to meet global demand.
We are looking for a proactive problem solver who thrives at the intersection of software engineering and operations. You will be instrumental in automating complex processes, driving incident response protocols, and championing a culture of continuous improvement through SRE best practices. If you are eager to influence architectural decisions and work with cutting-edge cloud-native technologies, we want to hear from you.
Responsibility
- Design, implement, and maintain highly available, scalable, and secure cloud infrastructure.
- Automate manual operational tasks using Infrastructure as Code (IaC) tools like Terraform or CloudFormation.
- Collaborate with cross-functional software development teams to optimize performance and deployment pipelines (CI/CD).
- Lead incident management and blameless post-mortems to improve system architecture and reduce MTTR.
- Establish and monitor Service Level Objectives (SLOs) and Error Budgets to balance feature velocity with system reliability.
- Optimize resource utilization and cloud costs without compromising performance.
- Mentor junior engineers and promote DevOps/SRE best practices across the organization.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related technical field.
- Minimum 5+ years of experience in SRE, DevOps, or Systems Engineering roles.
- Strong proficiency in Linux system administration and troubleshooting.
- Advanced experience with Cloud platforms (AWS, Azure, or GCP).
- Hands-on experience with container orchestration platforms such as Kubernetes (EKS/GKE).
- Solid programming skills in Go, Python, or Java for automation and tooling development.
- Deep understanding of observability tools like Prometheus, Grafana, Datadog, or ELK Stack.
- Strong communication skills with the ability to articulate technical debt and reliability strategies to stakeholders.