Beranda Job Details
B
Information & Communication Technology 🏢 Full Time ⭐️ Terverifikasi

Site Reliability Engineer (SRE) - Global System Service

ByteDance
Canggu, Ubud, Denpasar, Jimbaran, Nusa Dua, Kuta, Badung
Salary Estimate
Rp 30.000.000 – Rp 50.000.000
Newest
Live Update
27 Juli 2026
Deadline
27 Jul 2027

job description

About the Team

The Global System Service team at ByteDance is at the heart of our infrastructure, owning and optimizing the critical services and management solutions that power our data centers worldwide. As a Site Reliability Engineer (SRE), you will play a pivotal role in ensuring the reliability, scalability, and performance of our global systems. This is a unique opportunity to work with cutting-edge technology, collaborate with top-tier engineers, and make a tangible impact on one of the world's most innovative tech companies.

Why Join Us?

  • Work on large-scale, high-impact systems that serve millions of users globally.
  • Collaborate with a diverse, talented team in a dynamic and fast-paced environment.
  • Enjoy the flexibility of remote or hybrid work from beautiful Bali, with opportunities for global travel and collaboration.
  • Access to continuous learning and professional growth in a company that values innovation and excellence.

What You’ll Do

Responsibility

  • Design, build, and maintain scalable and reliable infrastructure services to support ByteDance's global operations.
  • Develop and implement automation tools and processes to improve system efficiency, reliability, and performance.
  • Monitor system health, identify potential issues, and proactively resolve them to minimize downtime.
  • Collaborate with cross-functional teams to ensure seamless integration and deployment of new features and services.
  • Conduct root cause analysis for system outages and implement preventive measures to avoid recurrence.
  • Optimize system performance through capacity planning, load balancing, and resource management.
  • Participate in on-call rotations to provide 24/7 support for critical systems.
  • Document system architectures, processes, and best practices to ensure knowledge sharing and team alignment.

Qualifications

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field.
  • 3+ years of experience in Site Reliability Engineering, DevOps, or a similar role.
  • Strong proficiency in scripting languages (e.g., Python, Bash) and automation tools (e.g., Ansible, Terraform).
  • Experience with cloud platforms (e.g., AWS, GCP) and containerization technologies (e.g., Docker, Kubernetes).
  • Solid understanding of networking, security, and system administration in Linux environments.
  • Familiarity with monitoring and observability tools (e.g., Prometheus, Grafana, ELK Stack).
  • Excellent problem-solving skills and the ability to troubleshoot complex system issues.
  • Strong communication and collaboration skills to work effectively in a global team.

Required Skills

Site Reliability Engineering DevOps Python Bash Automation Cloud Computing Kubernetes Docker Linux Monitoring Observability Networking Security Problem-Solving

Ready to Take This Challenge?

Make sure your resume is ready. Submit your application now before the deadline..

Apply Now

Lowongan Terkait

Rekomendasi pekerjaan serupa untuk Anda

Lihat Semua