Job Description
Are you a technical expert seeking a role that offers flexibility and high-impact challenges? We are looking for a dedicated Site Reliability Engineer to join our elite critical infrastructure team. This is a unique opportunity to manage high-stakes systems during off-hours, offering a perfect balance of technical depth and work-life flexibility through our weekend and night shift rotations.
As part of our global support structure, you will ensure 99.99% uptime for our core services while enjoying the autonomy of a specialized technical role.
Responsibilities
- Monitor system health and performance during night shift and weekend windows.
- Respond to and resolve critical incidents with minimal supervision.
- Execute automated deployments and patch management during low-traffic windows.
- Conduct root cause analysis (RCA) for post-mortems and implement preventive measures.
- Collaborate with the 24/7 engineering team to optimize reliability and reduce MTTR.
- Ensure strict security compliance and data integrity during all operational activities.
Qualifications
- 3+ years of experience in Linux system administration or DevOps engineering.
- Proficiency in scripting languages (Python, Bash, or Go).
- Experience with cloud platforms (AWS, GCP, or Azure).
- Strong background in containerization technologies (Docker, Kubernetes).
- Proven track record of working in night shift or on-call rotation environments.
- Bachelor’s degree in Computer Science, Engineering, or related field.