Job Description
Are you a technical expert looking for a role that balances autonomy with high-impact work? Apex Systems is seeking a Senior Site Reliability Engineer to join our elite infrastructure team, specifically focused on managing critical weekend operations.
In this role, you will ensure the stability and scalability of our cloud-native platforms, working in a full-time capacity but with the flexibility of a dedicated weekend rotation. This is not just a job; it is an opportunity to shape the future of our engineering culture.
Why Join Us?
We offer a competitive benefits package, a collaborative environment, and the chance to work with cutting-edge technologies in a stable, full-time capacity.
Responsibilities
- Lead weekend maintenance windows to deploy updates, patches, and infrastructure changes without disrupting business operations.
- Respond to and resolve critical incidents during off-peak hours with minimal latency.
- Design and implement automated scaling solutions to handle traffic spikes.
- Conduct post-mortem analyses to drive continuous improvement in system reliability.
- Collaborate with development teams to define best practices for CI/CD pipelines.
- Optimize database performance and cloud resource utilization.
Qualifications
- 5+ years of experience in DevOps, SRE, or Systems Engineering roles.
- Strong proficiency in scripting languages such as Python, Bash, or Go.
- Expert knowledge of AWS, Azure, or Google Cloud Platform.
- Experience with containerization technologies (Docker, Kubernetes).
- Proven track record of managing on-call rotations and incident response.
- Bachelor’s degree in Computer Science or related field.