Job Description
Are you ready to architect the digital landscape of 2026? Nexus Horizon is seeking a visionary Senior AI Infrastructure Engineer to lead our next-generation machine learning operations. In this pivotal role, you will bridge the gap between theoretical AI research and production-grade scalability, ensuring our systems are robust, efficient, and ready for the future.
We are building the infrastructure that will power autonomous systems, predictive analytics, and next-gen user experiences. If you are passionate about optimizing deep learning pipelines and working with state-of-the-art cloud technologies, we want to hear from you.
Why Join Us?
- Work with a world-class team pushing the boundaries of artificial intelligence.
- Competitive compensation package and equity options.
- Flexible remote-first culture with a hub in the heart of San Francisco.
Responsibilities
- Design, deploy, and maintain scalable machine learning pipelines and infrastructure on cloud platforms (AWS/GCP/Azure).
- Optimize deep learning models for low latency and high throughput, specifically targeting real-time inference requirements.
- Implement robust CI/CD pipelines for model training and deployment, ensuring automated testing and version control.
- Collaborate with data scientists and software engineers to integrate AI models into core product features seamlessly.
- Ensure high availability and disaster recovery protocols for all AI infrastructure assets.
- Monitor system performance, troubleshoot complex bottlenecks, and implement cost-reduction strategies for GPU clusters.
- Stay ahead of the curve on emerging AI hardware and software trends to drive architectural innovation.
Qualifications
- Bachelor’s degree in Computer Science, Engineering, or a related field; Master’s degree preferred.
- 5+ years of experience in software engineering, DevOps, or AI infrastructure.
- Strong proficiency in Python, C++, or Rust, with deep experience in ML frameworks like TensorFlow or PyTorch.
- Expert knowledge of containerization technologies (Docker, Kubernetes) and orchestration tools.
- Experience with cloud providers, specifically in managing GPU instances and serverless architectures.
- Proven track record of optimizing large-scale distributed systems and database performance.
- Excellent problem-solving skills and ability to communicate complex technical concepts to non-technical stakeholders.