Home Job Details
N
Information Technology 🏢 Full Time ⭐️ Verified

Lead AI Infrastructure Architect (2026 Vision)

Nexus Horizon Labs
San Francisco
Estimated Salary
USD 180.000 – USD 260.000
Live Update
4 Juli 2026
Deadline
4 Jul 2027

Job Description

We are pioneering the architecture for the 2026 AI revolution. As a Lead AI Infrastructure Architect at Nexus Horizon Labs, you will be responsible for designing and deploying the high-performance computing systems that power the next generation of artificial intelligence. You will bridge the gap between cutting-edge machine learning research and scalable, resilient production infrastructure. If you are ready to build the backbone of the future, we want to hear from you.

Why Join Us?

  • Work on state-of-the-art Large Language Models (LLMs).
  • Shape the infrastructure landscape for the 2026 technological paradigm shift.
  • Competitive compensation and equity packages.
  • Flexible remote-first culture with a hub in San Francisco.

Responsibilities

  • Architect and implement scalable distributed computing systems for training and inference of large-scale AI models.
  • Optimize hardware utilization, including GPU clusters, TPUs, and custom ASICs, to maximize throughput and minimize latency.
  • Collaborate with research scientists to translate theoretical models into efficient, production-ready pipelines.
  • Develop and enforce best practices for data pipeline management, model versioning, and CI/CD for ML workloads.
  • Ensure high availability and fault tolerance of critical AI infrastructure components.
  • Lead a team of DevOps and MLOps engineers to drive technical innovation and infrastructure modernization.

Qualifications

  • 7+ years of experience in systems architecture, software engineering, or DevOps, with a specific focus on machine learning infrastructure.
  • Deep expertise in containerization (Docker, Kubernetes) and orchestration platforms.
  • Proficiency in programming languages such as Python, Go, or C++.
  • Experience with cloud providers (AWS, GCP, or Azure) and serverless architectures.
  • Strong understanding of distributed systems principles, networking, and database technologies.
  • Experience with MLOps tools (MLflow, Kubeflow, TFX) and data processing frameworks (Spark, Flink).
  • Demonstrated ability to lead cross-functional teams and drive technical strategy in a fast-paced environment.

Required Skills

Kubernetes Docker Python Machine Learning Infrastructure MLOps Distributed Systems AWS GCP C++ GPU Optimization

Ready to Take This Challenge?

Make sure your resume is ready. Submit your application now before the deadline.

Apply Now

Related Jobs

Similar job recommendations for you

View All