Job Description
About Nexus Horizon
Nexus Horizon is pioneering the next generation of autonomous systems. We're building the infrastructure for a world where AI agents collaborate, reason, and execute complex workflows in real-time. We're looking for a visionary Senior Machine Learning Engineer to join our core AI 2026 initiative.
Role Overview
We are seeking a highly skilled Senior Machine Learning Engineer to lead the development of our proprietary large-scale reasoning models. In this role, you will define the architecture for multi-agent systems, optimize inference pipelines for latency and throughput, and collaborate with product and research teams to deploy robust, scalable AI solutions. You will work on cutting-edge problems at the intersection of reasoning, planning, and tool use.
What You Will Do
Design and implement scalable, distributed training pipelines for large language models with a focus on reasoning and tool-use capabilities.
Optimize model inference for ultra-low latency and high concurrency in production environments.
Build and maintain robust evaluation frameworks to measure reasoning accuracy, hallucination rates, and safety metrics.
Collaborate with product teams to integrate AI agents into complex workflows and edge cases.
Research and prototype novel techniques in chain-of-thought reasoning, planning, and multi-agent collaboration.
Ensure system safety, reliability, and alignment with ethical guidelines for autonomous agents.
Mentor junior engineers and foster a culture of technical excellence within the ML team.
Requirements
PhD or MS in Computer Science, Mathematics, or a related field with 8+ years of industry experience in ML.
Deep expertise in PyTorch, JAX, or TensorFlow, with experience in distributed training frameworks (Ray, Spark MLlib).
Proven track record of deploying large-scale models at production scale, optimizing for inference speed and memory footprint.
Strong background in NLP, specifically in reasoning models, tool use, and agent architectures.
Experience with MLOps, model versioning, and A/B testing frameworks.
Excellent communication skills and ability to translate technical concepts to cross-functional stakeholders.
Experience in safety, RLHF, or alignment techniques for AI agents.
Why Nexus Horizon?
We offer competitive compensation, equity, and a mission-driven culture. You'll have the autonomy to shape the future of AI, working with the brightest minds in the industry.
Responsibilities
- Design and implement scalable, distributed training pipelines for large language models with a focus on reasoning and tool-use capabilities.
- Optimize model inference for ultra-low latency and high concurrency in production environments.
- Build and maintain robust evaluation frameworks to measure reasoning accuracy, hallucination rates, and safety metrics.
- Collaborate with product teams to integrate AI agents into complex workflows and edge cases.
- Research and prototype novel techniques in chain-of-thought reasoning, planning, and multi-agent collaboration.
- Ensure system safety, reliability, and alignment with ethical guidelines for autonomous agents.
- Mentor junior engineers and foster a culture of technical excellence within the ML team.
Qualifications
- PhD or MS in Computer Science, Mathematics, or a related field with 8+ years of industry experience in ML.
- Deep expertise in PyTorch, JAX, or TensorFlow, with experience in distributed training frameworks (Ray, Spark MLlib).
- Proven track record of deploying large-scale models at production scale, optimizing for inference speed and memory footprint.
- Strong background in NLP, specifically in reasoning models, tool use, and agent architectures.
- Experience with MLOps, model versioning, and A/B testing frameworks. Excellent communication skills and ability to translate technical concepts to cross-functional stakeholders.
- Experience in safety, RLHF, or alignment techniques for AI agents.