You will focus on expanding our ML research platform to benchmark, rapidly prototype, and stress-test both software and hardware layers across our entire distributed ML stack. By leveraging AI agents and auto-research capabilities, you will push our systems to their limits, identify bottlenecks, and create a frictionless environment to test novel machine learning models on realistic, large-scale data.
Responsibilities: Platform Validation & Infrastructure Benchmarking:
  • Serve as the primary feedback loop for the entire ML stack.
  • Actively run complex models through our full ML pipeline to comprehensively test both the training and inference environments.
  • Validate the central infrastructure in practice, seeing exactly how new research ideas fare and identifying system bottlenecks before broader rollout to research teams.
  • Streamline Rapid Prototyping for ML Research:
  • Build high-level abstractions that allow users to bypass setup friction.
  • Integrate our core ML tooling directly with our underlying simulation and data frameworks, providing a unified entry point to access our full tech stack.
  • Enable rapid iteration on real-world data and seamless distributed training via Ray.
  • Agentic Workflows for ML Research:
  • Leverage AI agents and auto-research workflows to autonomously generate experiments, stress-test our distributed clusters, and provide data-driven, actionable feedback on what infrastructure needs to be optimized or built next.
  • Research Platform Feedback & Insights Sharing:
  • Act as the critical bridge between infrastructure builders and ML researchers.
  • Be the first to exhaustively test new models and push the platform's limits.
  • Document and publish empirical findings on system capabilities and hardware performance.
  • Take your validated insights to assist engineering teams with platform improvements and advise researchers on how to best leverage the stack.
Qualifications: Strong Software Engineering Foundation:
  • Deep proficiency in Python and software design principles.
  • Ability to build clean, scalable APIs and abstractions that other developers and researchers are enthusiastic about using.
  • Applied Machine Learning:
  • Hands-on experience with modern frameworks (PyTorch, TensorFlow, etc.)Strong practical understanding of how to train, evaluate, and deploy models at scale.
  • Distributed Compute:
  • Experience scaling ML workloads across GPUs and multi-node clusters using frameworks like Ray, Dask, or PyTorch Distributed.
  • AI Agent Workflows:
  • Familiarity with LLM tooling, agentic frameworks, and using AI to automate coding, research, or testing tasks.
  • System Profiling & Optimization:
  • Ability to debug and identify bottlenecks across hardware and software layers (e.g., memory limits, GPU utilization, data pipeline latency).Comp: $200-300K + Bonus

Also on the board Same function, level within a rung

Level

Lead

Salary

$200,000 - $300,000 per year

Location

New York, NY

Occupation

Computer and Information Research Scientists

Industry

Computer Systems Design Services

Posted

yesterday

Apply for this role →