Principal AI/ML Platform Engineer - Remote

Principal AI/ML Platform Engineer - Remote

optumSeattle, WA

yesterday

$164,600 - $282,200 annually

Computer Systems Engineers/ArchitectsComputer and Information Systems ManagersSoftware Developers
Computer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesSoftware Publishers
Apply for this role →
Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care’s most complex challenges. Your contributions here have the potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together. As a Principal AI/ML Platform Engineer on the United Health Group (UHG) enterprise team, you will serve as the AI Architect across IaaS and PaaS environments, owning the technical direction, reference architecture, and long-term evolution of our multi-tenant AI compute platform end to end. Our team builds and maintains an advanced compute estate spanning on-premises bare-metal Red Hat Open Shift AI clusters equipped with high-performance NVIDIA GPUs and Infini Band/RoCE training fabrics, alongside public-cloud managed AI services including Azure AI Foundry, AWS Bedrock, and GCP Vertex AI. In this role, you will define architecture standards, optimize high-throughput model training and inference pipelines, establish cost and utilization economics, and enforce strict HIPAA, security, and data-governance standards for regulated healthcare workloads. You’ll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges. For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities: Own the end-to-end reference architecture for multi-tenant AI compute platforms across hybrid on-premises bare-metal Open Shift AI clusters and public-cloud managed AI platforms (Azure AI Foundry, AWS Bedrock, GCP Vertex AI)Set network, latency, and topology standards for distributed training (NVLink, Infini Band, RoCEv 2, GPUDirect RDMA, NCCL), ensuring interconnect boundaries are strictly maintained Establish cluster governance, Git Ops workflows (Argo CD), RHACM policies, and automated lifecycle management for bare-metal accelerated compute nodes Design cost and utilization models including capex amortization, accelerator-sharing strategies (MIG/time-slicing), and cost-per-token/training-run math to inform accelerator procurement roadmaps Standardize model-serving platforms (vLLM, KServe) and inference gateways, setting quantization policies, provenance review gates, and intelligent model routing rules Define workload placement frameworks to determine self-hosted versus managed cloud deployment based on data residency, latency, cost, and compliance requirements Design and enforce identity, access, and security controls for AI workloads and autonomous agents, including least-privilege RBAC, Vault secret management, short-lived credentials, and mTLSEstablish enterprise Service Level Objectives (SLOs), disaster recovery plans, and upgrade strategies for Open Shift, Open Shift AI, GPU operators, drivers, and firmware Partner with cross-functional AI teams, LLM gateway engineers, privacy, and security stakeholders to ensure seamless integration and HIPAA compliance
Required Qualifications: Bachelor’s degree or 4+ years of equivalent software/platform engineering experience in lieu of a degree 10+ years of experience in infrastructure, Dev Ops, SRE, or ML platform engineering 5+ years of experience operating Kubernetes or Open Shift at scale in production bare-metal or enterprise cloud environments 3+ years of experience designing and managing accelerated-compute (AI/GPU) infrastructure utilizing NVIDIA GPU Operator, NFD, MIG/time-slicing, and DCGM3+ years of experience architecting hybrid AI platforms spanning self-hosted IaaS and public-cloud PaaS managed AI services (e.g., Azure AI Foundry, AWS Bedrock, or GCP Vertex AI)3+ years of experience with HPC/AI networking technologies, including Infini Band or RoCEv 2, GPUDirect RDMA, and NCCL collective communication limits 3+ years of experience managing multi-cluster fleets using RHACM (or equivalent) and Git Ops tooling (Argo CD or Flux)3+ years of experience implementing enterprise security and IAM controls for software workloads (RBAC, OIDC/OAuth, Vault secrets management, mTLS)
Preferred Qualifications: Experience with distributed training frameworks (PyTorch DDP/FSDP, Deep Speed, Ray, JAX) and batch scheduling systems (Kueue, Volcano)Hands-on experience with LLM inference serving technologies (vLLM, TensorRT-LLM, KServe) and platform tooling such as Open Shift AI (RHOAI) or Kubeflow pipelines Experience with high-performance parallel storage systems (Ceph/ODF, Lustre, IBM Storage Scale, VAST, WEKA)Active Red Hat certifications (e.g., Red Hat Certified Architect / RHCA) or open-source contributions to CNCF, Open Shift, or AI infrastructure projects Experience operating AI/ML platforms within regulated healthcare environments under HIPAA and UHG data privacy controls Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you’ll find a far-reaching choice of benefits and incentives. The salary for this role will range from $164,600 - $282,200 annually based on full-time employment. We comply with all minimum wage laws as applicable. At United Health Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone–of every race, gender, sexuality, age, location and income–deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes — an enterprise priority reflected in our mission. United Health Group is an Equal Employment Opportunity employer under applicable law and qualified applicants will receive consideration for employment without regard to race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations. United Health Group is a drug - free workplace. Candidates are required to pass a drug test before beginning employment.

Also on the board Same function, level within a rung

Level

Manager

Salary

$164,600 - $282,200 annually

Location

Seattle, WA

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

yesterday

Apply for this role →