Scale Low-Latency ML Inference Engineer (GPU/CUDA)
Scale Low-Latency ML Inference Engineer (GPU/CUDA)
ai breaking wireBrooklyn, NY
yesterday
Computer Systems Engineers/ArchitectsSoftware DevelopersComputer and Information Research Scientists
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesSoftware PublishersComputer Systems Design Services
Apply for this role →AI Breaking Wire seeks a Machine Learning Engineer to join our Inference Infrastructure team. You will build and optimize the high-throughput, low-latency distributed systems that power models like GPT-4 and Sora for millions of users worldwide.
The role emphasizes C++ and Python proficiency, GPU optimization with CUDA and Triton, and collaboration with research teams to productionize new architectures. This hybrid position is based in San Francisco.
Also on the board Same function, level within a rung
Level
Mid
Location
Brooklyn, NY
Occupation
Computer Systems Engineers/Architects
Industry
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
Posted
yesterday