comunidade metodistaPalo Alto, CA
LLM Inference Engineer: Scale & Optimize Production
LLM Inference Engineer: Scale & Optimize Production
LLM Inference Engineer: Scale & Optimize Production
comunidade metodistaPalo Alto, CA
yesterday
Computer Systems Engineers/ArchitectsSoftware DevelopersComputer and Information Research Scientists
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesSoftware PublishersElectronic Computer Manufacturing
Apply for this role →Hippocratic AI is seeking an experienced LLM Inference Engineer to optimize its LLM serving infrastructure in Palo Alto, CA. You will design multi-node architectures, implement multi-LoRA serving, and apply quantization techniques to reduce footprint while maintaining quality.
The role emphasizes practical production deployment, benchmarking, and latency-focused optimizations across GPUs, with a strong emphasis on Python, C++, and CUDA.
Also on the board Same function, level within a rung
Level
Lead
Location
Palo Alto, CA
Occupation
Computer Systems Engineers/Architects
Industry
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services
Posted
yesterday