Hippocratic AI is seeking an experienced LLM Inference Engineer to optimize its LLM serving infrastructure in Palo Alto, CA. You will design multi-node architectures, implement multi-LoRA serving, and apply quantization techniques to reduce footprint while maintaining quality. The role emphasizes practical production deployment, benchmarking, and latency-focused optimizations across GPUs, with a strong emphasis on Python, C++, and CUDA.

Also on the board Same function, level within a rung

Level

Lead

Location

Palo Alto, CA

Occupation

Computer Systems Engineers/Architects

Industry

Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services

Posted

yesterday

Apply for this role →