DRW is seeking an AI Inference Platform Engineer to build, operate, and optimize systems that serve large language, vision, multimodal, and embedding models across DRW. You’ll own the serving platform end-to-end, onboarding models, measuring quality and performance across configurations, and scaling workloads for multiple tenants. This role requires hands-on GPU literacy and deep experience with modern inference runtimes, KV cache architectures, and multi-tenant scheduling, all within a trading

Also on the board Same function, level within a rung

Level

Manager

Location

Chicago, IL

Occupation

Computer Systems Engineers/Architects

Industry

Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services

Posted

yesterday

Apply for this role →