Edge Inference Inc. is building a nationwide AI inference network, a distributed fabric of Inference Points that delivers latency, throughput, and token-cost commitments customers rely on. The Inference Solutions Architect converts a customer's model-serving workload into a concrete deployment: GPU platform, rack counts, serving-stack topology, KV-cache and batching strategy tailored to latency goals. You own the capacity math under every ICA — from MW reserved to GPU-hours committed — and

Also on the board Same function, level within a rung

Level

Mid

Location

Albany, NY

Occupation

Computer Systems Engineers/Architects

Industry

Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services

Posted

yesterday

Apply for this role →