Infra Backend Engineering, Security
Infra Backend Engineering, Security
gmi cloudChicago, IL
yesterday
Computer Systems Engineers/ArchitectsSoftware DevelopersNetwork and Computer Systems Administrators
Computer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesSoftware Publishers
Apply for this role →About GMIGMI Cloud is a fast-growing, AI-native infrastructure company delivering high-performance GPU compute, inference services, and infrastructure for AI agents. Following 8x ARR growth, GMI Cloud continues to scale rapidly across the U.S. and APAC. As a Reference Platform NVIDIA Cloud Partner (NCP) and a validated leading NCP across both markets, we power production AI for leading AI-native companies including Fireworks AI, Cartesia, Reflection, and Open Router. From large-scale compute to optimized inference and agentic workloads, GMI Cloud gives AI teams the infrastructure they need to build, deploy, and scale on one unified cloud. One cloud for compute, inference, and agents.
Role Overview:
We are seeking a talented and highly skilled Infrastructure Backend Engineering Development Engineer to design, build, and maintain the scalable infrastructure that supports GMI AI/ML initiatives. The ideal candidate will have a strong background in cloud computing, distributed systems, and Dev Ops practices to enable efficient AI infrastructure operations.
Responsibilities:
Design, implement, and maintain AI/ML infrastructure optimized for large-scale training and inference. Develop automation pipelines for GPU/CPU resource provisioning and workload scheduling using Dev Ops best practices and methodologies. Develop observability and telemetry solutions to pro-actively monitor hardware performance, utilization, and health to ensure cluster reliability and efficiency. Optimize infrastructure for high-throughput data transfer and low-latency communication. Manage infrastructure security, access controls, and compliance standards for on-prem GPU cluster environments. Collaborate with relevant engineering teams to configure and troubleshoot GPU clusters and hardware resources. Document infrastructure architecture, deployment procedures, automation workflow, and operational best practices. Stay current with the latest GPU technology developments, infrastructure engineering and integrate new hardware/software solutions as appropriate.
Qualifications:
Bachelor’s degree in Computer Science or related field. Proficiency in at least one programming language (Golang, Python, Bash) with strong coding practices and system design skills. Extensive experience with infrastructure orchestration platforms, especially Open Stack and Kubernetes. Strong proficiency in automation, configuration management, and CI/CD pipelines using Ansible, Jenkins, Git Lab CI, or similar. Proven experience implementing telemetry and observability solutions (e.g. Prometheus, Grafana, Mimir and related technologies).Knowledge of networking, security, and performance tuning in GPU clusters. Hands-on experience in deploying GPU clusters and managing GPU workloads. Experience with secret management using Hashi Corp Vault. Experience with distributed and high performance storage systems (e.g. Vastdata, Weka, DDN, Ceph)Strong system thinking and abstraction skills, capable of designing complex distributed systems from an end-to-end perspective. Strong understanding of Dev Ops principles, automation, and cloud-native architectures. Meeting every qualification is not required—if you’re excited about this role, we’d love to hear from you. We believe diverse perspectives and experiences strengthen our team.
Also on the board Same function, level within a rung
Level
Lead
Location
Chicago, IL
Occupation
Computer Systems Engineers/Architects
Industry
Computer Systems Design Services
Posted
yesterday