Senior Site Reliability Engineering– AI Infrastructure

Senior Site Reliability Engineering– AI Infrastructure

hcltechChicago, IL

yesterday

$78,000/Annum

Computer Systems Engineers/ArchitectsNetwork and Computer Systems AdministratorsComputer and Information Systems Managers
Computer Systems Design ServicesComputing Infrastructure Providers, Data Processing, Web Hosting, and Related ServicesOther Computer Related Services
Apply for this role →
HCLTech is looking for a highly talented and self- motivated Senior Site Reliability Engineering– AI Infrastructureto join it in advancing the technological world through innovation and creativity. Job Title: Senior Site Reliability Engineering– AI Infrastructure
Job ID: 158107
Position Type: Full-time
Location: Remote Role/
Responsibilities: Engagement summary The Candidate will provide SRE services for AI platforms and supporting infrastructure with emphasis on reliability engineering, incident response, service health, and operational automation. This role is best suited to a senior hands-on engineer who can improve availability while remaining effective in detailed production troubleshooting. What this Candidate will be doing • Operate and improve reliability of AI platform services, cluster dependencies, and shared infrastructure components.• Lead or support incident triage for service degradation involving Kubernetes, Linux hosts, storage, network, scheduling, job orchestration, or dependency failures.• Define and refine SLIs, SLOs, alerting thresholds, runbooks, escalation paths, and post-incident actions.• Analyze recurring failure patterns and convert manual operations into automation and preventive controls.• Build observability across system, service, workload, and dependency layers using metrics, logs, traces, and event correlation.• Troubleshoot performance and availability issues affecting training jobs, inference services, internal platforms, and support tooling. • Partner with infrastructure and validation teams to improve production readiness and change safety.• Drive operational reviews, readiness criteria, and resilience testing. What we need to see • 7+ years in SRE, production operations, or reliability-focused infrastructure engineering.• Strong hands-on troubleshooting across Linux, Kubernetes, networking, and distributed systems.• Experience building observability, alerting, and response workflows in complex production environments.• Ability to balance urgent operational response with medium-term reliability engineering improvements.• Strong scripting and automation skills, with experience reducing toil through tooling.• Experience participating in incident management, root cause analysis, and post-incident follow through.• Strong communication skill with the ability to summarize technical issues clearly for cross functional teams.
Preferred experience: • Experience in AI platforms, ML infrastructure, or large-scale HPC-like service environments.• Familiarity with Prometheus, Grafana, ELK/Open Search, Loki, Pager Duty, and incident tooling.• Experience defining error budgets and applying SRE practices in environments with heavy batch and service traffic.
Pay and Benefits: Pay Range Minimum: $78,000/Annum Pay Range Maximum: $148,000/AnnumHCLTech is an equal opportunity employer, committed to providing equal employment opportunities to all applicants and employees regardless of race, religion, sex, color, age, national origin, pregnancy, sexual orientation, physical disability or genetic information, military or veteran status, or any other protected classification, in accordance with federal, state, and/or local law. Should any applicant have concerns about discrimination in the hiring process, they should provide a detailed report of those concerns to for investigation.
Compensation and Benefits: A candidate’s pay within the range will depend on their work location, skills, experience, education, and other factors permitted by law. This role may also be eligible for performance-based bonuses subject to company policies. In addition, this role is eligible for the following benefits subject to company policies: medical, dental, vision, pharmacy, life, accidental death & dismemberment, and disability insurance; employee assistance program; 401(k) retirement plan; 10 days of paid time off per year (some positions are eligible for need-based leave with no designated number of leave days per year); and 10 paid holidays per year. How You’ll Grow At HCLTech, we offer continuous opportunities for you to find your spark and grow with us. We want you to be happy and satisfied with your role and to really learn what type of work sparks your brilliance the best. Throughout your time with us, we offer transparent communication with senior-level employees, learning and career development programs at every level, and opportunities to experiment in different roles or even pivot industries. We believe that you should be in control of your career with unlimited opportunities to find the role that fits you best.

Also on the board Same function, level within a rung

Level

Lead

Salary

$78,000/Annum

Location

Chicago, IL

Occupation

Computer Systems Engineers/Architects

Industry

Computer Systems Design Services

Posted

yesterday

Apply for this role →
Senior Site Reliability Engineering– AI Infrastructure at hcltech...