Top Reliability Engineer Jobs in Chicago, IL

16 Days AgoSaved
Hybrid
Chicago, IL
Mid level
Mid level
Financial Services
Design, implement, and maintain reliable, scalable cloud-native platforms using infrastructure-as-code and CI/CD. Build observability, define SLOs/SLIs, troubleshoot incidents, reduce toil, and collaborate with engineering teams to improve availability and performance.
Top Skills: AnsibleAWSDatadogDockerDynatraceEcsGitlabGrafanaJavaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
Reposted 6 Days AgoSaved
Easy Apply
Remote
Chicago, IL
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
7 Days AgoSaved
Remote or Hybrid
Chicago, IL
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
12 Days AgoSaved
In-Office
Chicago, IL
100K-120K Annually
Senior level
100K-120K Annually
Senior level
Real Estate • Financial Services
Provide reliability engineering support for building and equipment assets: integrate BAS and energy systems, implement RCM/CbM/PdM strategies, perform RCFA, optimize PMs in CMMS, conduct condition assessments, support capital planning, and drive continuous reliability improvements across client facilities.
Top Skills: Automated Fault Detection And DiagnosticsBomBuilding Automation Systems (Bas)Computerized Maintenance Management Systems (Cmms)Condition-Based Maintenance (Cbm)Energy Management SystemsInfrared ThermographyExcelMotor Current AnalysisMtbfOil AnalysisPdmPredictive MaintenanceRcfaRcmReliability Centered Maintenance (Rcm)UltrasoundVibration Analysis
Reposted 2 Days AgoSaved
In-Office or Remote
Chicago, IL
119K-178K Annually
Senior level
119K-178K Annually
Senior level
Automotive • Information Technology • Other • Transportation • Energy
Perform RAM and FMECA/FMEA analyses, develop fault trees and reliability predictions, support maintainability and logistics analyses, produce reliability growth test plans, contribute to systems engineering documentation, advise design engineers on R&M shortfalls, and present results to management and clients.
Top Skills: Fault Tree AnalysisFmeaFmecaIntegrated Logistics Support (Ils/Ilsa)Iso-9000Mil-Hdbk-217FRam ModellingRam SoftwareStatistical Methods
Reposted 2 Days AgoSaved
In-Office or Remote
Chicago, IL
Senior level
Senior level
Artificial Intelligence • Cloud • Information Technology • Software
The Site Reliability Engineer will provision and manage Kubernetes clusters, build automation tools, debug customer issues, and improve infrastructure reliability.
Top Skills: AnsibleBashDatadogGoGrafanaHelmKubernetesLokiPrometheusPythonTerraform
Reposted 2 Days AgoSaved
Remote
Chicago, IL
190K-240K Annually
Senior level
190K-240K Annually
Senior level
Artificial Intelligence • Insurance • Software • Automation
Lead design, automation, and optimization of database infrastructure (PostgreSQL/Aurora). Build monitoring, tuning, and scaling strategies, create automation tooling, drive performance and reliability initiatives, and expand into broader SRE responsibilities to improve availability and system health for a growing SaaS platform.
Top Skills: Amazon AuroraCi/CdDockerJavaScriptKubernetesNode.jsPostgresPrismaRedshiftTerraformTerragruntTypescript
Reposted 8 Days AgoSaved
Easy Apply
Remote or Hybrid
Chicago, IL
Easy Apply
126K-248K Annually
Senior level
126K-248K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills: AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Reposted 10 Days AgoSaved
Remote or Hybrid
Chicago, IL
200K-230K Annually
Senior level
200K-230K Annually
Senior level
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills: Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
Reposted 10 Days AgoSaved
Remote or Hybrid
Chicago, IL
175K-200K Annually
Senior level
175K-200K Annually
Senior level
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills: AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
16 Days AgoSaved
In-Office
Chicago, IL
75K-85K Annually
Mid level
75K-85K Annually
Mid level
Cloud • Information Technology • Consulting • Design • Generative AI
Build and maintain Python-first automation and tooling for hyperscale data center, backbone, and out-of-band networks. Operate multi-vendor routing/switching (BGP, IS-IS/OSPF, EVPN/VXLAN, MPLS), implement telemetry and observability (gNMI/OpenConfig, SNMP, flow), codify config-as-code (Jinja2, NETCONF/YANG, REST APIs), carry production on-call, lead incident response and RCA, and support site turn-ups, capacity expansion, and secure change governance.
Top Skills: 802.1XAnsibleArista EosBgpCienaCisco Ios-XrCisco Nx-OsDwdmEcmpEvpnFlow TelemetryGitGithub ActionsGitlab CiGnmiGrpcIpsecIs-IsIxiaJenkinsJinja2Juniper JunosLinuxLlmMacsecMplsNacNapalmNcclientNetboxNetconfNetmikoNornirOltOnuOpenconfigOpengearOspfPonPyats/GeniePythonRest ApiRestconfRs-232ScrapliSnmpSpirentSystemdVxlanYangZpe Nodegrid
17 Days AgoSaved
In-Office
Chicago, IL
120K-130K Annually
Senior level
120K-130K Annually
Senior level
Other • Industrial
Lead implementation and sustainment of Remote Monitoring Centre and APM system to improve reliability and lifecycle performance of fixed plant and mobile equipment. Develop asset health models, guide OT/IT integrations, enable condition‑based and predictive maintenance, validate analytics outputs, perform RCA/RCM/FMEA, translate findings into maintenance strategies, and mentor teams. Remote role with 20–40% travel to operations.
Top Skills: Apm SystemsCmmsCondition MonitoringData HistoriansData PipelinesMachine LearningOil AnalysisOt/It IntegrationPredictive AnalyticsRemote MonitoringTemperature/Pressure SensorsVibration Analysis
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
Expert/Leader
3D Printing
Develop and implement reliability and qualification test programs for AI server power modules; define component selection and qualification processes; create test plans and validation documentation; coordinate with mechanical, software, and systems teams; and work with internal and external manufacturing partners to prototype and transition designs to production.
Top Skills: PcbaPower ModulesPower Semiconductors
Reposted 24 Days AgoSaved
Hybrid
Chicago, IL
103K-193K Annually
Senior level
103K-193K Annually
Senior level
Automotive • Hardware • Internet of Things • Mobile • Software • App development • PropTech
Design, implement, and optimize global cloud infrastructure and platforms for an IoT service. Lead platform improvement initiatives, automate infrastructure (IaC/GitOps), ensure observability and security, troubleshoot incidents, mentor SRE team members, and collaborate with executives, architects, and security stakeholders to execute the infrastructure roadmap.
Top Skills: Active DirectoryArgocdAWSBashDatadogEdge FirewallsGitopsGoGrafanaIacKubernetesLinuxNew RelicPowershellPrometheusPythonSIEMTerraformVpcWindows
Reposted 24 Days AgoSaved
Hybrid
Chicago, IL
113K-188K Annually
Senior level
113K-188K Annually
Senior level
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
The Staff Site Reliability Engineer will lead reliability strategies, manage high-risk initiatives, and enhance engineering standards while ensuring system reliability and operational excellence within a hybrid work environment.
Top Skills: BashCi/CdDatabase ArchitectureGoGoogle Cloud PlatformInfrastructure-As-CodeKubernetesMonitoring PlatformsPulumiPythonTerraform
25 Days AgoSaved
Hybrid
Chicago, IL
Mid level
Mid level
Financial Services
Design, implement, and maintain reliable, scalable cloud infrastructure and deployment pipelines. Monitor and optimize application availability using observability, SLOs, and telemetry. Automate infrastructure/configuration as code, troubleshoot containers and networking, collaborate across teams, and apply enterprise-authorized AI to accelerate incident triage and reliability improvements while ensuring data sensitivity.
Top Skills: .NetCi/CdCloudContainersDockerEnterprise AiJavaKubernetesMonitoringNetworkingObservabilityPythonService Level Objectives (Slo)Spring BootTelemetry
Reposted 10 Days AgoSaved
Remote
Chicago, IL
145K-180K Annually
Senior level
145K-180K Annually
Senior level
Legal Tech • Software
Lead automation and optimization of Filevine's data platform: performance tune MSSQL/Postgres, optimize Snowflake, provision infrastructure with Terraform/AWS, run stateful containers on Kubernetes, integrate AI/LLM and MCP for operational automation, manage CI/CD, capacity planning, documentation, and serve in 24/7 on-call rotation.
Top Skills: AWSC#DapperDockerDynamoDBEntity FrameworkGitlabKubernetesLlmsMcp (Model Context Protocol)Microsoft Sql Server (Mssql)Octopus DeployOpensearchPostgresPowershellPythonRedisSnowflakeTerraform
Reposted 20 Days AgoSaved
In-Office
Chicago, IL
Junior
Junior
Information Technology • Consulting
The Customer Reliability Engineer will analyze and provide predictive analytics for power generation and mining equipment, ensuring customer satisfaction and monitoring solutions.
Top Skills: Computer ProgrammingIndustrial EquipmentPredictive Analytics Software
Reposted 20 Days AgoSaved
In-Office
Chicago, IL
Junior
Junior
Information Technology • Software
The role involves delivering predictive analytics solutions for customer accounts, analyzing data, and mastering related software tools. Requires teamwork and customer management skills.
Top Skills: Computer ProgrammingIndustrial EquipmentPredictive Analytics SoftwareProcess EquipmentScripting
Reposted 20 Days AgoSaved
In-Office
Chicago, IL
Junior
Junior
Information Technology • Software
The Customer Reliability Engineer leverages expertise in mechanical engineering and IT to provide predictive analytics for industrial equipment maintenance and performance, primarily in power generation, oil and gas, and mining industries.
Top Skills: Computer ProgrammingMiningOil And GasPower GenerationPredictive Analytics SoftwareScripting
17 Days AgoSaved
Easy Apply
Remote
Chicago, IL
Easy Apply
173K-255K Annually
Senior level
173K-255K Annually
Senior level
Big Data • Fintech • Mobile • Payments • Financial Services
Lead design and delivery of a reliability platform for production systems. Own quarterly goals, guide engineers through ambiguity, collaborate with PM/design/analytics, monitor and operate services (on-call), set quality and code standards, and mentor teammates. Integrate AI/LLM tooling to improve automation, debugging, and service health.
Top Skills: Ai FrameworksAWSKotlinKubernetesLlmsMySQLPythonReactVue
Reposted 17 Days AgoSaved
Remote
Chicago, IL
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills: BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Reposted 18 Days AgoSaved
Easy Apply
Remote or Hybrid
Chicago, IL
Easy Apply
Internship
Internship
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills: AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 23 Days AgoSaved
In-Office
Chicago, IL
160K-220K Annually
Senior level
160K-220K Annually
Senior level
Cloud
The role involves designing, optimizing, and maintaining PostgreSQL and MySQL databases, ensuring high availability, reliability, and performance for mission-critical systems, while automating operational tasks and responding to incidents.
Top Skills: AnsibleAWSDatadogGCPGoGrafanaKubernetesMySQLPostgresPrometheusPythonTerraform
24 Days AgoSaved
In-Office
Chicago, IL
91K-137K Annually
Mid level
91K-137K Annually
Mid level
Fintech • Payments • Financial Services
Lead infrastructure resilience for cloud and SaaS systems by implementing observability, IaC, CI/CD automation, incident triage and self-healing. Build tooling, alerts, AI-driven anomaly detection and runbooks, partner with architecture and infrastructure teams to improve reliability, performance, and security.
Top Skills: Ai/Ml FrameworksAWSCi/CdCloudFormationCloudwatchDevsecopsDynatraceInfrastructure As CodeJavaKubernetesLlmObservabilityOraclePythonSplunkSQL ServerTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account