Maximum of 25 job preferences reached.
Top Reliability Engineer Jobs in Chicago, IL
Artificial Intelligence • Cloud • Information Technology • Legal Tech • Productivity • Software
Lead incident response and reliability improvements for iManage Cloud. Triage large-scale production issues, build observability and automation, run postmortems, partner with product and engineering, and proactively detect and eliminate systemic problems to improve uptime and customer experience.
Top Skills:
AzureAzure Kubernetes Service (Aks)BashGrafanaKibanaPowershellPrometheusPythonRest ApisShellSplunkSQL
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Design, deploy, automate, and support large-scale cloud network infrastructure across AWS, Azure, and GCP. Manage firewalls, Kubernetes networking, VPNs, DNS, BGP, monitoring, infrastructure as code, and network performance. Lead technical roadmaps, design reviews, troubleshooting, engineer onboarding, and team direction while participating in on-call support and occasional travel.
Top Skills:
AnsibleAWSAzureBgpCheck PointChefCloudFormationCniDnsFortinetGCPGoJavaKubernetesLinuxPalo Alto NetworksPythonRubySpaceliftTcp/IpTerraformTls
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Support and improve reliability, scalability, and performance of large-scale databases across cloud and on-prem. Build automation, Kubernetes operators, GitOps workflows, and tooling in Go or Python; implement observability, capacity planning, failover, backups, schema migrations, and self-healing systems; partner with application teams and leverage AI to improve operations and engineering productivity.
Top Skills:
AerospikeArgocdAurora MysqlClaudeCursorEksFluxcdGithub CopilotGitopsGkeGoKubernetesKubernetes OperatorsMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills:
AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Real Estate • PropTech
Lead technical strategy and implementation for Redfin's production databases and storage. Architect, scale, and support cloud database/storage solutions (self-managed and AWS managed), drive reliability, observability, backup/recovery, DR planning, incident response and root cause analysis, mentor engineers, and participate in on-call rotation.
Top Skills:
Amazon AuroraAmazon RdsAmazon S3Anthropic Claude CodeAWSConfiguration ManagementCursorDynamoDBElasticacheGithub CopilotInfrastructure As CodeLinuxOpensearchPostgresPython
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead operational reliability and platform enablement for Databricks: build monitoring, CI/CD, deployment standards, compute and job policies, observability, runbooks, and governance to support secure, cost-aware, production data workloads across regulated environments. Mentor engineers and align platform with cloud/infrastructure and compliance requirements.
Top Skills:
Ci/CdDatabricksDatabricks Asset BundlesDatabricks WorkflowsDelta LakeInfrastructure-As-CodeService PrincipalsUnity CatalogVersion Control (Git)
Artificial Intelligence • Big Data • Healthtech • Machine Learning • Analytics • Biotech • Generative AI
Join the SRE team to design, deploy, and operate resilient cloud infrastructure. Recommend solutions, automate workflows, configure Terraform and CI, implement monitoring and alerts, and support developers and users.
Top Skills:
AnsibleAurora MysqlAWSAzureBashChefCloudFormationComposerConcourseDataprocDockerGCPGoHipaaHitrustIsoKubernetesPackerPostgresPuppetPythonRubySaltSlackTerraform
Reposted 5 Hours AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
2 Days AgoSaved
Easy Apply
Easy Apply
Fintech • News + Entertainment • Software • Financial Services
Define tastytrade’s SRE practice, including customer-focused SLOs, error budgets, burn-rate alerts, observability standards, and production readiness reviews. Embed reliability patterns in Ruby, Java, and Elixir services running on HashiCorp Nomad. Extend Prometheus, Honeycomb, and OpenTelemetry observability; conduct fault-injection and tabletop exercises; strengthen on-call and incident-review processes; and mentor engineering teams in reliability practices.
Top Skills:
ConsulElixirGrafanaHashicorp NomadHoneycombJavaLinuxMulticastOpentelemetryPacket CapturePrometheusPythonRubyTcp/IpUdpVault
Reposted 2 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Fintech • Software • Financial Services
Lead the SRE function for NinjaTrader’s trading platform, ensuring availability, scalability, performance, and 99.95% uptime. Responsibilities include managing Kubernetes services, resolving production incidents, participating in a 12x7 on-call rotation, automating deployments and operational tasks, designing monitoring and alerting systems, establishing SLIs/SLOs, using Terraform for infrastructure automation, and implementing cloud security and compliance practices. The role also mentors engineers and collaborates cross-functionally on reliable platform delivery.
Top Skills:
AnsibleAWSAzureBashDatadogDockerGCPGithub ActionsGoGrafanaHelmKubernetesPci DssPrometheusPythonSoc 2Terraform
Digital Media • Information Technology • News + Entertainment
Responsible for ensuring reliability, scalability, and performance of data platforms. Design monitoring and alerting, automate deployments and recovery, optimize storage and query performance, troubleshoot incidents, plan capacity and scaling, document operations, enforce security/compliance, and collaborate with data engineering, product, and data science teams to maintain high availability of large-scale data systems.
Top Skills:
AnsibleAWSAzureCi/CdDockerElk StackGCPGoGrafanaJavaKubernetesMySQLNoSQLPostgresPrometheusPythonScalaTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Artificial Intelligence • Big Data • Enterprise Web • Fintech • Software • Financial Services
Lead technical ownership and resolution of complex production issues; coordinate cross-functional incident response; troubleshoot C#/.NET and JavaScript systems; analyze logs/metrics (Splunk, New Relic); drive reliability, observability, and customer-first improvements while mentoring teams.
Top Skills:
.Net.Net CoreAIApm ToolsAWSC#JavaScriptNew RelicSplunk
Aerospace • Logistics • Security • Software • Cybersecurity
Performs reliability and maintainability engineering for avionics mission systems. Responsibilities include managing FRACAS, documenting and trending failure data, conducting root-cause and corrective-action investigations, performing reliability predictions, and analyzing FMECA and MTTR. The role supports subsystem reliability requirements and serves as an individual contributor and deliverable lead within a multidisciplinary engineering team.
Top Skills:
FmecaFracasExcelMicrosoft PowerpointMicrosoft Word
Fintech
Leads observability engineering for critical customer journeys, defining SLIs, SLOs, error budgets, dashboards, alerts, synthetic monitoring, and telemetry standards. Partners with product, engineering, SRE, operations, and business teams to improve production readiness and service reliability. Analyzes incidents, telemetry, alert performance, and customer impact to identify gaps, reduce alert fatigue, and drive continuous improvement. Provides technical leadership and mentorship across observability practices.
Top Skills:
ApmCloud PlatformsDatadogDistributed SystemsDynatraceElasticGrafanaKubernetesLoggingMicroservicesNew RelicOpentelemetryPrometheusRumSplunkSynthetic MonitoringTelemetry FrameworksTracing
Reposted 9 Days AgoSaved
Easy Apply
Easy Apply
AdTech
Design, build, and scale cloud-native infrastructure with automation-first approach. Develop Terraform modules, Helm charts, Istio routing, observability (Prometheus/Grafana/Datadog), maintain GCP databases, improve CI/CD, and use AI agents to automate and operationalize reliability and developer experience.
Top Skills:
Amazon KinesisAWSAws LambdaAws SnsCi/CdClaude CodeCloudsqlCursorDatadogDockerGCPGitlabGoogle BigqueryGoogle Cloud FunctionsGoogle Cloud RunGoogle Pub/SubGoogle SpannerGrafanaHelmIstioKafkaKubernetesMySQLPrometheusSQLTerraform
Fintech • Payments • Financial Services
The Staff Reliability Engineer will enhance data platform reliability through automation, incident management, and observability in a hybrid work setting.
Top Skills:
AiopsAnsibleAWSCi/CdCloudFormationDatadogDynatraceEksEmrGCPGrafanaHadoopOpensearchPrometheusPythonSnowflakeSplunkTerraform
Financial Services
Independently execute small-to-medium reliability projects, write maintainable code, triage and resolve incidents, remove operational toil, maintain cloud infrastructure, implement observability and SLOs, use enterprise-authorized AI for troubleshooting and post-incident analysis, and collaborate across teams to improve reliability and CI/CD practices.
Top Skills:
Ci/Cd ToolingCloud InfrastructureContainersEnterprise-Authorized Ai CapabilitiesLinuxObservability (Slo/SliTelemetry)Windows
Artificial Intelligence • Cloud • Information Technology • Legal Tech • Productivity • Software
The Senior Site Reliability Engineer will focus on automating infrastructure, enhancing cloud resilience, supporting deployments, and mentoring teams in reliability best practices, while participating in on-call rotations.
Top Skills:
AzureBashCi/CdDockerGoGrafanaJavaKubernetesPowershellPrometheusPythonRubyTerraform
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own Garner’s cloud reliability strategy across AWS and Kubernetes, including SLOs, observability, incident response, infrastructure automation, cost optimization, and security compliance. Lead complex incident resolution, architect Terraform-based infrastructure, establish deployment and monitoring standards, mentor engineers, and use AI tools to automate operational work. Support high-scale AI/ML workloads while setting technical direction for platform reliability and production quality.
Top Skills:
AWSClaudeDatadogGitlabGoIstioKubernetesNatsPostgresPythonTerraformTypescript
8 Days AgoSaved
Aerospace • Logistics • Security • Software • Cybersecurity
Lead individual-contributor supporting product supportability, reliability and specialty engineering for avionic mission systems. Perform LORA, LCCA, provisioning/logistics data, packaging, manpower and sparing analyses; support reliability/maintainability predictions, FMECA, testability, human factors and system safety; present supportability at design reviews and coordinate systems engineering tasks and technical budgets.
Top Skills:
Built-In-Test (Bit) DesignContract Data Requirements List (Cdrl)Data Item Description (Did)Failure Modes Effects And Criticality Analysis (Fmeca)Level Of Repair Analysis (Lora)Life Cycle Cost Analysis (Lcca)ExcelMicrosoft PowerpointMicrosoft Word
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead the design and roadmap for global Active Directory and identity infrastructure, implement Identity-as-Code and GitOps automation, own incident escalation and observability, define delegation/tiered administration, integrate applications with Okta and cloud identity, mentor teams, and publish identity architecture and security best practices.
Top Skills:
Active Directory Domain Services (Ad Ds)AnsibleAWSAws Directory ServiceAzureAzure Active Directory (Entra Id)Azure SentinelCertificate ServicesChefDhcpDnsGCPGitopsGroup Policy Objects (Gpo)New RelicOktaPowershellPowershell DscPythonTerraform
Renewable Energy
Own reliability, performance, and scalability of Postgres and ClickHouse databases. Build scalable data pipelines, design analytical schemas and DBT models, migrate data to ClickHouse, implement data quality checks, eliminate duplicates, and manage database infrastructure via IaC.
Top Skills:
Aws CdkAws Step FunctionsCi/CdClickhouseDagsterDbtPostgresPulumiPythonSQL
Software
As a Senior DevOps / Platform Reliability Engineer, you will manage CI/CD pipelines, automate infrastructure, operate Kubernetes, and enhance observability while ensuring security and compliance for enterprise systems.
Top Skills:
Argo CdAurora MysqlAWSBashCloudFormationEksElasticacheGithub ActionsGrafanaKubernetesLinuxMskOpentelemetryPrometheusPythonS3Terraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Chicago, IL Companies Hiring Reliability Engineers
See AllPopular Chicago, IL Engineering Job Searches
Engineering Jobs in Chicago, IL
.NET Developer Jobs in Chicago, IL
Android Developer Jobs in Chicago, IL
Application Engineer Jobs in Chicago, IL
Automation Engineer Jobs in Chicago, IL
Backend Engineer Jobs in Chicago, IL
C# Jobs in Chicago, IL
C++ Jobs in Chicago, IL
Cloud Engineer Jobs in Chicago, IL
Controls Engineer Jobs in Chicago, IL
CTO Jobs in Chicago, IL
Design Engineer Jobs in Chicago, IL
DevOps Engineer Jobs in Chicago, IL
DevOps Jobs in Chicago, IL
Director of Engineering Jobs in Chicago, IL
Director of Software Engineering Jobs in Chicago, IL
Electrical Engineering Jobs in Chicago, IL
Embedded Software Engineer Jobs in Chicago, IL
Engineering Manager Jobs in Chicago, IL
Enterprise Architect Jobs in Chicago, IL
FPGA Engineer Jobs in Chicago, IL
Front End Developer Jobs in Chicago, IL
Full-Stack Engineer Jobs in Chicago, IL
Golang Jobs in Chicago, IL
Hardware Engineer Jobs in Chicago, IL
Infrastructure Engineer Jobs in Chicago, IL
iOS Developer Jobs in Chicago, IL
Java Developer Jobs in Chicago, IL
Java Full-Stack Engineer Jobs in Chicago, IL
Javascript Jobs in Chicago, IL
Lead Software Engineer Jobs in Chicago, IL
Linux Jobs in Chicago, IL
Manufacturing Engineer Jobs in Chicago, IL
Mechanical Design Engineer Jobs in Chicago, IL
Mechanical Engineering Jobs in Chicago, IL
Network Engineer Jobs in Chicago, IL
PHP Developer Jobs in Chicago, IL
Platform Engineer Jobs in Chicago, IL
Principal Engineer Jobs in Chicago, IL
Principal Software Engineer Jobs in Chicago, IL
Process Engineer Jobs in Chicago, IL
Project Engineer Jobs in Chicago, IL
Python Jobs in Chicago, IL
QA Engineer Jobs in Chicago, IL
QA Jobs in Chicago, IL
Reliability Engineer Jobs in Chicago, IL
Robotics Engineer Jobs in Chicago, IL
Ruby Jobs in Chicago, IL
Salesforce Developer Jobs in Chicago, IL
Scala Jobs in Chicago, IL
Security Engineer Jobs in Chicago, IL
Software Engineer Jobs in Chicago, IL
Software Engineering Manager Jobs in Chicago, IL
Software Test Engineer Jobs in Chicago, IL
Solutions Architect Jobs in Chicago, IL
Solutions Engineer Jobs in Chicago, IL
SRE Engineer Jobs in Chicago, IL
Staff Engineer Jobs in Chicago, IL
Staff Software Engineer Jobs in Chicago, IL
Systems Engineer Jobs in Chicago, IL
All Filters
Total selected ()
No Results
No Results















.png)














