Top SRE Engineer Jobs in Chicago, IL

5 Days AgoSaved
Hybrid
Chicago, IL
Mid level
Mid level
Financial Services
Builds, maintains, monitors, and optimizes applications and cloud infrastructure for availability, reliability, and scalability. Designs automated CI/CD and infrastructure-as-code solutions, implements observability and SLOs, troubleshoots incidents and networking issues, and reduces operational toil. The role also guides teammates in SRE best practices, uses authorized AI tools for incident analysis, and collaborates across teams to proactively resolve complex platform and business problems.
Top Skills: Amazon EcsAnsibleAWSDatadogDockerDynatraceGitlabGrafanaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
5 Days AgoSaved
Hybrid
Chicago, IL
Mid level
Mid level
Financial Services
Builds, maintains, monitors, and optimizes applications and infrastructure for availability, reliability, scalability, and performance. Designs automated CI/CD and infrastructure-as-code solutions, implements observability and SLOs, responds to incidents, reduces toil, troubleshoots cloud and networking issues, and applies authorized AI tools to improve incident analysis and reliability. Guides teammates in adopting SRE practices across data management and migration platforms.
Top Skills: Amazon EcsAnsibleAWSDatadogDockerDynatraceGitlabGrafanaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
7 Days AgoSaved
Hybrid
Chicago, IL
128K-209K Annually
Senior level
128K-209K Annually
Senior level
Artificial Intelligence • Cloud • Internet of Things • Software • Cybersecurity • Industrial
Lead reliability, availability, and performance efforts for eCommerce platforms and infrastructure. Monitor and troubleshoot production and QA systems, support 24/7 incident response, automate deployment and release processes, maintain monitoring dashboards, resolve application performance issues with developers, ensure security and compliance, and continuously improve operational practices.
Top Skills: AWSAzure DevopsCi/CdCloudFormationDockerGitInfrastructure As CodeJavaScriptKubernetesLoad BalancingNext.JsNode.jsPythonTerraform
Reposted 8 Days AgoSaved
Hybrid
Chicago, IL
Junior
Junior
Financial Services
Independently execute small-to-medium reliability projects, write maintainable code, triage and resolve incidents, remove operational toil, maintain cloud infrastructure, implement observability and SLOs, use enterprise-authorized AI for troubleshooting and post-incident analysis, and collaborate across teams to improve reliability and CI/CD practices.
Top Skills: Ci/Cd ToolingCloud InfrastructureContainersEnterprise-Authorized Ai CapabilitiesLinuxObservability (Slo/SliTelemetry)Windows
Reposted 13 Days AgoSaved
Hybrid
Chicago, IL
118K-176K Annually
Senior level
118K-176K Annually
Senior level
Digital Media • Information Technology • News + Entertainment
Responsible for ensuring reliability, scalability, and performance of data platforms. Design monitoring and alerting, automate deployments and recovery, optimize storage and query performance, troubleshoot incidents, plan capacity and scaling, document operations, enforce security/compliance, and collaborate with data engineering, product, and data science teams to maintain high availability of large-scale data systems.
Top Skills: AnsibleAWSAzureCi/CdDockerElk StackGCPGoGrafanaJavaKubernetesMySQLNoSQLPostgresPrometheusPythonScalaTerraform
6 Days AgoSaved
Easy Apply
Remote
Chicago, IL
Easy Apply
191K-226K Annually
Senior level
191K-226K Annually
Senior level
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills: AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Reposted 15 Days AgoSaved
Hybrid
Chicago, IL
113K-188K Annually
Senior level
113K-188K Annually
Senior level
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
The Staff Site Reliability Engineer will lead reliability strategies, manage high-risk initiatives, and enhance engineering standards while ensuring system reliability and operational excellence within a hybrid work environment.
Top Skills: BashCi/CdDatabase ArchitectureGoGoogle Cloud PlatformInfrastructure-As-CodeKubernetesMonitoring PlatformsPulumiPythonTerraform
Reposted 15 Days AgoSaved
Hybrid
Chicago, IL
103K-193K Annually
Senior level
103K-193K Annually
Senior level
Automotive • Hardware • Internet of Things • Mobile • Software • App development • PropTech
Design, implement, and optimize global cloud infrastructure and platforms for an IoT service. Lead platform improvement initiatives, automate infrastructure (IaC/GitOps), ensure observability and security, troubleshoot incidents, mentor SRE team members, and collaborate with executives, architects, and security stakeholders to execute the infrastructure roadmap.
Top Skills: Active DirectoryArgocdAWSBashDatadogEdge FirewallsGitopsGoGrafanaIacKubernetesLinuxNew RelicPowershellPrometheusPythonSIEMTerraformVpcWindows
Reposted 16 Days AgoSaved
Hybrid
Chicago, IL
Mid level
Mid level
Financial Services
Design, implement, and maintain reliable, scalable cloud infrastructure and deployment pipelines. Monitor and optimize application availability using observability, SLOs, and telemetry. Automate infrastructure/configuration as code, troubleshoot containers and networking, collaborate across teams, and apply enterprise-authorized AI to accelerate incident triage and reliability improvements while ensuring data sensitivity.
Top Skills: .NetCi/CdCloudContainersDockerEnterprise AiJavaKubernetesMonitoringNetworkingObservabilityPythonService Level Objectives (Slo)Spring BootTelemetry
Reposted 17 Days AgoSaved
Easy Apply
Remote or Hybrid
Chicago, IL
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills: AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
19 Days AgoSaved
Easy Apply
Hybrid
Chicago, IL
Easy Apply
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Fintech • News + Entertainment • Software • Financial Services
Define tastytrade’s SRE practice, including customer-focused SLOs, error budgets, burn-rate alerts, observability standards, and production readiness reviews. Embed reliability patterns in Ruby, Java, and Elixir services running on HashiCorp Nomad. Extend Prometheus, Honeycomb, and OpenTelemetry observability; conduct fault-injection and tabletop exercises; strengthen on-call and incident-review processes; and mentor engineering teams in reliability practices.
Top Skills: ConsulElixirGrafanaHashicorp NomadHoneycombJavaLinuxMulticastOpentelemetryPacket CapturePrometheusPythonRubyTcp/IpUdpVault
Reposted 19 Days AgoSaved
Easy Apply
Remote or Hybrid
Chicago, IL
Easy Apply
127K-249K Annually
Senior level
127K-249K Annually
Senior level
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills: AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
20 Days AgoSaved
Easy Apply
Hybrid
Chicago, IL
Easy Apply
160K-210K Annually
Senior level
160K-210K Annually
Senior level
Fintech • Software • Financial Services
Lead the SRE function for NinjaTrader’s trading platform, ensuring availability, scalability, performance, and 99.95% uptime. Responsibilities include managing Kubernetes services, resolving production incidents, participating in a 12x7 on-call rotation, automating deployments and operational tasks, designing monitoring and alerting systems, establishing SLIs/SLOs, using Terraform for infrastructure automation, and implementing cloud security and compliance practices. The role also mentors engineers and collaborates cross-functionally on reliable platform delivery.
Top Skills: AnsibleAWSAzureBashDatadogDockerGCPGithub ActionsGoGrafanaHelmKubernetesPci DssPrometheusPythonSoc 2Terraform
Reposted 16 Days AgoSaved
Remote
Chicago, IL
180K-220K Annually
Senior level
180K-220K Annually
Senior level
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills: AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 26 Days AgoSaved
Easy Apply
In-Office
Chicago, IL
Easy Apply
Senior level
Senior level
AdTech
Design, build, and scale cloud-native infrastructure with automation-first approach. Develop Terraform modules, Helm charts, Istio routing, observability (Prometheus/Grafana/Datadog), maintain GCP databases, improve CI/CD, and use AI agents to automate and operationalize reliability and developer experience.
Top Skills: Amazon KinesisAWSAws LambdaAws SnsCi/CdClaude CodeCloudsqlCursorDatadogDockerGCPGitlabGoogle BigqueryGoogle Cloud FunctionsGoogle Cloud RunGoogle Pub/SubGoogle SpannerGrafanaHelmIstioKafkaKubernetesMySQLPrometheusSQLTerraform
Reposted 27 Days AgoSaved
Hybrid
Chicago, IL
130K-180K Annually
Senior level
130K-180K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Legal Tech • Productivity • Software
The Senior Site Reliability Engineer will focus on automating infrastructure, enhancing cloud resilience, supporting deployments, and mentoring teams in reliability best practices, while participating in on-call rotations.
Top Skills: AzureBashCi/CdDockerGoGrafanaJavaKubernetesPowershellPrometheusPythonRubyTerraform
Reposted 21 Days AgoSaved
Remote or Hybrid
Chicago, IL
Senior level
Senior level
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills: .NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Reposted 3 Days AgoSaved
In-Office
Chicago, IL
Senior level
Senior level
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Engineer Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with CSPM tools, and lead major-incident response as Incident Commander. Own remediation through closure, communicate incident updates to technical and executive stakeholders, and develop incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and support compliant platforms subject to PCI-DSS and SOC 2 requirements.
Top Skills: AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
One Month AgoSaved
Hybrid
Chicago, IL
Mid level
Mid level
Financial Services
Design, implement, and maintain reliable, scalable cloud-native platforms using infrastructure-as-code and CI/CD. Build observability, define SLOs/SLIs, troubleshoot incidents, reduce toil, and collaborate with engineering teams to improve availability and performance.
Top Skills: AnsibleAWSDatadogDockerDynatraceEcsGitlabGrafanaJavaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
Reposted 23 Days AgoSaved
Easy Apply
Remote
Chicago, IL
Easy Apply
100K-110K Annually
Mid level
100K-110K Annually
Mid level
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills: Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
24 Days AgoSaved
Remote or Hybrid
Chicago, IL
140K-215K Annually
Senior level
140K-215K Annually
Senior level
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills: Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
6 Days AgoSaved
In-Office
Chicago, IL
180K-200K Annually
Senior level
180K-200K Annually
Senior level
Financial Services
Define and implement the company’s SRE practice, including customer-focused SLOs, error budgets, burn-rate alerting, reliability patterns, and observability standards. Embed circuit breakers, retries, bulkheads, and load shedding into Ruby, Java, and Elixir services. Extend Prometheus, Honeycomb, and OpenTelemetry monitoring, support scaling on HashiCorp Nomad, mentor engineers, and establish blameless incident review practices.
Top Skills: ConsulElixirFlow AnalysisGrafanaHashicorp NomadHoneycombJavaLinuxMulticastOpentelemetryPacket CapturePrometheusPythonRubyTcp/IpUdpVault
6 Days AgoSaved
Remote or Hybrid
Chicago, IL
100K-120K Annually
Senior level
100K-120K Annually
Senior level
Information Technology • Insurance • Professional Services • Software • Analytics
Leads post-incident investigations, root-cause analyses, and preventive strategies to improve reliability and time to resolution. Maintains observability tools, develops client dashboards and alerts, monitors performance metrics, and contributes to incident-response automation. Collaborates with Cloud Operations, SRE, and Engineering teams to enhance SaaS scalability and stability, troubleshoots .NET applications, and provides stakeholder visibility and technical feedback.
Top Skills: .NetAWSAzureC#Ci/CdDatadogInfrastructure As Code (Iac)JavaScriptNew RelicSaaSSQLSQL ServerSumo LogicWindows
Reposted 25 Days AgoSaved
Easy Apply
Remote or Hybrid
Chicago, IL
Easy Apply
126K-248K Annually
Senior level
126K-248K Annually
Senior level
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills: AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Reposted 27 Days AgoSaved
Remote or Hybrid
Chicago, IL
175K-200K Annually
Senior level
175K-200K Annually
Senior level
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills: AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account