Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Chicago, IL
Financial Services
Builds, maintains, monitors, and optimizes applications and cloud infrastructure for availability, reliability, and scalability. Designs automated CI/CD and infrastructure-as-code solutions, implements observability and SLOs, troubleshoots incidents and networking issues, and reduces operational toil. The role also guides teammates in SRE best practices, uses authorized AI tools for incident analysis, and collaborates across teams to proactively resolve complex platform and business problems.
Top Skills:
Amazon EcsAnsibleAWSDatadogDockerDynatraceGitlabGrafanaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
Financial Services
Builds, maintains, monitors, and optimizes applications and infrastructure for availability, reliability, scalability, and performance. Designs automated CI/CD and infrastructure-as-code solutions, implements observability and SLOs, responds to incidents, reduces toil, troubleshoots cloud and networking issues, and applies authorized AI tools to improve incident analysis and reliability. Guides teammates in adopting SRE practices across data management and migration platforms.
Top Skills:
Amazon EcsAnsibleAWSDatadogDockerDynatraceGitlabGrafanaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
Artificial Intelligence • Cloud • Internet of Things • Software • Cybersecurity • Industrial
Lead reliability, availability, and performance efforts for eCommerce platforms and infrastructure. Monitor and troubleshoot production and QA systems, support 24/7 incident response, automate deployment and release processes, maintain monitoring dashboards, resolve application performance issues with developers, ensure security and compliance, and continuously improve operational practices.
Top Skills:
AWSAzure DevopsCi/CdCloudFormationDockerGitInfrastructure As CodeJavaScriptKubernetesLoad BalancingNext.JsNode.jsPythonTerraform
Financial Services
Independently execute small-to-medium reliability projects, write maintainable code, triage and resolve incidents, remove operational toil, maintain cloud infrastructure, implement observability and SLOs, use enterprise-authorized AI for troubleshooting and post-incident analysis, and collaborate across teams to improve reliability and CI/CD practices.
Top Skills:
Ci/Cd ToolingCloud InfrastructureContainersEnterprise-Authorized Ai CapabilitiesLinuxObservability (Slo/SliTelemetry)Windows
Digital Media • Information Technology • News + Entertainment
Responsible for ensuring reliability, scalability, and performance of data platforms. Design monitoring and alerting, automate deployments and recovery, optimize storage and query performance, troubleshoot incidents, plan capacity and scaling, document operations, enforce security/compliance, and collaborate with data engineering, product, and data science teams to maintain high availability of large-scale data systems.
Top Skills:
AnsibleAWSAzureCi/CdDockerElk StackGCPGoGrafanaJavaKubernetesMySQLNoSQLPostgresPrometheusPythonScalaTerraform
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own the reliability, performance, resilience, observability, and security of AWS and Kubernetes infrastructure supporting products and AI/ML workloads. Define SLOs, lead incident response and root-cause analysis, build Terraform automation, optimize cloud costs, reduce operational toil, and establish deployment standards that help engineers ship reliably. Participate in on-call rotations and maintain HIPAA-compliant infrastructure.
Top Skills:
AWSClaudeDatadogGitlabGoHipaaIstioKubernetesNatsPostgresPythonSoc 2TerraformTypescript
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
The Staff Site Reliability Engineer will lead reliability strategies, manage high-risk initiatives, and enhance engineering standards while ensuring system reliability and operational excellence within a hybrid work environment.
Top Skills:
BashCi/CdDatabase ArchitectureGoGoogle Cloud PlatformInfrastructure-As-CodeKubernetesMonitoring PlatformsPulumiPythonTerraform
Automotive • Hardware • Internet of Things • Mobile • Software • App development • PropTech
Design, implement, and optimize global cloud infrastructure and platforms for an IoT service. Lead platform improvement initiatives, automate infrastructure (IaC/GitOps), ensure observability and security, troubleshoot incidents, mentor SRE team members, and collaborate with executives, architects, and security stakeholders to execute the infrastructure roadmap.
Top Skills:
Active DirectoryArgocdAWSBashDatadogEdge FirewallsGitopsGoGrafanaIacKubernetesLinuxNew RelicPowershellPrometheusPythonSIEMTerraformVpcWindows
Financial Services
Design, implement, and maintain reliable, scalable cloud infrastructure and deployment pipelines. Monitor and optimize application availability using observability, SLOs, and telemetry. Automate infrastructure/configuration as code, troubleshoot containers and networking, collaborate across teams, and apply enterprise-authorized AI to accelerate incident triage and reliability improvements while ensuring data sensitivity.
Top Skills:
.NetCi/CdCloudContainersDockerEnterprise AiJavaKubernetesMonitoringNetworkingObservabilityPythonService Level Objectives (Slo)Spring BootTelemetry
Reposted 17 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
19 Days AgoSaved
Easy Apply
Easy Apply
Fintech • News + Entertainment • Software • Financial Services
Define tastytrade’s SRE practice, including customer-focused SLOs, error budgets, burn-rate alerts, observability standards, and production readiness reviews. Embed reliability patterns in Ruby, Java, and Elixir services running on HashiCorp Nomad. Extend Prometheus, Honeycomb, and OpenTelemetry observability; conduct fault-injection and tabletop exercises; strengthen on-call and incident-review processes; and mentor engineering teams in reliability practices.
Top Skills:
ConsulElixirGrafanaHashicorp NomadHoneycombJavaLinuxMulticastOpentelemetryPacket CapturePrometheusPythonRubyTcp/IpUdpVault
Reposted 19 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Fintech • Software • Financial Services
Lead the SRE function for NinjaTrader’s trading platform, ensuring availability, scalability, performance, and 99.95% uptime. Responsibilities include managing Kubernetes services, resolving production incidents, participating in a 12x7 on-call rotation, automating deployments and operational tasks, designing monitoring and alerting systems, establishing SLIs/SLOs, using Terraform for infrastructure automation, and implementing cloud security and compliance practices. The role also mentors engineers and collaborates cross-functionally on reliable platform delivery.
Top Skills:
AnsibleAWSAzureBashDatadogDockerGCPGithub ActionsGoGrafanaHelmKubernetesPci DssPrometheusPythonSoc 2Terraform
Reposted 16 Days AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Reposted 26 Days AgoSaved
Easy Apply
Easy Apply
AdTech
Design, build, and scale cloud-native infrastructure with automation-first approach. Develop Terraform modules, Helm charts, Istio routing, observability (Prometheus/Grafana/Datadog), maintain GCP databases, improve CI/CD, and use AI agents to automate and operationalize reliability and developer experience.
Top Skills:
Amazon KinesisAWSAws LambdaAws SnsCi/CdClaude CodeCloudsqlCursorDatadogDockerGCPGitlabGoogle BigqueryGoogle Cloud FunctionsGoogle Cloud RunGoogle Pub/SubGoogle SpannerGrafanaHelmIstioKafkaKubernetesMySQLPrometheusSQLTerraform
Artificial Intelligence • Cloud • Information Technology • Legal Tech • Productivity • Software
The Senior Site Reliability Engineer will focus on automating infrastructure, enhancing cloud resilience, supporting deployments, and mentoring teams in reliability best practices, while participating in on-call rotations.
Top Skills:
AzureBashCi/CdDockerGoGrafanaJavaKubernetesPowershellPrometheusPythonRubyTerraform
Fintech • Software
Lead SRE efforts for DFIN SaaS: ensure availability, performance, scalability, and automation. Implement monitoring, CI/CD, IaC, container orchestration, AI-enhanced observability, incident response, RCA, and runbook automation while collaborating across engineering teams.
Top Skills:
.NetAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC#Ci/CdCloud Ai ServicesContainersCosmosDatadogDynatraceEksFirewallHarnessIdera Sql Diagnostic ManagerInfrastructure As Code (Iac)JavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows
Software
Maintain operational resilience across Azure, AWS, and GCP in a 24x7 environment. Engineer Terraform-based security baselines, optimize CI/CD pipelines, monitor workloads with CSPM tools, and lead major-incident response as Incident Commander. Own remediation through closure, communicate incident updates to technical and executive stakeholders, and develop incident-management playbooks, runbooks, and escalation procedures. Mentor SREs and support compliant platforms subject to PCI-DSS and SOC 2 requirements.
Top Skills:
AWSAzureCi/CdCnappCspmGCPGoKubernetesPagerdutyPci-DssPythonServicenowSoc 2TerraformWiz
Financial Services
Design, implement, and maintain reliable, scalable cloud-native platforms using infrastructure-as-code and CI/CD. Build observability, define SLOs/SLIs, troubleshoot incidents, reduce toil, and collaborate with engineering teams to improve availability and performance.
Top Skills:
AnsibleAWSDatadogDockerDynatraceEcsGitlabGrafanaJavaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
6 Days AgoSaved
Financial Services
Define and implement the company’s SRE practice, including customer-focused SLOs, error budgets, burn-rate alerting, reliability patterns, and observability standards. Embed circuit breakers, retries, bulkheads, and load shedding into Ruby, Java, and Elixir services. Extend Prometheus, Honeycomb, and OpenTelemetry monitoring, support scaling on HashiCorp Nomad, mentor engineers, and establish blameless incident review practices.
Top Skills:
ConsulElixirFlow AnalysisGrafanaHashicorp NomadHoneycombJavaLinuxMulticastOpentelemetryPacket CapturePrometheusPythonRubyTcp/IpUdpVault
Information Technology • Insurance • Professional Services • Software • Analytics
Leads post-incident investigations, root-cause analyses, and preventive strategies to improve reliability and time to resolution. Maintains observability tools, develops client dashboards and alerts, monitors performance metrics, and contributes to incident-response automation. Collaborates with Cloud Operations, SRE, and Engineering teams to enhance SaaS scalability and stability, troubleshoots .NET applications, and provides stakeholder visibility and technical feedback.
Top Skills:
.NetAWSAzureC#Ci/CdDatadogInfrastructure As Code (Iac)JavaScriptNew RelicSaaSSQLSQL ServerSumo LogicWindows
Reposted 25 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Chicago, IL Companies Hiring SRE Engineers
See AllPopular Chicago, IL Engineering Job Searches
Engineering Jobs in Chicago, IL
.NET Developer Jobs in Chicago, IL
Android Developer Jobs in Chicago, IL
Application Engineer Jobs in Chicago, IL
Automation Engineer Jobs in Chicago, IL
Backend Engineer Jobs in Chicago, IL
C# Jobs in Chicago, IL
C++ Jobs in Chicago, IL
Cloud Engineer Jobs in Chicago, IL
Controls Engineer Jobs in Chicago, IL
CTO Jobs in Chicago, IL
Design Engineer Jobs in Chicago, IL
DevOps Engineer Jobs in Chicago, IL
DevOps Jobs in Chicago, IL
Director of Engineering Jobs in Chicago, IL
Director of Software Engineering Jobs in Chicago, IL
Electrical Engineering Jobs in Chicago, IL
Embedded Software Engineer Jobs in Chicago, IL
Engineering Manager Jobs in Chicago, IL
Enterprise Architect Jobs in Chicago, IL
FPGA Engineer Jobs in Chicago, IL
Front End Developer Jobs in Chicago, IL
Full-Stack Engineer Jobs in Chicago, IL
Golang Jobs in Chicago, IL
Hardware Engineer Jobs in Chicago, IL
Infrastructure Engineer Jobs in Chicago, IL
iOS Developer Jobs in Chicago, IL
Java Developer Jobs in Chicago, IL
Java Full-Stack Engineer Jobs in Chicago, IL
Javascript Jobs in Chicago, IL
Lead Software Engineer Jobs in Chicago, IL
Linux Jobs in Chicago, IL
Manufacturing Engineer Jobs in Chicago, IL
Mechanical Design Engineer Jobs in Chicago, IL
Mechanical Engineering Jobs in Chicago, IL
Network Engineer Jobs in Chicago, IL
PHP Developer Jobs in Chicago, IL
Platform Engineer Jobs in Chicago, IL
Principal Engineer Jobs in Chicago, IL
Principal Software Engineer Jobs in Chicago, IL
Process Engineer Jobs in Chicago, IL
Project Engineer Jobs in Chicago, IL
Python Jobs in Chicago, IL
QA Engineer Jobs in Chicago, IL
QA Jobs in Chicago, IL
Reliability Engineer Jobs in Chicago, IL
Robotics Engineer Jobs in Chicago, IL
Ruby Jobs in Chicago, IL
Salesforce Developer Jobs in Chicago, IL
Scala Jobs in Chicago, IL
Security Engineer Jobs in Chicago, IL
Software Engineer Jobs in Chicago, IL
Software Engineering Manager Jobs in Chicago, IL
Software Test Engineer Jobs in Chicago, IL
Solutions Architect Jobs in Chicago, IL
Solutions Engineer Jobs in Chicago, IL
SRE Engineer Jobs in Chicago, IL
Staff Engineer Jobs in Chicago, IL
Staff Software Engineer Jobs in Chicago, IL
Systems Engineer Jobs in Chicago, IL
All Filters
Total selected ()
No Results
No Results



.png)














.png)











