Maximum of 25 job preferences reached.
Top Senior Site Reliability Engineer Jobs in Chicago, IL
Big Data • Fintech • Information Technology • Business Intelligence • Financial Services • Cybersecurity • Big Data Analytics
The Staff Site Reliability Engineer will lead reliability strategies, manage high-risk initiatives, and enhance engineering standards while ensuring system reliability and operational excellence within a hybrid work environment.
Top Skills:
BashCi/CdDatabase ArchitectureGoGoogle Cloud PlatformInfrastructure-As-CodeKubernetesMonitoring PlatformsPulumiPythonTerraform
Automotive • Hardware • Internet of Things • Mobile • Software • App development • PropTech
Design, implement, and optimize global cloud infrastructure and platforms for an IoT service. Lead platform improvement initiatives, automate infrastructure (IaC/GitOps), ensure observability and security, troubleshoot incidents, mentor SRE team members, and collaborate with executives, architects, and security stakeholders to execute the infrastructure roadmap.
Top Skills:
Active DirectoryArgocdAWSBashDatadogEdge FirewallsGitopsGoGrafanaIacKubernetesLinuxNew RelicPowershellPrometheusPythonSIEMTerraformVpcWindows
Financial Services
Design, implement, and maintain reliable, scalable cloud infrastructure and deployment pipelines. Monitor and optimize application availability using observability, SLOs, and telemetry. Automate infrastructure/configuration as code, troubleshoot containers and networking, collaborate across teams, and apply enterprise-authorized AI to accelerate incident triage and reliability improvements while ensuring data sensitivity.
Top Skills:
.NetCi/CdCloudContainersDockerEnterprise AiJavaKubernetesMonitoringNetworkingObservabilityPythonService Level Objectives (Slo)Spring BootTelemetry
Artificial Intelligence • Big Data • Healthtech • Machine Learning • Analytics • Biotech • Generative AI
Join the SRE team to design, deploy, and operate resilient cloud infrastructure. Recommend solutions, automate workflows, configure Terraform and CI, implement monitoring and alerts, and support developers and users.
Top Skills:
AnsibleAurora MysqlAWSAzureBashChefCloudFormationComposerConcourseDataprocDockerGCPGoHipaaHitrustIsoKubernetesPackerPostgresPuppetPythonRubySaltSlackTerraform
Reposted 2 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
Fintech • Software • Financial Services
Lead the SRE function for NinjaTrader’s trading platform, ensuring availability, scalability, performance, and 99.95% uptime. Responsibilities include managing Kubernetes services, resolving production incidents, participating in a 12x7 on-call rotation, automating deployments and operational tasks, designing monitoring and alerting systems, establishing SLIs/SLOs, using Terraform for infrastructure automation, and implementing cloud security and compliance practices. The role also mentors engineers and collaborates cross-functionally on reliable platform delivery.
Top Skills:
AnsibleAWSAzureBashDatadogDockerGCPGithub ActionsGoGrafanaHelmKubernetesPci DssPrometheusPythonSoc 2Terraform
Reposted 11 Days AgoSaved
Easy Apply
Easy Apply
AdTech
Design, build, and scale cloud-native infrastructure with automation-first approach. Develop Terraform modules, Helm charts, Istio routing, observability (Prometheus/Grafana/Datadog), maintain GCP databases, improve CI/CD, and use AI agents to automate and operationalize reliability and developer experience.
Top Skills:
Amazon KinesisAWSAws LambdaAws SnsCi/CdClaude CodeCloudsqlCursorDatadogDockerGCPGitlabGoogle BigqueryGoogle Cloud FunctionsGoogle Cloud RunGoogle Pub/SubGoogle SpannerGrafanaHelmIstioKafkaKubernetesMySQLPrometheusSQLTerraform
Financial Services
Independently execute small-to-medium reliability projects, write maintainable code, triage and resolve incidents, remove operational toil, maintain cloud infrastructure, implement observability and SLOs, use enterprise-authorized AI for troubleshooting and post-incident analysis, and collaborate across teams to improve reliability and CI/CD practices.
Top Skills:
Ci/Cd ToolingCloud InfrastructureContainersEnterprise-Authorized Ai CapabilitiesLinuxObservability (Slo/SliTelemetry)Windows
Big Data • Healthtech • HR Tech • Machine Learning • Software • Telehealth • Big Data Analytics
Own Garner’s cloud reliability strategy across AWS and Kubernetes, including SLOs, observability, incident response, infrastructure automation, cost optimization, and security compliance. Lead complex incident resolution, architect Terraform-based infrastructure, establish deployment and monitoring standards, mentor engineers, and use AI tools to automate operational work. Support high-scale AI/ML workloads while setting technical direction for platform reliability and production quality.
Top Skills:
AWSClaudeDatadogGitlabGoIstioKubernetesNatsPostgresPythonTerraformTypescript
Artificial Intelligence • Fintech • Information Technology • Logistics • Payments • Business Intelligence • Generative AI
Lead the design and roadmap for global Active Directory and identity infrastructure, implement Identity-as-Code and GitOps automation, own incident escalation and observability, define delegation/tiered administration, integrate applications with Okta and cloud identity, mentor teams, and publish identity architecture and security best practices.
Top Skills:
Active Directory Domain Services (Ad Ds)AnsibleAWSAws Directory ServiceAzureAzure Active Directory (Entra Id)Azure SentinelCertificate ServicesChefDhcpDnsGCPGitopsGroup Policy Objects (Gpo)New RelicOktaPowershellPowershell DscPythonTerraform
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Financial Services
Design, implement, and maintain reliable, scalable cloud-native platforms using infrastructure-as-code and CI/CD. Build observability, define SLOs/SLIs, troubleshoot incidents, reduce toil, and collaborate with engineering teams to improve availability and performance.
Top Skills:
AnsibleAWSDatadogDockerDynatraceEcsGitlabGrafanaJavaJenkinsKubernetesLinuxPrometheusPythonSplunkTerraformWindows
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Cloud • Computer Vision • Information Technology • Sales • Security • Cybersecurity
Senior SRE owning availability, automation, and observability for CI/CD platform services. Build and operate infrastructure, run on-call, lead incident response, mentor engineers, drive design/capacity planning, integrate AI-assisted workflows, and improve cross-team reliability.
Top Skills:
Active DirectoryAnsibleApache AirflowSparkAWSAzureBashBazelBitbucketCassandraChefDatadogDnsFirewall RulesGCPGitGithub ActionsGitlabGitlab CiGoGrafanaHoneycombHumio/LogscaleJenkinsKafkaKubernetesLoad BalancersMongoDBMySQLNasNew RelicNfsObject StorageOpensearchOraclePostgresPowershellPrometheusPulsarPuppetPythonRabbitMQRedis/ValkeyRedpandaRoutingSaltSanSplunkTerraformVarnishVipsWindows Server
Reposted 10 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will develop and support distributed storage services, ensuring reliability and operational safety, with a focus on automation and efficiency.
Top Skills:
AWSAzureDnsGoGoogle Cloud PlatformKubernetesLinuxPythonTcp/IpTls
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills:
Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
4 Days AgoSaved
Easy Apply
Easy Apply
Fintech • News + Entertainment • Software • Financial Services
Define tastytrade’s SRE practice, including customer-focused SLOs, error budgets, burn-rate alerts, observability standards, and production readiness reviews. Embed reliability patterns in Ruby, Java, and Elixir services running on HashiCorp Nomad. Extend Prometheus, Honeycomb, and OpenTelemetry observability; conduct fault-injection and tabletop exercises; strengthen on-call and incident-review processes; and mentor engineering teams in reliability practices.
Top Skills:
ConsulElixirGrafanaHashicorp NomadHoneycombJavaLinuxMulticastOpentelemetryPacket CapturePrometheusPythonRubyTcp/IpUdpVault
Reposted 4 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Information Technology • Professional Services • Consulting
Design, implement, and maintain scalable AWS infrastructure and CI/CD pipelines using Terraform and Harness. Build observability with Dynatrace/Datadog, troubleshoot production incidents, run blameless postmortems, optimize cost/performance, drive chaos engineering, embed SRE practices (SLAs/SLOs), and mentor junior engineers.
Top Skills:
AnsibleApp MeshAWSAws Well-Architected FrameworkCi/CdCloudfrontContainerizationDatadogDynatraceEc2EcsEksGoHarnessIamInfrastructure As CodeIstioKubernetesLambdaLinuxPythonRdsS3ShellTerraformVpc
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Digital Media • Information Technology • News + Entertainment
Responsible for ensuring reliability, scalability, and performance of data platforms. Design monitoring and alerting, automate deployments and recovery, optimize storage and query performance, troubleshoot incidents, plan capacity and scaling, document operations, enforce security/compliance, and collaborate with data engineering, product, and data science teams to maintain high availability of large-scale data systems.
Top Skills:
AnsibleAWSAzureCi/CdDockerElk StackGCPGoGrafanaJavaKubernetesMySQLNoSQLPostgresPrometheusPythonScalaTerraform
Reposted 20 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Reposted 22 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will lead security design and implementation for cloud infrastructures, mentor teams, and automate security solutions.
Top Skills:
AnsibleAWSAzureCloud Security ToolsCloudFormationGCPGoTerraform
Automotive
Leads SRE engineering leaders and engineers while defining enterprise observability, reliability, and platform strategy across GCP, on-premise, manufacturing, distribution, and campus environments. Oversees vendor-agnostic tooling, OpenTelemetry integrations, CI/CD observability, SRE maturity models, and Agentic AI initiatives. Drives adoption of SRE practices, develops technical roadmaps, partners with senior leadership and operational teams, and maintains hands-on architectural and technical credibility.
Top Skills:
Agentic AiAWSAzureCi/CdDatadogDynatraceGCPNew RelicOpentelemetryOtel Genai Semantic ConventionsSource Control PlatformsSplunkTerraform
Fintech
Lead SRE work partnering with development teams to design and implement availability, scalability, observability, and automation for production systems. Build tooling, manage incident response and RCAs, optimize capacity and performance, mentor engineers, maintain runbooks, and participate in a 24x7 on-call rotation.
Top Skills:
AuroraAWSChefCi/CdDockerDynamoDBGitGoIpJavaJavaScriptJenkinsJmsKafkaKubernetesLinuxMavenMemcachedMicroservicesObservabilityOraclePythonRedisRubySqsSwarmTcpUdp
Reposted YesterdaySaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Chicago, IL Companies Hiring Senior Site Reliability Engineers
See AllPopular Job Searches
All Software Engineer Jobs in Chicago
.NET Developer Jobs in Chicago
Android Developer Jobs in Chicago
Application Engineer Jobs in Chicago
Artificial Intelligence Engineer Jobs in Chicago
Backend Engineer Jobs in Chicago
C# Jobs in Chicago
C++ Jobs in Chicago
Devops Engineer Jobs in Chicago
DevOps Jobs in Chicago
Director Of Software Engineering Jobs in Chicago
Electrical Engineering Jobs in Chicago
Engineering Jobs in Chicago
Engineering Manager Jobs in Chicago
Enterprise Architect Jobs in Chicago
Fpga Engineer Jobs in Chicago
Front-End Developer Jobs in Chicago
Full-Stack Engineer Jobs in Chicago
Golang Jobs in Chicago
Hardware Engineer Jobs in Chicago
Infrastructure Engineer Jobs in Chicago
iOS Developer Jobs in Chicago
Java Developer Jobs in Chicago
Java Full-Stack Engineer Jobs in Chicago
Javascript Jobs in Chicago
Lead Software Engineer Jobs in Chicago
Linux Jobs in Chicago
Perl Jobs in Chicago
PHP Developer Jobs in Chicago
Platform Engineer Jobs in Chicago
Principal Engineer Jobs in Chicago
Principal Software Engineer Jobs in Chicago
Project Engineer Jobs in Chicago
Python Jobs in Chicago
QA Engineer Jobs in Chicago
Reliability Engineer Jobs in Chicago
Ruby Jobs in Chicago
Sales Engineer Jobs in Chicago
Salesforce Developer Jobs in Chicago
Scala Jobs in Chicago
Senior Android Engineer Jobs in Chicago
Senior Devops Engineer Jobs in Chicago
Senior Engineer Jobs in Chicago
Senior Front-End Engineer Jobs in Chicago
Senior Full-Stack Engineer Jobs in Chicago
Senior Java Engineer Jobs in Chicago
Senior Network Engineer Jobs in Chicago
Senior Platform Engineer Jobs in Chicago
Senior Site Reliability Engineer Jobs in Chicago
Senior Software Architect Jobs in Chicago
Senior Solutions Architect Jobs in Chicago
Senior Systems Engineer Jobs in Chicago
Software Engineering Manager Jobs in Chicago
Software Test Engineer Jobs in Chicago
Solutions Architect Jobs in Chicago
Solutions Engineer Jobs in Chicago
Staff Engineer Jobs in Chicago
Staff Software Engineer Jobs in Chicago
Systems Engineer Jobs in Chicago
Web Developer Jobs in Chicago
All Filters
Total selected ()
No Results
No Results







.png)























