Top SRE Engineer Jobs in Chicago, IL

Reposted 4 Days AgoSaved
Remote
Chicago, IL
125K-165K Annually
Senior level
125K-165K Annually
Senior level
Healthtech
Design, scale, and operate secure AWS cloud infrastructure (EKS, IAM, RBAC); build and maintain IaC (Terraform/Terragrunt), GitHub Actions CI/CD, Datadog observability, and Python automation; document runbooks, participate in on-call rotations, postmortems, and Agile workflows to improve reliability and security.
Top Skills: AWSDatadogEc2EksFargateGithub ActionsGithub Advanced SecurityHelmIamJIRAKubernetesLambdaPythonRbacSecrets ManagerServerlessTerraformTerragruntVpc
Reposted 4 Days AgoSaved
Remote
Chicago, IL
135K-170K Annually
Senior level
135K-170K Annually
Senior level
Big Data • Analytics
Own production reliability for customer-facing radar and weather data services across Azure, colocation, and edge Kubernetes. Refactor C#/.NET services for multi-replica safety, design multi-cluster HA, operate self-managed Kubernetes, improve observability and automation, lead incident response and postmortems, and drive operational excellence and capacity planning.
Top Skills: .NetAnsibleC#DatadogGpu-Enabled WorkloadsGrafanaHelmIstioKubernetesLokiLonghornAzureNatsOctopus DeployOpentelemetryPostgisPostgresPrometheusRabbitMQRancherRke2Terraform
Reposted 4 Days AgoSaved
Remote
Chicago, IL
150K-210K Annually
Senior level
150K-210K Annually
Senior level
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Founding Staff SRE for Volcano: define SLOs/error budgets, architect multi-region Kubernetes infrastructure, build GitOps/CI-CD with ArgoCD/Helm/Terraform, scale managed Postgres/Redis/object storage, implement observability with Datadog/Prometheus/Grafana, lead incident response and SRE culture, and mentor cross-functional teams.
Top Skills: ArgocdCanary DeploymentsCi/CdCniDatadogGitopsGrafanaHelmIngressKubernetesObject StoragePostgresPrometheusRedisService MeshTerraformTerragrunt
Reposted 4 Days AgoSaved
Remote
Chicago, IL
Senior level
Senior level
Artificial Intelligence
Own operational excellence for cloud infrastructure: run incident management, improve reliability through automation, own a platform domain (e.g., Kubernetes, Temporal, observability), manage vendor and cost relationships, and deliver measurable reductions in incidents and costs within 12 months.
Top Skills: AWSKubernetesLlm ApisMongoDBObservabilityPythonTemporal
5 Days AgoSaved
Remote
Chicago, IL
173K-321K Annually
Senior level
173K-321K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Senior SRE to build and run Veeam's Government/Sovereign-cloud reliability practice. Responsibilities include mapping platform workloads, writing runbooks, defining SLIs/SLOs, designing HA on Azure Government, incident response and postmortems, closing observability gaps, automation and IaC in compliance-restricted environments, CI/CD/GitOps pipelines, on-call rotations, and cross-team collaboration and mentoring.
Top Skills: Api ManagementApplication InsightsArgocdArm TemplatesAWSAws CloudformationAws GovcloudAzureAzure DevopsAzure FunctionsAzure GovernmentAzure MonitorAzure StorageBitbucketC#Ci/CdCosmos DbDaggerElastic Stack (Elk)Entra IdFluxcdGitGithub ActionsGitlab CiGitopsGoGrafanaJavaJavaScriptKubernetesMicrosoft TfsOpentelemetryPrometheusPulumiServerless FrameworkTerraformTerragruntTypescript
5 Days AgoSaved
Remote
Chicago, IL
103K-287K Annually
Senior level
103K-287K Annually
Senior level
3D Printing • Artificial Intelligence • Software • Design
Lead design and operation of scalable, multi-tenant spatial streaming platforms. Build Terraform-based cloud infrastructure, optimize CDN/content delivery, implement observability (SLI/SLO), run incident response/on-call, conduct post-mortems, enforce compliance and security practices, and mentor DevOps engineers to improve reliability and production readiness.
Top Skills: Aws FargateCdnCoreweaveGrafanaKubernetesPrometheusTerraform
Reposted 5 Days AgoSaved
Remote
Chicago, IL
175K-275K Annually
Mid level
175K-275K Annually
Mid level
Software
As a Site Reliability Engineer, you'll enhance system reliability, collaborate on production readiness, define SLIs/SLOs, and improve incident response.
Top Skills: AWSDatadogGrafanaKubernetesOpentelemetryPrometheusTypescript
6 Days AgoSaved
Remote
Chicago, IL
160K-200K Annually
Senior level
160K-200K Annually
Senior level
Digital Media • Edtech
Drive reliability and observability of Epic's GCP-based platform. Own cloud infrastructure, container platform (Kubernetes/GKE), CI/CD, observability, and IaC (Terraform). Define SLOs/SLIs, reduce toil, manage security and compliance practices, participate in on-call rotations, lead incident response and post-mortems, and partner with product and data teams to troubleshoot and improve platform reliability.
Top Skills: ArgocdBashCloud MonitoringDockerGceGCPGcsGithub ActionsHelmIamJenkinsKubernetes (Gke)New RelicPythonTerraformVpc
6 Days AgoSaved
Remote
Chicago, IL
190K-220K Annually
Expert/Leader
190K-220K Annually
Expert/Leader
Information Technology • Cybersecurity
Lead global SRE team to design, implement, and operate scalable, highly available cloud infrastructure. Drive observability, automation (IaC/CI-CD), cost optimization (FinOps), security hygiene, and post-incident improvements while remaining hands-on and exploring AI tooling for reliability.
Top Skills: Ai ToolingAlloyAWSAzureCi/CdCloudFormationContainer OrchestrationDatadogEdge ComputingFinopsGCPGrafanaGrafana IrnInfrastructure-As-CodeKubernetesLokiPrometheusPulumiServerlessSplunkSpot InstancesTerraform
6 Days AgoSaved
Remote
Chicago, IL
Senior level
Senior level
Logistics • Software
Own and operate scalable infrastructure on GCP (GKE, Cloud Run, AlloyDB); author Terraform modules; manage containerized workloads; build observability in Datadog; design CI/CD in GitHub Actions; automate operational workflows; lead incident response and post-mortems; partner with engineers to improve reliability, cost, and automation.
Top Skills: AlloydbClickhouseCloud RunCloudflare WorkersDatadogDockerGCPGithub ActionsGkeGoGrafanaIamKafkaKubernetesNetworkingPostgresPrometheusPub/SubPythonRedisRedpandaTerraformTypescript
6 Days AgoSaved
In-Office or Remote
Chicago, IL
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills: AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
6 Days AgoSaved
In-Office or Remote
Chicago, IL
102K-219K Annually
Junior
102K-219K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Develop and maintain code and automation for highly scalable M365 sovereign cloud services. Operate live sites, participate in on-call rotations, troubleshoot incidents, improve observability, design automation for deployments, and collaborate with product and security stakeholders to ensure reliability and compliance.
Top Skills: AzureCC#C++CopilotExchange Online ProtectionExchange TransportGenerative AiJavaJavaScriptMicrosoft 365Microsoft Defender For OfficeOffice 365OnedrivePurviewPythonSharepointTeams
New

Cut your apply time in half.

Use ourAI Assistantto automatically fill your job applications.

Use For Free
Application Tracker Preview
6 Days AgoSaved
In-Office or Remote
Chicago, IL
120K-261K Annually
Senior level
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead qualification, performance validation, and production readiness for new Azure Storage hardware and firmware. Drive test planning, automation frameworks, large-scale telemetry and benchmark analysis, root-cause investigations across software/hardware/firmware, and partner with engineering and vendors to resolve reliability and performance issues.
Top Skills: Automation FrameworksAzureAzure StorageBenchmarkingFirmwareNetworkingSsdsTelemetry
6 Days AgoSaved
Remote
Chicago, IL
102K-219K Annually
Junior
102K-219K Annually
Junior
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, develop, deploy, manage, and monitor Azure infrastructure and services. Serve as on-call DRI, analyze metrics, automate production deployments, ensure security/compliance, collaborate with partner teams, respond to incidents, and run postmortems. Support physical infrastructure including GPUs and InfiniBand.
Top Skills: AzureGpusInfiniband
7 Days AgoSaved
Remote
Chicago, IL
Senior level
Senior level
Insurance
Lead reliability and observability for the financial data platform: define SLOs/SLIs, build metrics pipelines, extend instrumentation across Velocity, Redpanda, MuleSoft, Snowflake, Fabric, and AWS; implement incident management (Datadog -> Incident.io -> ServiceNow), scale automation and remediation, and design AI-assisted SRE agents using Cursor for triage and root-cause analysis.
Top Skills: Ai Coding AssistantsAWSCursorD365DatadogFabricGrafanaIncident.IoKafkaLlmsMulesoftOpentelemetryPower AppsPrometheusRedpandaServicenowSnowflakeVelocity
7 Days AgoSaved
Remote
Chicago, IL
165K-230K Annually
Senior level
165K-230K Annually
Senior level
Information Technology • Security
Lead technical strategy and architecture for SimSpace's infrastructure, evolving CI/CD and multi-cluster Kubernetes platforms using Jsonnet and Grafana Tanka. Define SLIs/SLOs, build observability with the Grafana stack, embed security and compliance into pipelines, enable self-service developer tooling, command major incidents, and mentor engineering teams to improve reliability and scalability across cloud, on-prem, VMware, and air-gapped deployments.
Top Skills: ArgocdCi/CdGithub ActionsGitopsGoGrafanaGrafana TankaJsonnetKubernetesKustomizePythonVMware
8 Days AgoSaved
Remote
Chicago, IL
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Software
Join as the company's first SRE to design reliability processes, build observability/CI/CD/IaC tooling, define SLOs, embed with product teams, run incident response, and support compliance for patient data.
Top Skills: Access ControlsAlerting)Audit TrailsCi/CdGoHipaaInfrastructure As CodeLoggingObservability (MetricsPythonSlosSoc 2TracingTypescript
8 Days AgoSaved
Remote
Chicago, IL
160K-185K Annually
Senior level
160K-185K Annually
Senior level
Healthtech • Telehealth
Lead the creation of SRE practices across six product teams: define SLOs/SLIs, implement observability, reduce outages, automate toil with IaC and tooling, run production readiness reviews, and train teams to improve incident detection and reliability.
Top Skills: AWSAws EksDatadogGrafanaKubernetesNode.jsPrometheusPythonRdsReactTerraformTypescript
Reposted 17 Days AgoSaved
In-Office
Chicago, IL
145K-175K Annually
Senior level
145K-175K Annually
Senior level
Fintech • Marketing Tech • Financial Services
Support and improve deployment pipelines, troubleshoot Kubernetes/Docker/Linux environments, maintain AWS infrastructure with Terraform, manage observability (Grafana/Prometheus/Elasticsearch) and secrets (Vault), build automation tooling, participate in on-call incident response, and collaborate with engineering teams to improve reliability and developer experience.
Top Skills: AtlantisAws Ec2DockerElasticsearchGitlab CiGoGrafanaHashicorp VaultIamKafkaKubernetesLinuxLogstashPrometheusRdsS3TeamcityTerraform
Reposted 17 Days AgoSaved
In-Office
Chicago, IL
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
The Site Reliability Engineer will manage Kubernetes platforms, optimize AWS cloud infrastructure, ensure high availability, and automate deployment while handling troubleshooting and security compliance.
Top Skills: AWSBashCi/CdCloudwatchElk StackGoGrafanaHelmIstioKubernetesPrometheusPythonTerraform
Reposted 17 Days AgoSaved
In-Office
Chicago, IL
194K-267K Annually
Senior level
194K-267K Annually
Senior level
Cloud
The Senior Site Reliability Engineer will enhance the Splunk ecosystem and develop an Observability Platform by automating infrastructure and managing complex distributed systems, while optimizing log collection and incident response.
Top Skills: AWSGCPGoKubernetesLinuxOpentelemetryPythonRubySplunkTerraform
Reposted 8 Days AgoSaved
Remote
Chicago, IL
113K-176K Annually
Senior level
113K-176K Annually
Senior level
Other • Social Impact
The Senior Site Reliability Engineer is responsible for maintaining Wikimedia's infrastructure, improving reliability, automating processes, and collaborating with teams. The role involves troubleshooting, managing deployments, and leading incident responses while working remotely.
Top Skills: AnsibleBashCassandraDebianGoGrafanaHhvmKubernetesMariadbMemcachedPHPPrometheusPuppetPythonRedisRubyShell
8 Days AgoSaved
Remote
Chicago, IL
152K-195K Annually
Senior level
152K-195K Annually
Senior level
Information Technology • Security • Cybersecurity
Design, build, and scale Kubernetes-based, multi-tenant infrastructure and CI/CD systems. Own AI tooling infrastructure (MCP servers) and secure AI access patterns. Optimize CI/CD, streaming analytics (Kafka, Flink, ClickHouse), observability, and incident response. Implement IaC (Terraform, Helm, Pulumi), GitOps (Argo CD), progressive delivery, automated testing, and mentor engineering teams.
Top Skills: Ai AgentsAi/Llm ToolingAksArgo CdBashClickhouseDatadogEksFlinkGithub ActionsGitlab CiGitopsGkeGoGrafanaHelmJenkinsKafkaKubernetesLangfuseLangsmithMcp ServersMlopsOpentelemetryPrometheusPulumiPythonTerraform
Reposted 8 Days AgoSaved
Remote
Chicago, IL
110K-140K Annually
Senior level
110K-140K Annually
Senior level
Real Estate • Financial Services • PropTech
Support and optimize products migrated to AWS, implement cloud best practices, maintain operational coverage, enhance automation, observability, CI/CD/GitOps, and security. Collaborate with development and platform teams to scale, troubleshoot, and ensure reliable SaaS operations.
Top Skills: AmisArgocdAWSAws Elastic BeanstalkAws Transfer FamilyAzure DevopsBashCloudwatchCurlDockerEc2EksFluxcdGitGitopsHTTPIstioKubernetesLinkerdLoad BalancerPowershellPythonRdsSQLTerraformWget
9 Days AgoSaved
Remote
Chicago, IL
Expert/Leader
Expert/Leader
Artificial Intelligence • Software • Cybersecurity
Design, build, and maintain scalable, highly available cloud infrastructure for an AI-native cybersecurity platform. Automate deployments and incident response, optimize performance for AI workloads, manage IaC across cloud environments, lead incident management and post-mortems, and collaborate with engineering and security teams to embed reliability.
Top Skills: AWSAzureDatadogDistributed DatabasesEksElkEvent-Driven SystemsGCPGkeGrafanaKubernetesMicroservicesNetworkingPrometheusPulumiStorageTerraform
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account