Maximum of 25 job preferences reached.
Top SRE Engineer Jobs in Chicago, IL
Healthtech
Design, scale, and operate secure AWS cloud infrastructure (EKS, IAM, RBAC); build and maintain IaC (Terraform/Terragrunt), GitHub Actions CI/CD, Datadog observability, and Python automation; document runbooks, participate in on-call rotations, postmortems, and Agile workflows to improve reliability and security.
Top Skills:
AWSDatadogEc2EksFargateGithub ActionsGithub Advanced SecurityHelmIamJIRAKubernetesLambdaPythonRbacSecrets ManagerServerlessTerraformTerragruntVpc
Big Data • Analytics
Own production reliability for customer-facing radar and weather data services across Azure, colocation, and edge Kubernetes. Refactor C#/.NET services for multi-replica safety, design multi-cluster HA, operate self-managed Kubernetes, improve observability and automation, lead incident response and postmortems, and drive operational excellence and capacity planning.
Top Skills:
.NetAnsibleC#DatadogGpu-Enabled WorkloadsGrafanaHelmIstioKubernetesLokiLonghornAzureNatsOctopus DeployOpentelemetryPostgisPostgresPrometheusRabbitMQRancherRke2Terraform
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Founding Staff SRE for Volcano: define SLOs/error budgets, architect multi-region Kubernetes infrastructure, build GitOps/CI-CD with ArgoCD/Helm/Terraform, scale managed Postgres/Redis/object storage, implement observability with Datadog/Prometheus/Grafana, lead incident response and SRE culture, and mentor cross-functional teams.
Top Skills:
ArgocdCanary DeploymentsCi/CdCniDatadogGitopsGrafanaHelmIngressKubernetesObject StoragePostgresPrometheusRedisService MeshTerraformTerragrunt
Artificial Intelligence
Own operational excellence for cloud infrastructure: run incident management, improve reliability through automation, own a platform domain (e.g., Kubernetes, Temporal, observability), manage vendor and cost relationships, and deliver measurable reductions in incidents and costs within 12 months.
Top Skills:
AWSKubernetesLlm ApisMongoDBObservabilityPythonTemporal
Cloud • Security • Software • Cybersecurity
Senior SRE to build and run Veeam's Government/Sovereign-cloud reliability practice. Responsibilities include mapping platform workloads, writing runbooks, defining SLIs/SLOs, designing HA on Azure Government, incident response and postmortems, closing observability gaps, automation and IaC in compliance-restricted environments, CI/CD/GitOps pipelines, on-call rotations, and cross-team collaboration and mentoring.
Top Skills:
Api ManagementApplication InsightsArgocdArm TemplatesAWSAws CloudformationAws GovcloudAzureAzure DevopsAzure FunctionsAzure GovernmentAzure MonitorAzure StorageBitbucketC#Ci/CdCosmos DbDaggerElastic Stack (Elk)Entra IdFluxcdGitGithub ActionsGitlab CiGitopsGoGrafanaJavaJavaScriptKubernetesMicrosoft TfsOpentelemetryPrometheusPulumiServerless FrameworkTerraformTerragruntTypescript
3D Printing • Artificial Intelligence • Software • Design
Lead design and operation of scalable, multi-tenant spatial streaming platforms. Build Terraform-based cloud infrastructure, optimize CDN/content delivery, implement observability (SLI/SLO), run incident response/on-call, conduct post-mortems, enforce compliance and security practices, and mentor DevOps engineers to improve reliability and production readiness.
Top Skills:
Aws FargateCdnCoreweaveGrafanaKubernetesPrometheusTerraform
Software
As a Site Reliability Engineer, you'll enhance system reliability, collaborate on production readiness, define SLIs/SLOs, and improve incident response.
Top Skills:
AWSDatadogGrafanaKubernetesOpentelemetryPrometheusTypescript
Digital Media • Edtech
Drive reliability and observability of Epic's GCP-based platform. Own cloud infrastructure, container platform (Kubernetes/GKE), CI/CD, observability, and IaC (Terraform). Define SLOs/SLIs, reduce toil, manage security and compliance practices, participate in on-call rotations, lead incident response and post-mortems, and partner with product and data teams to troubleshoot and improve platform reliability.
Top Skills:
ArgocdBashCloud MonitoringDockerGceGCPGcsGithub ActionsHelmIamJenkinsKubernetes (Gke)New RelicPythonTerraformVpc
Information Technology • Cybersecurity
Lead global SRE team to design, implement, and operate scalable, highly available cloud infrastructure. Drive observability, automation (IaC/CI-CD), cost optimization (FinOps), security hygiene, and post-incident improvements while remaining hands-on and exploring AI tooling for reliability.
Top Skills:
Ai ToolingAlloyAWSAzureCi/CdCloudFormationContainer OrchestrationDatadogEdge ComputingFinopsGCPGrafanaGrafana IrnInfrastructure-As-CodeKubernetesLokiPrometheusPulumiServerlessSplunkSpot InstancesTerraform
Logistics • Software
Own and operate scalable infrastructure on GCP (GKE, Cloud Run, AlloyDB); author Terraform modules; manage containerized workloads; build observability in Datadog; design CI/CD in GitHub Actions; automate operational workflows; lead incident response and post-mortems; partner with engineers to improve reliability, cost, and automation.
Top Skills:
AlloydbClickhouseCloud RunCloudflare WorkersDatadogDockerGCPGithub ActionsGkeGoGrafanaIamKafkaKubernetesNetworkingPostgresPrometheusPub/SubPythonRedisRedpandaTerraformTypescript
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills:
AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Develop and maintain code and automation for highly scalable M365 sovereign cloud services. Operate live sites, participate in on-call rotations, troubleshoot incidents, improve observability, design automation for deployments, and collaborate with product and security stakeholders to ensure reliability and compliance.
Top Skills:
AzureCC#C++CopilotExchange Online ProtectionExchange TransportGenerative AiJavaJavaScriptMicrosoft 365Microsoft Defender For OfficeOffice 365OnedrivePurviewPythonSharepointTeams
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead qualification, performance validation, and production readiness for new Azure Storage hardware and firmware. Drive test planning, automation frameworks, large-scale telemetry and benchmark analysis, root-cause investigations across software/hardware/firmware, and partner with engineering and vendors to resolve reliability and performance issues.
Top Skills:
Automation FrameworksAzureAzure StorageBenchmarkingFirmwareNetworkingSsdsTelemetry
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Design, develop, deploy, manage, and monitor Azure infrastructure and services. Serve as on-call DRI, analyze metrics, automate production deployments, ensure security/compliance, collaborate with partner teams, respond to incidents, and run postmortems. Support physical infrastructure including GPUs and InfiniBand.
Top Skills:
AzureGpusInfiniband
Insurance
Lead reliability and observability for the financial data platform: define SLOs/SLIs, build metrics pipelines, extend instrumentation across Velocity, Redpanda, MuleSoft, Snowflake, Fabric, and AWS; implement incident management (Datadog -> Incident.io -> ServiceNow), scale automation and remediation, and design AI-assisted SRE agents using Cursor for triage and root-cause analysis.
Top Skills:
Ai Coding AssistantsAWSCursorD365DatadogFabricGrafanaIncident.IoKafkaLlmsMulesoftOpentelemetryPower AppsPrometheusRedpandaServicenowSnowflakeVelocity
Information Technology • Security
Lead technical strategy and architecture for SimSpace's infrastructure, evolving CI/CD and multi-cluster Kubernetes platforms using Jsonnet and Grafana Tanka. Define SLIs/SLOs, build observability with the Grafana stack, embed security and compliance into pipelines, enable self-service developer tooling, command major incidents, and mentor engineering teams to improve reliability and scalability across cloud, on-prem, VMware, and air-gapped deployments.
Top Skills:
ArgocdCi/CdGithub ActionsGitopsGoGrafanaGrafana TankaJsonnetKubernetesKustomizePythonVMware
Software
Join as the company's first SRE to design reliability processes, build observability/CI/CD/IaC tooling, define SLOs, embed with product teams, run incident response, and support compliance for patient data.
Top Skills:
Access ControlsAlerting)Audit TrailsCi/CdGoHipaaInfrastructure As CodeLoggingObservability (MetricsPythonSlosSoc 2TracingTypescript
Healthtech • Telehealth
Lead the creation of SRE practices across six product teams: define SLOs/SLIs, implement observability, reduce outages, automate toil with IaC and tooling, run production readiness reviews, and train teams to improve incident detection and reliability.
Top Skills:
AWSAws EksDatadogGrafanaKubernetesNode.jsPrometheusPythonRdsReactTerraformTypescript
Fintech • Marketing Tech • Financial Services
Support and improve deployment pipelines, troubleshoot Kubernetes/Docker/Linux environments, maintain AWS infrastructure with Terraform, manage observability (Grafana/Prometheus/Elasticsearch) and secrets (Vault), build automation tooling, participate in on-call incident response, and collaborate with engineering teams to improve reliability and developer experience.
Top Skills:
AtlantisAws Ec2DockerElasticsearchGitlab CiGoGrafanaHashicorp VaultIamKafkaKubernetesLinuxLogstashPrometheusRdsS3TeamcityTerraform
Cloud
The Site Reliability Engineer will manage Kubernetes platforms, optimize AWS cloud infrastructure, ensure high availability, and automate deployment while handling troubleshooting and security compliance.
Top Skills:
AWSBashCi/CdCloudwatchElk StackGoGrafanaHelmIstioKubernetesPrometheusPythonTerraform
Cloud
The Senior Site Reliability Engineer will enhance the Splunk ecosystem and develop an Observability Platform by automating infrastructure and managing complex distributed systems, while optimizing log collection and incident response.
Top Skills:
AWSGCPGoKubernetesLinuxOpentelemetryPythonRubySplunkTerraform
Other • Social Impact
The Senior Site Reliability Engineer is responsible for maintaining Wikimedia's infrastructure, improving reliability, automating processes, and collaborating with teams. The role involves troubleshooting, managing deployments, and leading incident responses while working remotely.
Top Skills:
AnsibleBashCassandraDebianGoGrafanaHhvmKubernetesMariadbMemcachedPHPPrometheusPuppetPythonRedisRubyShell
Information Technology • Security • Cybersecurity
Design, build, and scale Kubernetes-based, multi-tenant infrastructure and CI/CD systems. Own AI tooling infrastructure (MCP servers) and secure AI access patterns. Optimize CI/CD, streaming analytics (Kafka, Flink, ClickHouse), observability, and incident response. Implement IaC (Terraform, Helm, Pulumi), GitOps (Argo CD), progressive delivery, automated testing, and mentor engineering teams.
Top Skills:
Ai AgentsAi/Llm ToolingAksArgo CdBashClickhouseDatadogEksFlinkGithub ActionsGitlab CiGitopsGkeGoGrafanaHelmJenkinsKafkaKubernetesLangfuseLangsmithMcp ServersMlopsOpentelemetryPrometheusPulumiPythonTerraform
Real Estate • Financial Services • PropTech
Support and optimize products migrated to AWS, implement cloud best practices, maintain operational coverage, enhance automation, observability, CI/CD/GitOps, and security. Collaborate with development and platform teams to scale, troubleshoot, and ensure reliable SaaS operations.
Top Skills:
AmisArgocdAWSAws Elastic BeanstalkAws Transfer FamilyAzure DevopsBashCloudwatchCurlDockerEc2EksFluxcdGitGitopsHTTPIstioKubernetesLinkerdLoad BalancerPowershellPythonRdsSQLTerraformWget
Artificial Intelligence • Software • Cybersecurity
Design, build, and maintain scalable, highly available cloud infrastructure for an AI-native cybersecurity platform. Automate deployments and incident response, optimize performance for AI workloads, manage IaC across cloud environments, lead incident management and post-mortems, and collaborate with engineering and security teams to embed reliability.
Top Skills:
AWSAzureDatadogDistributed DatabasesEksElkEvent-Driven SystemsGCPGkeGrafanaKubernetesMicroservicesNetworkingPrometheusPulumiStorageTerraform
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Chicago, IL Engineering Job Searches
Engineer Jobs in Chicago, IL
.NET Developer Jobs in Chicago, IL
Android Developer Jobs in Chicago, IL
Application Engineer Jobs in Chicago, IL
Automation Engineer Jobs in Chicago, IL
Backend Engineer Jobs in Chicago, IL
C# Jobs in Chicago, IL
C++ Jobs in Chicago, IL
Cloud Engineer Jobs in Chicago, IL
Controls Engineer Jobs in Chicago, IL
CTO Jobs in Chicago, IL
Design Engineer Jobs in Chicago, IL
DevOps Engineer Jobs in Chicago, IL
DevOps Jobs in Chicago, IL
Director of Engineering Jobs in Chicago, IL
Director of Software Engineering Jobs in Chicago, IL
Electrical Engineer Jobs in Chicago, IL
Embedded Software Engineer Jobs in Chicago, IL
Engineering Manager Jobs in Chicago, IL
Enterprise Architect Jobs in Chicago, IL
FPGA Engineer Jobs in Chicago, IL
Front End Developer Jobs in Chicago, IL
Full-Stack Engineer Jobs in Chicago, IL
Golang Jobs in Chicago, IL
Hardware Engineer Jobs in Chicago, IL
Infrastructure Engineer Jobs in Chicago, IL
iOS Developer Jobs in Chicago, IL
Java Developer Jobs in Chicago, IL
Java Full-Stack Engineer Jobs in Chicago, IL
Javascript Jobs in Chicago, IL
Lead Software Engineer Jobs in Chicago, IL
Linux Jobs in Chicago, IL
Manufacturing Engineer Jobs in Chicago, IL
Mechanical Design Engineer Jobs in Chicago, IL
Mechanical Engineer Jobs in Chicago, IL
Network Engineer Jobs in Chicago, IL
PHP Developer Jobs in Chicago, IL
Platform Engineer Jobs in Chicago, IL
Principal Engineer Jobs in Chicago, IL
Principal Software Engineer Jobs in Chicago, IL
Process Engineer Jobs in Chicago, IL
Project Engineer Jobs in Chicago, IL
Python Jobs in Chicago, IL
QA Engineer Jobs in Chicago, IL
QA Jobs in Chicago, IL
Reliability Engineer Jobs in Chicago, IL
Robotics Engineer Jobs in Chicago, IL
Ruby Jobs in Chicago, IL
Salesforce Developer Jobs in Chicago, IL
Scala Jobs in Chicago, IL
Security Engineer Jobs in Chicago, IL
Software Engineer Jobs in Chicago, IL
Software Engineering Manager Jobs in Chicago, IL
Software Test Engineer Jobs in Chicago, IL
Solutions Architect Jobs in Chicago, IL
Solutions Engineer Jobs in Chicago, IL
SRE Engineer Jobs in Chicago, IL
Staff Engineer Jobs in Chicago, IL
Staff Software Engineer Jobs in Chicago, IL
Systems Engineer Jobs in Chicago, IL
All Filters
Total selected ()
No Results
No Results






























