Top Remote Reliability Engineer Jobs in Chicago, IL

Reposted 2 Days AgoSaved
Remote
United States
190K-220K Annually
Expert/Leader
190K-220K Annually
Expert/Leader
Information Technology • Cybersecurity
Lead global SRE team to design, implement, and operate scalable, highly available cloud infrastructure. Drive observability, automation (IaC/CI-CD), cost optimization (FinOps), security hygiene, and post-incident improvements while remaining hands-on and exploring AI tooling for reliability.
Top Skills: Ai ToolingAlloyAWSAzureCi/CdCloudFormationContainer OrchestrationDatadogEdge ComputingFinopsGCPGrafanaGrafana IrnInfrastructure-As-CodeKubernetesLokiPrometheusPulumiServerlessSplunkSpot InstancesTerraform
Reposted 2 Days AgoSaved
Remote
United States
Senior level
Senior level
Logistics • Software
Own and operate scalable infrastructure on GCP (GKE, Cloud Run, AlloyDB); author Terraform modules; manage containerized workloads; build observability in Datadog; design CI/CD in GitHub Actions; automate operational workflows; lead incident response and post-mortems; partner with engineers to improve reliability, cost, and automation.
Top Skills: AlloydbClickhouseCloud RunCloudflare WorkersDatadogDockerGCPGithub ActionsGkeGoGrafanaIamKafkaKubernetesNetworkingPostgresPrometheusPub/SubPythonRedisRedpandaTerraformTypescript
Reposted 2 Days AgoSaved
In-Office or Remote
United States
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Ensure reliability, scalability, and usability of network infrastructure for Akamai Connected Cloud. Define requirements and SLOs, build automation and CI/CD pipelines, collaborate with dev/QA to improve code and stability, troubleshoot complex network issues (on-call), and mentor teammates while driving architectural standards.
Top Skills: AnsibleArgocdBashBirdChefFrrGithub ActionsGoGobgpJenkinsLinux NetworkingPuppetPythonSalt Stack
Reposted 2 Days AgoSaved
Remote or Hybrid
US
138K-221K Annually
Senior level
138K-221K Annually
Senior level
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills: .NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
3 Days AgoSaved
Remote
United States
145K-200K Annually
Senior level
145K-200K Annually
Senior level
Software
Lead SRE to define strategy and roadmap for reliability, scalability, observability, and automation across cloud and hybrid environments. Design and operate containerized production workloads, infrastructure-as-code, monitoring/alerting, incident management, and compliance for regulated domains. Mentor SREs, partner with security and product teams, manage cloud costs and capacity, and build a developer platform to improve delivery and production stability.
Top Skills: AWSAws MarketplaceAzureAzure MarketplaceGCPGoogle Cloud MarketplaceGrafanaKubernetesPrometheusTerraform
Reposted 3 Days AgoSaved
Remote
United States
147K-168K Annually
Senior level
147K-168K Annually
Senior level
Legal Tech • Software
Lead observability and incident management efforts: define SLIs/SLOs, build monitoring/alerting, dashboards, logging, and tracing. Drive incident response, postmortems, and reliability improvements to reduce MTTD/MTTR. Integrate observability into CI/CD, maintain AWS and Kubernetes infrastructure, automate operations, and mentor engineers on SRE best practices.
Top Skills: AWSBashCi/CdDatadogDistributed TracingDynatraceGrafanaKubernetesNew RelicOpentelemetryPowershellPrometheusPython
Reposted 3 Days AgoSaved
Remote
USA
235K-275K Annually
Expert/Leader
235K-275K Annually
Expert/Leader
Legal Tech • Software
Senior technical leader for SRE driving observability, platform infrastructure, SLIs/SLOs, incident response, automation, and self-service platform capabilities. Shapes reliability strategy, mentors engineers, and ensures production-scale operational excellence.
Top Skills: AiopsBashDatadogGoInfrastructure As CodeKubernetesNew RelicObservabilityPython
Reposted 3 Days AgoSaved
Remote
United States
175K-195K Annually
Senior level
175K-195K Annually
Senior level
Legal Tech • Software
Own and improve platform reliability, availability, and performance for Filevines systems. Build AWS infrastructure with Terraform, develop CI/CD pipelines, monitoring, and automation, define SLOs/SLIs/error budgets, lead incident response, mentor engineers, and participate in a 24/7 on-call rotation to support highly available, low-latency production systems.
Top Skills: AutomationAWSBashCi/CdCloudwatchEc2EcsEksIamLambdaObservabilityPowershellPythonRoute 53S3TerraformVpc
Reposted 3 Days AgoSaved
Remote
United States
Senior level
Senior level
Insurance
Lead reliability and observability for the financial data platform: define SLOs/SLIs, build metrics pipelines, extend instrumentation across Velocity, Redpanda, MuleSoft, Snowflake, Fabric, and AWS; implement incident management (Datadog -> Incident.io -> ServiceNow), scale automation and remediation, and design AI-assisted SRE agents using Cursor for triage and root-cause analysis.
Top Skills: Ai Coding AssistantsAWSCursorD365DatadogFabricGrafanaIncident.IoKafkaLlmsMulesoftOpentelemetryPower AppsPrometheusRedpandaServicenowSnowflakeVelocity
Reposted 13 Days AgoSaved
In-Office or Remote
5 Locations
Expert/Leader
Expert/Leader
Agency • Information Technology • Professional Services • Software
Lead development and implementation of preventive and predictive maintenance for onshore mechanical equipment, use CMMS to plan and monitor maintenance, analyze reliability data, perform RCA, support operations and maintenance teams, ensure safety and compliance, and recommend improvements to reduce downtime and costs.
Top Skills: CmmsPredictive MaintenancePreventive MaintenanceRoot Cause Analysis
Reposted 13 Days AgoSaved
In-Office or Remote
5 Locations
Expert/Leader
Expert/Leader
Agency • Information Technology • Professional Services • Software
Lead development and implementation of preventive and predictive maintenance programs for offshore mechanical equipment, use CMMS to plan and track work, perform RCA for failures, support offshore teams in troubleshooting, monitor equipment reliability, and ensure compliance with safety and maintenance standards.
Top Skills: CmmsPredictive MaintenancePreventive MaintenanceRoot Cause Analysis
Reposted 4 Days AgoSaved
Remote
USA
160K-185K Annually
Senior level
160K-185K Annually
Senior level
Healthtech • Telehealth
Lead the creation of SRE practices across six product teams: define SLOs/SLIs, implement observability, reduce outages, automate toil with IaC and tooling, run production readiness reviews, and train teams to improve incident detection and reliability.
Top Skills: AWSAws EksDatadogGrafanaKubernetesNode.jsPrometheusPythonRdsReactTerraformTypescript
New

Track Smarter, Apply Better.

Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.

Use For Free
Application Tracker Preview
Reposted 4 Days AgoSaved
Remote
USA
150K-200K Annually
Senior level
150K-200K Annually
Senior level
Software
Join as the company's first SRE to design reliability processes, build observability/CI/CD/IaC tooling, define SLOs, embed with product teams, run incident response, and support compliance for patient data.
Top Skills: Access ControlsAlerting)Audit TrailsCi/CdGoHipaaInfrastructure As CodeLoggingObservability (MetricsPythonSlosSoc 2TracingTypescript
Reposted 4 Days AgoSaved
Remote
USA
113K-176K Annually
Senior level
113K-176K Annually
Senior level
Other • Social Impact
The Senior Site Reliability Engineer is responsible for maintaining Wikimedia's infrastructure, improving reliability, automating processes, and collaborating with teams. The role involves troubleshooting, managing deployments, and leading incident responses while working remotely.
Top Skills: AnsibleBashCassandraDebianGoGrafanaHhvmKubernetesMariadbMemcachedPHPPrometheusPuppetPythonRedisRubyShell
Reposted 5 Days AgoSaved
In-Office or Remote
United States
120K-261K Annually
Senior level
120K-261K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Lead qualification, performance validation, and production readiness for new Azure Storage hardware and firmware. Drive test planning, automation frameworks, large-scale telemetry and benchmark analysis, root-cause investigations across software/hardware/firmware, and partner with engineering and vendors to resolve reliability and performance issues.
Top Skills: Automation FrameworksAzureAzure StorageBenchmarkingFirmwareNetworkingSsdsTelemetry
Reposted 5 Days AgoSaved
In-Office or Remote
United States
146K-264K Annually
Senior level
146K-264K Annually
Senior level
Cloud • Security • Software • Cybersecurity
Lead and mentor SRE teams; partner with engineering, operations and product; apply statistical analysis and networking expertise to diagnose performance and reliability issues; define and implement data feeds; influence technical decisions and investments; build tooling to automate analytical workflows and increase platform reliability.
Top Skills: CCloudDistributed SystemsDnsEdgeHTTPJavaPerlPythonRSQLTcpTls
Reposted 5 Days AgoSaved
Remote
USA
Expert/Leader
Expert/Leader
Artificial Intelligence • Software • Cybersecurity
Design, build, and maintain scalable, highly available cloud infrastructure for an AI-native cybersecurity platform. Automate deployments and incident response, optimize performance for AI workloads, manage IaC across cloud environments, lead incident management and post-mortems, and collaborate with engineering and security teams to embed reliability.
Top Skills: AWSAzureDatadogDistributed DatabasesEksElkEvent-Driven SystemsGCPGkeGrafanaKubernetesMicroservicesNetworkingPrometheusPulumiStorageTerraform
Reposted 5 Days AgoSaved
Remote
USA
110K-140K Annually
Senior level
110K-140K Annually
Senior level
Real Estate • Financial Services • PropTech
Support and optimize products migrated to AWS, implement cloud best practices, maintain operational coverage, enhance automation, observability, CI/CD/GitOps, and security. Collaborate with development and platform teams to scale, troubleshoot, and ensure reliable SaaS operations.
Top Skills: AmisArgocdAWSAws Elastic BeanstalkAws Transfer FamilyAzure DevopsBashCloudwatchCurlDockerEc2EksFluxcdGitGitopsHTTPIstioKubernetesLinkerdLoad BalancerPowershellPythonRdsSQLTerraformWget
Reposted 5 Days AgoSaved
Remote
United States
152K-195K Annually
Senior level
152K-195K Annually
Senior level
Information Technology • Security • Cybersecurity
Design, build, and scale Kubernetes-based, multi-tenant infrastructure and CI/CD systems. Own AI tooling infrastructure (MCP servers) and secure AI access patterns. Optimize CI/CD, streaming analytics (Kafka, Flink, ClickHouse), observability, and incident response. Implement IaC (Terraform, Helm, Pulumi), GitOps (Argo CD), progressive delivery, automated testing, and mentor engineering teams.
Top Skills: Ai AgentsAi/Llm ToolingAksArgo CdBashClickhouseDatadogEksFlinkGithub ActionsGitlab CiGitopsGkeGoGrafanaHelmJenkinsKafkaKubernetesLangfuseLangsmithMcp ServersMlopsOpentelemetryPrometheusPulumiPythonTerraform
Mid level
Artificial Intelligence • Hardware • Software • Semiconductor
Operate and scale production AI inference infrastructure, run releases and capacity changes, build self-service CD pipelines and automation, extend telemetry and observability, collaborate on SLOs, post-mortems, and capacity planning to reduce operational toil.
Top Skills: Argo CdBazelFluxGitopsGoGrafanaInfluxdbKubernetesPrometheusPython
Reposted 6 Days AgoSaved
Remote
US
110K-130K Annually
Senior level
110K-130K Annually
Senior level
Software
Support and improve production SaaS infrastructure across AWS, colocation, and hosted platforms. Administer Windows and Linux systems, virtualization, storage, networking, and database support. Lead incident troubleshooting, root cause analysis, monitoring improvements, vulnerability remediation, automation initiatives, and medium-sized infrastructure projects. Collaborate on compliance (SOX/PCI/HIPAA), disaster recovery, and documentation to increase operational reliability.
Top Skills: AWSBackup And RecoveryBashDatabasesFirewallsLinuxMonitoring PlatformsNetworkingPowershellPythonStorage SystemsVirtualizationWindows Server
7 Days AgoSaved
Remote
US
104K-178K Annually
Senior level
104K-178K Annually
Senior level
Software • Financial Services
Lead reliability, observability, and resilience for cloud-based financial SaaS. Define SLOs/SLIs, design monitoring/tracing, own incident response and runbooks, build IaC and automation, implement AIOps, perform chaos and load testing, and write production-grade Python tooling while ensuring security and compliance.
Top Skills: AnsibleAWSAzureBashCi/CdCloudFormationDatadogElkGitGrafanaNew RelicPowershellPrometheusPythonTerraform
Reposted 7 Days AgoSaved
Remote
US
101K-161K Annually
Senior level
101K-161K Annually
Senior level
Cloud • Software • Analytics
Join Arista Networks as a Site Reliability Engineer to manage CloudVision service reliability, scalability, and stability in a FedRAMP environment, focusing on areas like architecture, security, and performance optimization.
Top Skills: AnsibleBashGCPGkeGoKubernetesPulumiPython
Reposted 7 Days AgoSaved
Remote
United States
Senior level
Senior level
Big Data
You will manage AWS infrastructure, automate deployments, debug application issues, and improve the operational health of Metabase Cloud.
Top Skills: AWSDatadogGoGrafanaKubernetesPrometheusPythonTerraform
8 Days AgoSaved
Remote
USA
140K-170K Annually
Mid level
140K-170K Annually
Mid level
Artificial Intelligence • Software • Generative AI • Automation
Operate and harden Blitzy's self-hosted, Kubernetes-based AI platform inside customer-controlled secure cloud environments. Own deployments, upgrades, capacity planning, observability, incident response, and customer-facing technical coordination while championing security and feeding operational learnings back into the product roadmap.
Top Skills: Alerting)BashCloud (Aws/Gcp/Azure)Container OrchestrationGoInfrastructure-As-CodeKubernetesMetricsObservability (LoggingPulumiPythonTerraformTracing
All Filters
JobType
New Jobs
Job Category
Experience
Industry
Company Name
Company Size

Sign up now Access later

Create Free Account