Maximum of 25 job preferences reached.
Top Remote Reliability Engineer Jobs in Chicago, IL
Real Estate • PropTech
Lead technical strategy and implementation for Redfin's production databases and storage. Architect, scale, and support cloud database/storage solutions (self-managed and AWS managed), drive reliability, observability, backup/recovery, DR planning, incident response and root cause analysis, mentor engineers, and participate in on-call rotation.
Top Skills:
Amazon AuroraAmazon RdsAmazon S3Anthropic Claude CodeAWSConfiguration ManagementCursorDynamoDBElasticacheGithub CopilotInfrastructure As CodeLinuxOpensearchPostgresPython
Aerospace • Artificial Intelligence • Machine Learning • Robotics • Software
Lead operational reliability and platform enablement for Databricks: build monitoring, CI/CD, deployment standards, compute and job policies, observability, runbooks, and governance to support secure, cost-aware, production data workloads across regulated environments. Mentor engineers and align platform with cloud/infrastructure and compliance requirements.
Top Skills:
Ci/CdDatabricksDatabricks Asset BundlesDatabricks WorkflowsDelta LakeInfrastructure-As-CodeService PrincipalsUnity CatalogVersion Control (Git)
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Support and improve reliability, scalability, and performance of large-scale databases across cloud and on-prem. Build automation, Kubernetes operators, GitOps workflows, and tooling in Go or Python; implement observability, capacity planning, failover, backups, schema migrations, and self-healing systems; partner with application teams and leverage AI to improve operations and engineering productivity.
Top Skills:
AerospikeArgocdAurora MysqlClaudeCursorEksFluxcdGithub CopilotGitopsGkeGoKubernetesKubernetes OperatorsMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills:
AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Fintech • Financial Services
Ensure high availability and reliability of OnePay integrations by troubleshooting distributed systems, driving root-cause analysis, enhancing observability (e.g., Splunk/New Relic), monitoring production, supporting CI/CD (Jenkins), coordinating releases, participating in on-call rotation, and implementing reliability and process improvements.
Top Skills:
AnsibleBashChefGoJavaScriptJenkinsNew RelicPowershellPuppetPythonSplunkTerraformUnix
Artificial Intelligence • Machine Learning
Lead development of AI-assisted reliability tooling, own incident response end-to-end, improve observability and SLO/SLI frameworks, scale single-tenant SaaS operations, mentor engineers, and reduce recurring operational toil through engineering and automation.
Top Skills:
Cloud PlatformsGoKubernetesLinuxLlm/Ai ToolingLogs And TracingObservability ToolingPythonSlo/Sli Frameworks
eCommerce • Fintech • Payments • Software
The role involves ensuring software reliability and performance, managing incidents, developing infrastructure automation, and mentoring junior engineers within a platform team.
Top Skills:
AWSCloudFormationDatadogKubernetesOpentelemetryRubyRuby On RailsTerraform
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead long-term strategy and architecture for cloud and on‑prem platform infrastructure, driving Kubernetes and multi‑cloud reliability, IaC/GitOps automation, observability, SLO/SLI/error‑budget practices, incident leadership, AI‑augmented tooling adoption, and mentorship of senior engineers to improve platform resilience and developer experience.
Top Skills:
Amazon Elastic Kubernetes Service (Eks)AutoscalingAWSCapacity PlanningCi/CdGitopsGoGoogle Cloud PlatformGoogle Kubernetes Engine (Gke)Identity And Access ManagementInfrastructure As CodeKubernetesLinuxNetworkingObservabilityOperatorsPulumiPythonRke2StorageTerraform
Legal Tech • Software
Lead automation and optimization of Filevine's data platform: performance tune MSSQL/Postgres, optimize Snowflake, provision infrastructure with Terraform/AWS, run stateful containers on Kubernetes, integrate AI/LLM and MCP for operational automation, manage CI/CD, capacity planning, documentation, and serve in 24/7 on-call rotation.
Top Skills:
AWSC#DapperDockerDynamoDBEntity FrameworkGitlabKubernetesLlmsMcp (Model Context Protocol)Microsoft Sql Server (Mssql)Octopus DeployOpensearchPostgresPowershellPythonRedisSnowflakeTerraform
Reposted 16 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
As a Senior Site Reliability Engineer, you'll design and build complex systems, support Atlas platform operations, automate processes, and ensure high availability of services.
Top Skills:
AWSAzureDnsGCPGoHTTPLinuxPythonRubyTls
8 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Fintech • Mobile • Payments • Financial Services
Lead design and delivery of a reliability platform for production systems. Own quarterly goals, guide engineers through ambiguity, collaborate with PM/design/analytics, monitor and operate services (on-call), set quality and code standards, and mentor teammates. Integrate AI/LLM tooling to improve automation, debugging, and service health.
Top Skills:
Ai FrameworksAWSKotlinKubernetesLlmsMySQLPythonReactVue
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Operate and improve large-scale HPC GPU clusters for AI model training and inference. Build observability, automation, CI/CD, and incident response tooling; lead on-call rotations, ensure security/compliance, collaborate with ML engineers, and optimize capacity and costs for production ML workloads.
Top Skills:
AWSAzureBashCi/CdContainer OrchestrationDatadogDockerGCPGoGpu ClustersGrafanaHpcInfrastructure-As-CodeKubernetesKubernetes OperatorsOpentelemetryPython
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Cloud • Information Technology • Security • Software • Cybersecurity
This internship role focuses on SRE skills, requiring collaboration and problem-solving in dynamic environments for Zscaler's Zero Trust Exchange team.
Top Skills:
AnsibleAws EcsKubernetesLinuxPythonTerraform
Software
Lead resolution of high-severity post-sales escalations: investigate and reproduce complex networking and software issues, run customer calls, collect logs and evidence, create engineering-ready handoff packages, identify systemic patterns, and improve tooling and runbooks to raise Support's technical capability.
Top Skills:
Ci/CdCloud IaasContainersDnsFirewallsGoJSONKubernetesLinuxLoad BalancersmacOSNatPacket-Level AnalysisPprofPythonRoutingStunTcp/IpUdp Hole PunchingVpnWindowsWireguard
Reposted 19 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Reposted 11 Days AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will lead security design and implementation for cloud infrastructures, mentor teams, and automate security solutions.
Top Skills:
AnsibleAWSAzureCloud Security ToolsCloudFormationGCPGoTerraform
Healthtech • Software
Operate and maintain AWS-hosted MERN applications and large-scale data workflows. Manage serverless and Spark-based pipelines, perform incident response and on-call duties, engineer automation to eliminate operational toil, ensure HIPAA/SOC2/HITRUST compliance, build observability and lead blameless post-mortems.
Top Skills:
Amazon EcsAmazon EksAmazon EmrAthenaAws GlueAws LambdaAws SnsAws SqsCloudwatchEc2IamJavaScriptMernMySQLNode.jsOpentofuPysparkPythonRabbitMQTerraformTypescriptVpc
Software
Lead reliability engineering for Silicon Photonics hardware: define and validate reliability models, perform MTBF/MTBCF predictions, analyze field data, direct verification testing and root-cause analysis, drive corrective actions, and mentor cross-functional teams to improve product reliability.
Top Skills:
Derating AnalysisDfmeaMtbcfMtbfSherlockSilicon PhotonicsTelcordiaThermal DesignWindchill Qs
Database • Analytics
As a Database Reliability Engineer at ClickHouse, you'll improve reliability, manage escalation processes, support incident response, and enhance database performance while collaborating across teams.
Top Skills:
AWSAzureC++ClickhouseGoogle Cloud PlatformPythonShellSQL
Software
Drive reliability qualification and production monitoring for optical communication products. Track and analyze reliability stress tests, investigate failures with cross-functional teams, implement corrective actions, maintain dashboards and reports, support NPI, and apply statistical and AI tools to generate reliability insights and improve product quality.
Top Skills:
AIJmp/JslMinitabSQL
eCommerce • Healthtech • Kids + Family • Retail • Social Media
Own and evolve Babylist's AWS infrastructure and developer platform using Terraform and Kubernetes. Improve CI/CD reliability, support engineers across environments, define monitoring and alerting standards, lead incident response and postmortems, and shape platform architecture to scale for millions of users.
Top Skills:
AWSCdnCircleCICronitorDatadogDnsEksGithub ActionsKubernetesLoad BalancersMySQLPagerdutyRdsRedisRuby On RailsSentrySidekiqTerraform
Information Technology • Software • Database • Automation
Owner of on-prem reliability and escalations: reproduce and resolve L2/L3 issues across heterogeneous Kubernetes environments, build diagnostics and automation, improve CI and e2e test stability, establish performance baselines, harden install/upgrade flows, and write tooling in Python/Go/Rust to reduce repeat incidents.
Top Skills:
BenchmarkingCiCi/CdContainersE2E TestingGoHealth ChecksHelmInstallersIntegration TestingKubernetesLoad GenerationLogsMetricsNetworkingObservabilityPackagingProfilingPythonRbacRustStorageSupport BundlesTraces
16 Days AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Software
Lead ownership and support of production database systems (MySQL CloudSQL and Spanner). Drive schema change processes, build automated deploy/test pipelines, implement monitoring/alerting, manage backup/recovery and DR, perform production debugging/RCA, and evaluate database features for global scalability and uptime.
Top Skills:
Cloud SqlGoogle Cloud Platform (Gcp)Google Cloud SpannerInfrastructure As CodeMysql 8.4RedisShardingSQL
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Popular Job Searches
All Filters
Total selected ()
No Results
No Results





.png)

























