Maximum of 25 job preferences reached.
Top Remote Reliability Engineer Jobs in Chicago, IL
Information Technology • Productivity • Software • Infrastructure as a Service (IaaS)
Own the reliability, performance, architecture, and lifecycle of heterogeneous database platforms. Design and operate relational, NoSQL, caching, time-series, and warehouse systems; improve provisioning, developer tooling, CI/CD, backups, disaster recovery, security, and performance. Influence architecture and product decisions, lead migrations and upgrades, troubleshoot developer and production environments, and establish platform standards, documentation, and regulatory traceability.
Top Skills:
Amazon AuroraAmazon RdsAWSAzureBigQueryC++Ci/CdCloudFormationFargateGCPGoIamJavaKotlinKubernetesPostgresRedshiftSnowflakeTerraform
Fintech • Software • Financial Services
Architects and optimizes high-volume database platforms supporting trading applications. Responsibilities include MySQL schema design, SQL and query optimization, migrations, scalability, reliability, observability, production issue resolution, developer tooling, and database architecture. The role partners with engineering and platform teams, leads technical initiatives, establishes database best practices, mentors engineers, and uses Google Cloud services to improve performance, resilience, and operational efficiency.
Top Skills:
C#Cloud BuildCloud DeployCloud LoggingCloud MonitoringCloud SqlCloud StorageGoGoogle Cloud PlatformJavaKotlinMySQLPostgresPub/SubPythonScalaSecret ManagerSQL
Aerospace • Defense • Manufacturing
The Reliability Engineer develops reliability, availability, and maintainability strategies for defense systems and equipment. Responsibilities include supportability analysis, risk assessment, asset criticality ranking, failure and performance analysis, reliability metrics monitoring, root cause analysis, FMEA, fault-tree analysis, mitigation planning, maintenance optimization, and participation in design reviews. The role supports U.S. Department of Defense customers and requires collaboration with maintenance, operations, and engineering teams.
Top Skills:
ConfluenceExcelMicrosoft Office SuitePowerPointSharepointTeamcenterVisioWindchillWord
Fintech • Information Technology • Software • Financial Services
Own the durability, recoverability, performance, and security of a production PostgreSQL/RDS fleet supporting a live trading platform. Lead replication, failover, backup and restore, disaster-recovery drills, data lifecycle management, access control, encryption, and database observability. Investigate engine-level performance issues including WAL contention, replica lag, bloat, and locking. Build infrastructure and AI-assisted operational tooling while documenting runbooks and reliability decisions.
Top Skills:
AlloydbAmazon AuroraAmazon RdsAWSBashBigQueryElkGrafanaKafkaKubernetesLinuxPostgresPrometheusPythonSQLTerraform
Logistics • Transportation • 3PL: Third Party Logistics
Designs and operates Redwood’s internal engineering platform, including cloud infrastructure, Kubernetes, Terraform, Helm, GitOps, CI/CD, observability, and developer self-service capabilities. The role improves reliability, security, scalability, incident response, operational readiness, and cloud efficiency through automation and cross-functional collaboration with engineering, architecture, security, data, and MLOps teams.
Top Skills:
AksArgocdAWSAzureAzure DevopsBashCi/CdDatadogDockerEksFluxGithub ActionsGitopsGkeGoGrafanaHelmKubernetesLinuxLogicmonitorNew RelicOidcPowershellPrometheusPythonTerraform
Artificial Intelligence • Machine Learning
Own and modernize Domino's Tempest scale-testing platform; build repeatable automated validation, sizing guidance, and cloud-scale test automation; partner with platform teams to enable multi-cloud scale testing and improve test reliability and reporting.
Top Skills:
Ci SystemsCloud PlatformsCloud-Native ToolingEnd-To-End FrameworksKubernetesMulti-CloudPerformance/Load Testing FrameworksPythonTempest
Fintech • Information Technology • Insurance • Financial Services • Big Data Analytics
Leads architecture, modernization, optimization, and reliability initiatives for mainframe CICS, MQ, and z/OS Connect environments. Provides technical direction across development and operations teams, establishes governance and change processes, tunes performance using telemetry, resolves incidents, and develops modernization roadmaps. Collaborates with stakeholders and enterprise architects to deliver secure, scalable, high-availability solutions while evaluating automation, cloud integration, and AI technologies.
Top Skills:
AnsibleCicsCobolDevOpsIbm MqIbm Z/OsOpenshiftPythonRed Hat Ansible Automation PlatformZ/Os ConnectZlinux
18 Days AgoSaved
Easy Apply
Easy Apply
Cloud • Security • Software • Cybersecurity • Automation
Provide technical direction for GitLab Dedicated, a managed single-tenant SaaS platform. Lead architecture and transformation across resilience, failover, tenant orchestration, change management, automation, and platform integrations. Identify systemic reliability and scalability risks, establish reusable platform patterns, strengthen service ownership, and guide cross-team technical decisions. Mentor senior engineers and advance engineering excellence across the organization.
Top Skills:
Cloud InfrastructureDevsecopsDistributed SystemsGoInfrastructure As CodeObservabilityPythonRuby
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Operate and improve large-scale HPC GPU clusters for AI model training and inference. Build observability, automation, CI/CD, and incident response tooling; lead on-call rotations, ensure security/compliance, collaborate with ML engineers, and optimize capacity and costs for production ML workloads.
Top Skills:
AWSAzureBashCi/CdContainer OrchestrationDatadogDockerGCPGoGpu ClustersGrafanaHpcInfrastructure-As-CodeKubernetesKubernetes OperatorsOpentelemetryPython
Healthtech • Information Technology • Software • Telehealth
Develop, monitor, and maintain distributed production systems and AWS-based microservices infrastructure. Build automation, tooling, and repeatable processes that improve uptime, scalability, security, and operational efficiency. Support product engineering teams with performance, scaling, incident diagnosis, and production debugging. Analyze and tune systems, code, and networking while participating in on-call operations and blameless post-mortems.
Top Skills:
AWSDnsDockerGCPGenaiHttp/HttpsKubernetesLoad BalancersNtpReverse ProxiesTcp/IpTlsWeb Application Firewalls
Software • Defense
Own reliability, scalability, security, observability, and incident response for production applications across AWS and on-premises DoD environments. Build monitoring and alerting, define SLIs and SLOs, lead post-incident reviews, automate infrastructure with Terraform and Ansible, operate Kubernetes clusters, embed RMF and STIG controls, reduce operational toil, and support secure air-gapped deployments.
Top Skills:
AlloyAnsibleAWSAws GovcloudBashDatadogElk StackGithub ActionsGitlab Ci/CdGitopsGoGrafanaHyper-VIstioJenkinsKubernetesLinkerdLokiNutanixPrometheusProxmoxPythonRmfSecurity+StigsTerraformVMware
Blockchain • Fintech • Payments • Financial Services • Cryptocurrency • Web3 • Infrastructure as a Service (IaaS)
Manage AWS and GCP cloud environments, scale infrastructure globally, shape technical architecture, and build reliable CI/CD pipelines. Automate security and compliance controls, improve infrastructure performance, and develop systems interacting with smart contracts across multiple blockchains. The role requires Terraform, shell scripting, GitHub Actions, Docker, production cloud operations, and observability experience, with Kubernetes, networking, and fintech compliance knowledge preferred.
Top Skills:
AWSCi/CdDockerFirewallsGCPGithub ActionsGkeGoHelmInfrastructure As CodeKubernetesLoad BalancersMtlsNode.jsPciShellSsl/TlsTerraformTypescriptVpcZero Trust
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Fintech • Software • Automation
Own and improve the reliability of Kintsugi’s AWS and managed Kubernetes infrastructure. Build automation, internal tools, observability, CI/CD pipelines, and infrastructure-as-code workflows while reducing operational toil. Lead incident response, reliability architecture, developer enablement, cost optimization, security, compliance, and disaster recovery efforts. Work agent-first and partner with Platform Engineering, Product, QA, and software teams to deliver scalable, secure systems.
Top Skills:
AWSCi/CdCloudFormationInfrastructure As CodeKubernetesObservabilityPostgresRedisTerraform
Software
Own the operational health, observability, reliability, and failover of AI model providers and endpoints. Build SLO monitoring, canaries, quality regression detection, automated provider-operations tooling, and load-testing systems. Lead incident response, communicate with external providers, produce reliability scorecards, and drive postmortems and corrective actions. Support capacity planning and launch readiness for high-volume LLM inference traffic.
Top Skills:
ClickhouseCloudflare WorkersGCPPostgresPythonTypescriptVercel
Legal Tech • Software
Design and improve observability (monitoring, logging, tracing, SLIs/SLOs), build automation and CI/CD, lead incident response and reliability improvements, mentor SREs, run on-call, and apply AI/ML to operational signals to forecast and reduce risks.
Top Skills:
AWSBashCi/CdDistributed TracingGoInfrastructure As CodeKubernetesLoggingMonitoringPythonSlisSlos
eCommerce
Own the reliability, availability, security, and observability of Tradeweb’s global AWS platform. Responsibilities include infrastructure-as-code automation, monitoring, SLO development, incident triage and resolution, performance analysis, architecture collaboration, and regular on-call support. The role requires cloud-native engineering, scripting, networking, Linux/Unix, security, and reliability expertise while contributing to an Agile engineering organization.
Top Skills:
Amazon EksAmazon SmsAmazon SnsArgocdAWSAws LambdaGitsecopsKubernetesKustomizeLgtmLinuxPulumiPythonUnix
Big Data • Software
Owns production reliability for a B2B SaaS platform, including SLOs, error budgets, observability, alerting, incident response, recovery exercises, performance and capacity engineering, load testing, disaster recovery, and production readiness. The role writes automation and production code, leads incidents, coaches engineering teams, and improves reliability across Kubernetes, Azure, databases, messaging, and service infrastructure.
Top Skills:
.NetAksAsp.NetAWSAzureAzure Service BusC#Dora MetricsFluxGatlingGitopsGoGrafanaIncident.IoIstioJmeterK6KafkaKubernetesLokiMongoDBNist 800-171PagerdutyPrometheusPythonSentrySignalrSoc 2TempoTemporalTypescript
Cloud • Security • Software • Cybersecurity
Performs site reliability engineering for large-scale cloud infrastructure, focusing on application and network performance, reliability, security, scalability, and capacity. Deploys highly available systems, automates cloud service deployments, monitors and troubleshoots services, analyzes logs and events, maintains SLAs, and resolves infrastructure issues. The role requires expertise in microservices, container orchestration, cloud migration, deployment automation, and multiple DevOps technologies.
Top Skills:
AzureChefCloud ComputingContainer OrchestrationGitIntellij IdeaJenkinsKubernetesLinuxMicroservicesMySQLPostgresPythonTerraformVMware
Cloud • Security • Software • Cybersecurity
Analyzes and resolves availability and performance issues in large-scale production and lab environments. Develops automation, monitoring, alerting, log analysis, debugging tools, and systems programming solutions. Collaborates with software development and engineering teams on CI/CD, platform architecture, incident resolution, and operational best practices within an agile SDLC.
Top Skills:
AlertingAnsibleAutomation ScriptingBashContinuous DeliveryContinuous IntegrationDevOpsLinuxLog AnalysisMonitoringPowershellSystems Programming
Fintech • Real Estate • Software
Lead reliability and observability efforts across the org: design and maintain Kubernetes and AWS infrastructure, build CI/CD pipelines, drive IaC standards (Terraform/Crossplane), partner with 16+ teams to roll out tools and processes, participate in on-call rotation and incident response, and use AI tools to accelerate work.
Top Skills:
Ai ToolsArgoAurora PostgresAWSCi/Cd PipelinesCrossplaneDatadogDocumentdb (Mongo)EcsEksGithub ActionsHelmKubernetesMongoDBPostgresRdsTerraform
2 Days AgoSaved
Cloud • Security • Software • Cybersecurity
Develop and operate Akamai’s Linux-based cloud infrastructure, including bare-metal and virtual machine platforms, kernels, operating systems, and KVM/QEMU virtualization. Build software, automation, observability, and infrastructure enhancements at scale; troubleshoot complex distributed-system issues; collaborate across engineering and operations; mentor team members; and participate in on-call service restoration.
Top Skills:
AnsibleChefKubernetesKvm/QemuLinuxLinux KernelNested VirtualizationPuppetSaltstack
Software
Own production reliability and infrastructure operations for the GrayKey cloud platform. Responsibilities include AWS monitoring, incident response, vulnerability remediation, networking, IAM and Okta administration, EKS and Argo CD operations, backups, recovery testing, Terraform changes, cost optimization, automation, documentation, and compliance support. The role requires independent work during Pacific and Mountain time-zone hours and involves less than 5% travel.
Top Skills:
Active DirectoryAmazon EksAnsibleArgo CdAWSAzure AdBashDatadogDnsDockerEc2Entra IdFortigateGitlab Ci/CdHelmIamKubernetesLambdaLinuxOidcOktaPostgresPythonRdsRedisSAMLTcp/IpTerraformTlsVpcVpn
Artificial Intelligence • Information Technology • Consulting
The Linux Systems Administrator will maintain and troubleshoot Linux systems, support network services, and work on systems integration while collaborating with infrastructure teams.
Top Skills:
DhcpDnsLinuxNtpPython
Digital Media • Gaming • Information Technology • Software • Sports • Esports • Big Data Analytics
Lead reliability, scalability, and operational excellence of large-scale database platforms across cloud and on-prem. Build automation-first database infrastructure (Kubernetes operators, IaC, GitOps), drive monitoring/SLOs, incident leadership, performance and cost optimization, and partner with application teams on safe schema/migration practices. Mentor engineers and evaluate AI-assisted workflows to improve productivity and reliability.
Top Skills:
AerospikeArgocdAuroraClaudeCloud SqlCursorDatabase OperatorsEksFluxcdGithub CopilotGitopsGkeGoKubernetesMcpMongoDBMySQLPersistent VolumesPostgresPulumiPythonRedisScylladbStatefulsetsTerraform
Information Technology
Leads SRE and cloud operations for highly available, scalable platforms and AI-powered solutions. Designs AWS infrastructure, Kubernetes environments, CI/CD pipelines, Infrastructure as Code, observability, automation, and incident response processes. Develops AI-driven operational capabilities, supports MLOps and model lifecycle management, establishes reliability metrics and SLOs, and mentors engineering teams. Partners with software, data, machine learning, security, and product teams to improve platform performance, resilience, and operational excellence.
Top Skills:
Amazon SagemakerAWSAws BedrockAzure DevopsAzure OpenaiCi/CdCloudFormationCloudwatchDatadogDockerEc2EcsEksElk StackGithub ActionsGitlab Ci/CdGrafanaIamInfrastructure As CodeJenkinsKubernetesLambdaLangchainMlopsNvidia AiOpenai ApisPrometheusPythonRdsS3ShellSplunkTerraformVpc
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Chicago, IL Companies Hiring Remote Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results

































