Maximum of 25 job preferences reached.
Top Remote Senior Site Reliability Engineer Jobs in Chicago, IL
Reposted One Month AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
Develop and maintain Kubernetes runtime environments, support developers, resolve critical issues, and participate in on-call rotations for production systems.
Top Skills:
AWSAzureCert-ManagerCorednsCrdsCriCsiGatekeeperGCPGoHelmKubernetesKustomizeOperatorsPythonTerraform
Cloud • Information Technology • Cybersecurity • Infrastructure as a Service (IaaS)
Owns reliability, observability, and incident response for a GPUaaS platform. Defines SLOs, builds monitoring and alerting systems, leads major incidents and post-incident reviews, automates operational processes, maintains runbooks, manages on-call operations, coordinates with engineering teams, drives chaos testing, reports SLA performance, and mentors junior engineers.
Top Skills:
DatadogGoGpuaasGrafanaGremlinHpcKubernetesLitmusOpentelemetryPrometheusPython
Artificial Intelligence • Cloud • Software • Infrastructure as a Service (IaaS)
Ensure stability and resilience of Runpod's distributed AI platform by defining SLIs/SLOs, leading incident response, building observability and reliability tooling, automating operational workflows, and partnering with engineering teams to reduce toil and improve production readiness.
Top Skills:
BashCi/CdContainerized Production SystemsGoGpu Observability ToolingGrafanaInfrastructure As CodeLinuxPrometheusPython
Edtech
Lead infrastructure modernization and platform reliability across multiple cloud providers. Design infrastructure as code, operate Kubernetes and Linux environments, improve CI/CD and deployment tooling, establish SLI/SLO practices, strengthen observability, lead incident response, manage cloud costs, and partner on security and compliance. Provide technical leadership through architecture guidance, mentorship, engineering standards, and roadmap development while participating in on-call support.
Top Skills:
AWSCi/CdGCPJenkinsKubernetesLinuxPythonRubyRuby On RailsSoc 2SpinnakerTerraform
Cloud • Information Technology • Business Intelligence • Consulting
Design, build, and operate cloud infrastructure and SRE capabilities for an enterprise AI platform. Responsibilities include infrastructure-as-code, landing zones, networking, Kubernetes, CI/CD, observability, incident response, SLOs, production readiness, automation, cost optimization, and support for hybrid, edge, air-gapped, and customer-controlled environments. The role is remote, client-facing, and requires strong collaboration and reliability ownership.
Top Skills:
AlertingAzureAzure ArcAzure DevopsBicepCi/CdDashboardsDockerGithub ActionsGpu WorkloadsInfrastructure As CodeKubernetesLogsMetricsObservabilitySlosTerraformTraces
Cloud • Security • Software • Cybersecurity
Design, build, and operate scalable infrastructure and CI/CD/IaC systems. Implement observability (monitoring, logging, alerting), automate reliability improvements, mentor engineers, collaborate on incident response, and participate in on-call rotations to maintain Akamai Cloud services.
Top Skills:
AlertingAnsibleBashChefCi/CdGithub ActionsGitlab Ci/CdGoInfrastructure As CodeJenkinsLoggingMonitoringPuppetPythonSaltstackTelemetryTerraform
Cloud • Security • Software • Cybersecurity
The Site Reliability Engineer II ensures the reliability, availability, performance, and security of critical cloud systems and services. Responsibilities include developing automation for provisioning and configuration management, maintaining monitoring and alerting, optimizing infrastructure performance, supporting high availability, and enabling continuous integration and delivery. The role collaborates with security teams and drives operational improvements across cloud and network infrastructure.
Top Skills:
AnsibleAWSAzureChefContinuous DeliveryContinuous IntegrationDnsElk StackGCPGoGrafanaHTTPKubernetesLinuxPrometheusPuppetPythonShellTcp/IpUnix
Artificial Intelligence • Cloud • Information Technology • Software • Big Data Analytics
Own reliability for Kong’s Volcano internal developer platform by defining SLOs, incident practices, and observability. Design multi-region Kubernetes infrastructure, GitOps deployment automation, preview environments, managed PostgreSQL, Redis, and object storage. Lead reliability and compliance initiatives with engineering, security, and OCTO leadership while evaluating emerging edge, serverless, vector database, and AI infrastructure technologies.
Top Skills:
ArgocdAutoscalingCi/CdCniDatadogGitopsGrafanaHelmIngressKubernetesObject StoragePostgresPrometheusRedisService MeshTerraformTerragrunt
Cloud • Security • Software • Cybersecurity
Build and maintain reliable, scalable cloud compute platforms across distributed services. Troubleshoot Linux, networking, and production issues; develop automation and AI-assisted tooling; improve monitoring, alerting, SLIs, and SLOs; conduct incident response and root cause analysis; and partner with engineering teams on system design, deployment safety, and operational readiness.
Top Skills:
AnsibleDnsDockerElkGoGrafanaKubernetesLinuxLokiNomadOpensearchPodmanPrometheusPythonSaltTcp/IpTerraform
Artificial Intelligence • Healthtech • Software • Telehealth
Designs, deploys, and maintains resilient AWS and Kubernetes infrastructure. Builds automation, GitHub Actions components, internal AI-assisted operational tools, and observability systems. Leads incident response, postmortems, and SLO/SLI management while ensuring HIPAA compliance and high availability. Collaborates across teams on architecture reviews, risk reduction, clinical safety, and reliability best practices.
Top Skills:
Ai-Assisted OperationsAmazon Ec2Amazon EksAmazon RdsAmazon S3AWSBashDatadogGithub ActionsGoHelmKubernetesPythonTerraform
Cloud • Security • Software • Cybersecurity
Lead site reliability engineering for Akamai’s compute infrastructure and services. Develop automation, define reliability requirements and standards, establish SLOs and KPIs, troubleshoot complex distributed-system and hardware issues, manage incidents and postmortems, and participate in on-call rotations. Coordinate engineering teams, support strategic initiatives, mentor engineers, and guide restoration of service-impacting issues.
Top Skills:
Cloud InfrastructureDistributed SystemsInfrastructure AutomationLinux
Artificial Intelligence • Healthtech • Software
Build and maintain scalable infrastructure platforms supporting domestic and international workloads. Develop CI/CD, declarative lifecycle management, Kubernetes clusters, monitoring, and automation systems. Troubleshoot infrastructure issues, respond to alerts, minimize downtime, streamline delivery pipelines and database changes, and help guide SRE team direction. Collaborate with engineers, data scientists, and technology professionals while promoting a high-performance, cross-functional culture.
Top Skills:
AWSAzureContainerdContinuous DeploymentContinuous IntegrationDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
New
Cut your apply time in half.
Use ourAI Assistantto automatically fill your job applications.
Use For Free
Software
Operate and maintain highly available, secure, containerized SaaS applications across AWS and Azure. Responsibilities include rotating 24x7 on-call coverage, observability, incident and security response, disaster recovery, Terraform infrastructure-as-code, CI/CD automation, cloud integration, performance optimization, and developing AI agents to automate SRE and DevSecOps workflows.
Top Skills:
Amazon Web ServicesCheckovClaude CodeDatadogDockerDynatraceGithub ActionsGitlab CiGoJavaScriptKubernetesAzureNew RelicPrisma CloudPythonTerraformWiz
Reposted 2 Months AgoSaved
Easy Apply
Easy Apply
Big Data • Cloud • Software • Database
The Senior Site Reliability Engineer will lead security design and implementation for cloud infrastructures, mentor teams, and automate security solutions.
Top Skills:
AnsibleAWSAzureCloud Security ToolsCloudFormationGCPGoTerraform
Healthtech
Lead the migration from legacy Azure services to a Kubernetes-based, containerized microservices platform. Design, build, and scale infrastructure, implement observability (monitoring/alerting/logging), drive incident response and SLOs, automate with IaC and CI/CD, optimize cost and networking, mentor teams, and document systems to ensure reliable, scalable healthcare platform operations.
Top Skills:
.NetAWSAzureAzure Entra IdBashC#DatadogGCPGithub ActionsGitlab Ci/CdGrafanaHelmKubernetesPrometheusPythonTerraform
Cybersecurity
Owns the FedRAMP-authorized AWS GovCloud environment, ensuring reliability, security, compliance, patching, vulnerability remediation, continuous monitoring, deployments, hardening, audits, incident response, and on-call support. The role requires hands-on Kubernetes, AWS, infrastructure-as-code, CI/CD, vulnerability management, automation, and regulated-environment experience, while collaborating with SRE and security teams.
Top Skills:
AnsibleAws GovcloudCloudFormationFips-Validated CryptographyGithub ActionsGoJenkinsKubernetesNessusOpenscapPythonQualysRubySIEMTenableTerraform
Digital Media
Build, maintain, and operate Ookla’s globally distributed infrastructure platform at massive scale. Responsibilities include managing cloud instances, containers, serverless applications, databases, streaming systems, and big-data tooling; supporting 24/7 production operations and on-call rotations; implementing security programs; improving deployment pipelines, monitoring, observability, and reliability; and guiding software and data engineering teams on operational best practices and troubleshooting.
Top Skills:
Amazon AuroraAmazon RdsAnsibleSparkAWSChefCloudFormationDockerDynamoDBGitGitGoIds/IpsJavaKafkaKinesisKubernetesLinuxMongoDBMySQLPHPPostgresPythonRubySQLTerraformTypescript
Reposted One Month AgoSaved
Software • Defense
Work as an SRE embedded with product teams to improve reliability by fixing application code (primarily TypeScript), building observability (Prometheus, Loki, Grafana, Alloy), defining SLIs/SLOs, leading incident response and postmortems, automating toil, and supporting deployments across on‑prem DoD and AWS environments.
Top Skills:
AlloyAWSBashContainersDockerGithub ActionsGitlab Ci/CdGoGrafanaJenkinsKubectlKubernetesLokiNode.jsPrometheusPythonTypescript
Healthtech • Software
Support reliable deployment and daily operations across AWS accounts, Linux workloads, AWS data services, networking, and Snowflake connectivity. Investigate infrastructure incidents, deployment failures, outages, and access issues. Improve automation, configuration management, monitoring, alerting, CI/CD, infrastructure-as-code practices, deployment consistency, and operational documentation. Collaborate with application engineering, data engineering, security, and other stakeholders to maintain secure, available, and dependable environments.
Top Skills:
Amazon KinesisAmazon S3AWSAws GlueAws PrivatelinkBashCi/CdDnsInfrastructure As CodeLinuxPythonSnowflakeTcp/IpTls
Aerospace • Manufacturing
Build and lead a centralized observability platform for satellite, ground-station, and distributed network systems. Responsibilities include scaling metrics, logging, and tracing infrastructure; defining SLOs, SLIs, and error budgets; enabling application instrumentation; automating deployments with Terraform and ArgoCD; monitoring Kubernetes, GCP, and AWS environments; and developing incident response, alerting, and reliability practices. The role includes on-call responsibilities and requires an active Top Secret/SCI clearance.
Top Skills:
ArgocdAWSC++ElkGitlab CiGoGoogle Cloud PlatformGrafanaHoneycombIstioJaegerJavaKubernetesLinkerdLokiOpentelemetryPrometheusPythonTempoTerraform
Information Technology • Security • Software • Cybersecurity
Owns production reliability for high-throughput, low-latency systems, including observability, SLOs, incident response, capacity planning, infrastructure as code, progressive delivery, chaos testing, and operational tooling. The role requires hands-on software and infrastructure engineering, production Redis/ElastiCache expertise, cloud infrastructure knowledge, security awareness, mentoring, and participation in on-call operations.
Top Skills:
AWSDatadogEksElasticacheGoGrafanaKubernetesOpentelemetryPrometheusPythonRedisTerraform
Artificial Intelligence • Cloud • Machine Learning • Software • Database • App development • Generative AI
Design and maintain reliable, scalable infrastructure for Replit’s global platform. Responsibilities include building observability and alerting systems, automating infrastructure with infrastructure-as-code, managing CI/CD pipelines, defining SLOs and SLIs, leading incident response and postmortems, maintaining runbooks, optimizing performance, and improving capacity, availability, and recovery times.
Top Skills:
AnsibleCi/CdDatadogGoGoogle Cloud PlatformGrafanaKubernetesPrometheusPulumiPythonTerraform
Software
Own reliability, performance, scalability, and operational standards across on-premises, private-cloud, and AWS environments. Build infrastructure-as-code, deployment automation, monitoring, alerting, and observability tooling; define SLOs, SLIs, error budgets, and readiness standards. Lead incident response, root-cause analysis, disaster-recovery readiness, release coordination, and preventive automation. Mentor engineers, coach teams on operational practices, coordinate on-call coverage, and ensure infrastructure meets security and compliance requirements.
Top Skills:
AnsibleAuto ScalingAWSBashCi/CdClaudeCloudwatchCrowdstrikeDatadogDistributed SystemsDnsDockerEc2EcsGitGithub CopilotGitlabGitopsIamKubernetesLinuxLoad BalancingOracle LinuxPythonQualysRapid7RhelS3Tcp/IpTerraformVpcWireshark
Healthtech • Software
Design, automate, and maintain scalable infrastructure and SRE tooling. Manage Kubernetes clusters, CI/CD, monitoring, and incident response. Improve processes, reduce toil via automation, and collaborate with engineering and data teams to support domestic and international workloads.
Top Skills:
AWSAzureContainerdDnsDockerFirewallsGCPGoGrpcHelmKubernetesLinuxLoad BalancingPrometheusPythonRoutingShell ScriptingTcp/IpUdp
Healthtech • Software
Own production reliability across cloud environments by defining SLOs, SLIs, error budgets, observability practices, and actionable dashboards. Troubleshoot incidents, perform root cause analysis, manage capacity planning, and improve service performance and availability. Build automation to reduce toil, support release management, maintain runbooks, and collaborate with engineering teams on operational readiness. Manage tools including Datadog, Prometheus, and Grafana while advancing monitoring, logging, metrics, tracing, and reliability practices.
Top Skills:
.NetArgoAWSDatadogGithub ActionsGitlabGoGrafanaHelmJavaJenkinsKargoKubernetesLinuxOpentelemetryPrometheusPythonSplunkWindows
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Chicago, IL Companies Hiring Remote Senior Site Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results










.png)








.jpg)








.png)

