Maximum of 25 job preferences reached.
Top Remote Senior Site Reliability Engineer Jobs in Chicago, IL
Cloud • Security • Software • Cybersecurity
As a Site Reliability Engineer II, you'll automate tasks, monitor AI workloads, enhance dashboards, support CI/CD processes, and collaborate with engineering teams on complex issues while participating in on-call rotations.
Top Skills:
GoGrafanaKubernetesLinuxPrometheusPythonSaltstackTerraform
Automotive
Design and implement scalable cloud infrastructure, monitor performance, automate processes, ensure security and compliance, and lead a DevOps team.
Top Skills:
AWSBashCi/CdDockerElk StackGCPGrafanaKubernetesPrometheusPythonTerraform
Artificial Intelligence • Machine Learning • Software • Analytics
The role involves end-to-end ownership of AWS infrastructure, managing Kubernetes platforms, and ensuring system reliability through observability and automation. Responsibilities include incident response and maintaining CI/CD systems.
Top Skills:
ArgocdAWSDatadogGitGoKubernetesPythonTerraform
Fitness • Healthtech • Software
Own reliability and security of CI/CD and production services: define SLI/SLOs, lead incident response and postmortems, build observability (Datadog), operate Kubernetes and IaC (Terraform), harden pipelines with SAST/DAST/SCA and policy-as-code, and coach teams on reliability and operational best practices.
Top Skills:
AWSCi/CdConftestDastDatadogDockerGithub ActionsGoInfrastructure As CodeKubernetesKyvernoOpa/RegoPythonSastScaTerraformTypescript
Cloud • Security • Software • Cybersecurity
Design, develop, test, and operate scalable infrastructure and services for Akamai Cloud. Implement and manage Infrastructure-as-Code (Terraform and similar tools), CI/CD, and observability. Automate reliability improvements, mentor engineers, collaborate on incident response and root-cause remediation, and participate in on-call rotations.
Top Skills:
Alerting)AnsibleChefCi/CdInfrastructure As CodeLinuxLoggingObservability (MonitoringPuppetSaltstackTerraform
Other
Design, build, and maintain highly available cloud-native systems. Improve reliability through automation, CI/CD, Kubernetes, observability, and incident management. Collaborate with developers, security, and product teams to define SLOs, implement self-healing, debug production issues, and ensure secure deployments.
Top Skills:
AWSAzure Cloud ServicesDatadogGCPGithub ActionsGitlab CiGoInfrastructure As CodeKubernetesOpsgeniePagerdutyPythonRubySite Reliability Engineering Foundation
Artificial Intelligence • Information Technology • Consulting
Build and operate Nebius's network infrastructure: define SLIs/SLOs, improve site and inter-site reliability, lead incident response and postmortems, develop observability and alerting, automate change workflows, and collaborate with network and platform teams to embed operability.
Top Skills:
Ci/CdContainer PlatformsGoInfrastructure As CodeLinuxPython
Software
The role involves managing compute infrastructure for decentralized applications, requiring critical thinking, documentation skills, and experience in Kubernetes and blockchain management.
Top Skills:
BlockchainGitopsInfrastructure-As-CodeKubernetesProgramming Languages
Social Media • Software
Design, implement, and operate infrastructure for a federated social network. Own reliability, availability, observability, incident response, deployments, capacity planning, and cost management. Build automation and tooling, scale bare-metal and cloud systems for millions of users, lead incident reviews, mentor engineers, and manage vendor relationships to ensure operational excellence.
Top Skills:
Bare-MetalCapacity PlanningCloud ServicesColocationDatabasesDeployment And Rollback SystemsGoIncident ResponseKubernetesLinuxMonitoringNetworkingObservability SystemsProduction AutomationStorage
Fintech • Real Estate • Software
Lead reliability and observability efforts across the org: design and maintain Kubernetes and AWS infrastructure, build CI/CD pipelines, drive IaC standards (Terraform/Crossplane), partner with 16+ teams to roll out tools and processes, participate in on-call rotation and incident response, and use AI tools to accelerate work.
Top Skills:
Ai ToolsArgoAurora PostgresAWSCi/Cd PipelinesCrossplaneDatadogDocumentdb (Mongo)EcsEksGithub ActionsHelmKubernetesMongoDBPostgresRdsTerraform
Artificial Intelligence • Other • Sales • Software
The role involves designing and advancing infrastructure for the engineering team, ensuring the reliability of Kubernetes clusters, automating operations, and building machine learning infrastructure.
Top Skills:
ArgoAWSAzureCloudFormationFluxGithub ActionsGoGCPKubernetesPostgresPythonTerraform
Agency • Information Technology
Lead SRE role designing and maintaining CI/CD pipelines (GitHub Actions), containerized deployments (Docker, Kubernetes, AKS, Helm), web/mobile app releases, observability, automated testing, and DevOps best practices across cloud environments with cross-functional collaboration and regulatory compliance.
Top Skills:
AksAndroidAzure Application InsightsAzure Log AnalyticsAzure MonitorBashBranchingDockerDocker ComposeGitGit HooksGithub ActionsGoogle PlayHelmHerokuiOSIos App StoreJavaKubernetesNpmPowershellPull RequestsPythonSonarqubeVeracodeVercel
New
Track Smarter, Apply Better.
Ditch the spreadsheets. Organize your job search with our freeApplication Tracker.
Use For Free
Artificial Intelligence • Insurance • Software • Automation
The Staff Site Reliability Engineer will build and scale infrastructure for Assured's platform, automate delivery, enhance observability, and lead mentoring initiatives.
Top Skills:
AWSKubernetesPostgresTerraform
Cloud • Security • Software • Cybersecurity
Lead reliability for a serverless AI inference platform: own observability and SLO/SLI frameworks, build automation and tooling, manage incidents and on-call, define deployment safety (canaries, rollbacks), influence architecture with product teams, and mentor other SREs.
Top Skills:
AutoscalingCi/CdContainer OrchestrationContainerizationGoGpu WorkloadsInfrastructure-As-CodeKubernetesModel ServingPythonResource Scheduling
Information Technology • Legal Tech • Analytics
Partner with DBAs and SRE teams to ensure reliable, secure, and scalable database infrastructure. Lead incident management, vulnerability remediation, DR planning and resilience testing, infrastructure provisioning across VMs and Chainguard containers, cost optimization, and SRE capability building. Provide on-call support and coach junior staff.
Top Skills:
Amazon Web Services (Aws)ChainguardGitlabGrafanaIaasJenkinsLinuxMySQLPmm3PostgresPrometheusTerraformVm
Artificial Intelligence • Healthtech • Software • Telehealth
Design, deploy, and maintain AWS-hosted Kubernetes (EKS) infrastructure; build automation and AI-assisted runbooks; provide observability and incident response; ensure HIPAA-compliant, high-availability platform operations and mentor engineering teams.
Top Skills:
AWSBashDatadogEc2EksGitGithub ActionsGoHelmKubernetesPythonRdsS3Terraform
Software • Financial Services
Own day-to-day AWS and database operations for a serverless production platform, manage backups and disaster recovery, monitor and debug production, lead incident response and on-call, maintain infrastructure-as-code (SST/Pulumi), optimize costs, and mentor the team on operational best practices.
Top Skills:
AWSBashCloudwatchDnsDockerEventbridgeGithub ActionsLambdaLinuxPostgresPulumiPythonS3SqsSstTerraformTlsTypescript
Other
The Senior Site Reliability Engineer at Juul Labs ensures operational stability and performance of hybrid cloud infrastructure, leads automation, and handles critical incidents.
Top Skills:
AWSBashCloudFormationGCPNutanixPowershellPythonTerraform
Healthtech
Design, scale, and operate secure AWS cloud infrastructure (EKS, IAM, RBAC); build and maintain IaC (Terraform/Terragrunt), GitHub Actions CI/CD, Datadog observability, and Python automation; document runbooks, participate in on-call rotations, postmortems, and Agile workflows to improve reliability and security.
Top Skills:
AWSDatadogEc2EksFargateGithub ActionsGithub Advanced SecurityHelmIamJIRAKubernetesLambdaPythonRbacSecrets ManagerServerlessTerraformTerragruntVpc
Artificial Intelligence
Own operational excellence for cloud infrastructure: run incident management, improve reliability through automation, own a platform domain (e.g., Kubernetes, Temporal, observability), manage vendor and cost relationships, and deliver measurable reductions in incidents and costs within 12 months.
Top Skills:
AWSKubernetesLlm ApisMongoDBObservabilityPythonTemporal
Digital Media • Edtech
Drive reliability and observability of Epic's GCP-based platform. Own cloud infrastructure, container platform (Kubernetes/GKE), CI/CD, observability, and IaC (Terraform). Define SLOs/SLIs, reduce toil, manage security and compliance practices, participate in on-call rotations, lead incident response and post-mortems, and partner with product and data teams to troubleshoot and improve platform reliability.
Top Skills:
ArgocdBashCloud MonitoringDockerGceGCPGcsGithub ActionsHelmIamJenkinsKubernetes (Gke)New RelicPythonTerraformVpc
Logistics • Software
Own and operate scalable infrastructure on GCP (GKE, Cloud Run, AlloyDB); author Terraform modules; manage containerized workloads; build observability in Datadog; design CI/CD in GitHub Actions; automate operational workflows; lead incident response and post-mortems; partner with engineers to improve reliability, cost, and automation.
Top Skills:
AlloydbClickhouseCloud RunCloudflare WorkersDatadogDockerGCPGithub ActionsGkeGoGrafanaIamKafkaKubernetesNetworkingPostgresPrometheusPub/SubPythonRedisRedpandaTerraformTypescript
Artificial Intelligence • Cloud • Fintech • Machine Learning • Mobile • Software
Lead design, development, deployment, and scaling of cloud infrastructure and SRE tooling. Build automation, CI/CD, observability, capacity planning, and reliability improvements; collaborate with product teams to define non-functional requirements and resolve production issues.
Top Skills:
.NetApi GatewayAWSAzureC#Data LakehouseDatabricks DeltaDatadogElasticsearchElkEvent HubsFunctions/ServerlessGitGrafanaJavaJenkinsKafkaKibanaKubernetesLogstashPowershellSnowflakeSqsTeamcityVisual Basic
Digital Media • Social Media • Software • Sports
Lead the technical architecture and execution of migration to AWS, drive developer enablement, and automate infrastructure using code-first principles.
Top Skills:
Aws EksDatadogGithub ActionsGoIstioK6KubernetesNode.jsTerraform
Legal Tech • Software
Lead Site Reliability Engineer responsible for platform availability and reliability of RelativityOne. Drive SRE best practices, build tools, lead projects, coach SREs, work with stakeholders, support incidents, run postmortems, and improve monitoring, automation, and operational efficiency.
Top Skills:
Ci/CdDevOpsJenkinsJIRAKubernetesAzureMonitoring And AlertingNew RelicNoSQLPowershellRelativity ServerRelativityoneSQLTableau
Let Your Resume Do The Work
Upload your resume to be matched with jobs you're a great fit for.
Success! We'll use this to further personalize your experience.
Top Chicago, IL Companies Hiring Remote Senior Site Reliability Engineers
See AllPopular Job Searches
All Filters
Total selected ()
No Results
No Results




















.png)










