CentralReach Logo

CentralReach

Sr. Site Reliability Engineer

Posted 19 Days Ago
Be an Early Applicant
Remote or Hybrid
Hiring Remotely in US
160K-180K Annually
Senior level
Remote or Hybrid
Hiring Remotely in US
160K-180K Annually
Senior level
Own production reliability across cloud environments by defining SLOs, SLIs, error budgets, observability practices, and actionable dashboards. Troubleshoot incidents, perform root cause analysis, manage capacity planning, and improve service performance and availability. Build automation to reduce toil, support release management, maintain runbooks, and collaborate with engineering teams on operational readiness. Manage tools including Datadog, Prometheus, and Grafana while advancing monitoring, logging, metrics, tracing, and reliability practices.
The summary above was generated by AI

CentralReach is a leading provider of autism and IDD care software for Applied Behavior Analysis (ABA), multidisciplinary therapy, and special education. Trusted by more than 200,000 users, we enable therapy providers, educators, and employers to scale the way they deliver ABA and related therapies with innovative technology, market-leading industry expertise, and world-class customer satisfaction. 

The Platform Engineering group at CentralReach builds the underlying technologies that power our Public and Private Cloud Platforms worldwide. The group is responsible for storage, data infrastructure, IT, observability systems, DevOps, SRE, provisioning, compute, orchestration platform, internal tools, internal platforms (laptops, networks, systems etc.) and services - all the components that make up the CentralReach Platform. 

If you have a passion for the future, enjoy and thrive in an agile, fast-moving, ever-changing startup environment, welcome and take on technical challenges of all shapes and sizes, have excellent interpersonal skill and sense of humor and enjoy rolling up your sleeves and jumping in, then read on!  

As a Sr. SRE, you will work closely with the key stakeholders in Software Engineering to drive adoption of modern reliability practices like SLOs, error budget policies, actionable alerts, incident retrospectives, chaos testing, and end-to-end ownership. 

Key Accountabilities: 

  • Own production reliability, including availability, latency, performance, capacity planning, monitoring, emergency response, and uptime for production environments. 
  • Define, maintain, and improve SLOs, SLIs, error budgets, actionable dashboards, and observability practices. 
  • Analyze, troubleshoot, and resolve operational issues that affect service reliability and SLO performance. 
  • Build and automate multi-environment observability capabilities, including capacity forecasting based on usage patterns. 
  • Reduce toil and increase development velocity through automation and continuous improvement. 
  • Provide production support, including incident, change, and problem management; root cause analysis; service restoration; runbooks; and standard operating procedures. 
  • Identify data-driven opportunities to improve system architecture, availability, performance, and reliability. 
  • Collaborate with software engineering teams on release management, roadmap planning, and operational readiness. 
  • Implement and manage reliability and observability tools such as Datadog, Prometheus, and Grafana. 

 

Desired Skills and Experience: 

  • Experience with monitoring, APM, and observability tools such as Splunk, Prometheus, Datadog, and OpenTelemetry. 
  • Experience implementing observability strategies for logs, metrics, and traces. 
  • Strong understanding of CI/CD practices and tools such as Jenkins, GitHub Actions, GitLab, Argo, and Kargo. 
  • Strong understanding of major cloud providers, preferably AWS, and cloud-native infrastructure concepts. 
  • Strong understanding of containerization technologies, including Kubernetes and Helm. 
  • Experience with one or more programming languages, such as Java, Python, or Go, and familiarity with .NET application development. 
  • Strong understanding of Linux, Windows, software development, systems, networking, and cloud concepts. 
  • Experience using AI to improve productivity and amplify technical skills. 
Base Salary Range
$160,000—$180,000 USD

Backed by Roper Technologies, Inc. (Nasdaq: ROP), CentralReach is entering an exciting phase of growth, innovation, and scale.  

Recognized as one of the best places to work over 10 times by organizations such as Inc, Built In, and NJBIZ, our culture is centered around impact, inclusion, and flexibility. As a hybrid company with collaborative offices in Ft. Lauderdale, FL; Holmdel, NJ; and Verona, Italy, we foster a workplace where top talent can thrive and make a real difference in the lives of those we serve.
We offer competitive compensation, comprehensive health benefits, generous PTO, 401(k) matching, and paid parental leave to our full-time employees. Our team members also enjoy hybrid work schedules, career development support, wellness programs, and opportunities to give back through CR Cares™, our community engagement initiative.

Be part of a market leader driving the future of care. Explore opportunities at centralreach.com/careers.  

Protecting your information is important to us.  Please take a moment to review our Notice of Privacy Practices for Job Applicants. Applicant-Privacy-Notice-110725 to understand how we collect, use, store, and protect your personal information during the recruitment process. 

Similar Jobs

6 Days Ago
In-Office or Remote
75K-195K Annually
Senior level
75K-195K Annually
Senior level
Cloud • Software
Own NetBox’s build and release pipeline from image creation through Cloud and Enterprise deployment. Improve Django and PostgreSQL performance, establish observability and SLOs, strengthen software supply chain security, and support SOC 2 compliance. Participate in on-call rotations, incident response, and postmortems. Drive cross-team migrations and release processes while contributing fixes to NetBox Core when reliability issues originate in the application.
Top Skills: ArgocdAws Ec2Aws IamAws RdsAws VpcClaude CodeCosignDjangoFluxcdGithub ActionsGrafanaHelmKubernetesPostgresPrometheusPythonSigstoreSlsaTerraform
10 Days Ago
In-Office or Remote
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Designs and operates secure, reliable Azure cloud platforms using Terraform, GitHub Actions, containers, and automation. Responsibilities include CI/CD, observability, incident response, platform security, vulnerability remediation, disaster recovery, infrastructure troubleshooting, and SRE practices. The role supports production workloads, improves reliability and delivery processes, participates in on-call activities, and mentors engineers while partnering across development, security, architecture, and operations teams.
Top Skills: BashCi/CdCloud SecurityDockerGitGithub ActionsGitopsInfrastructure As CodeKubernetesAzureObservabilityPowershellPythonTerraform
18 Days Ago
Remote or Hybrid
United States
Senior level
Senior level
Fintech • Software
The Senior Site Reliability Engineer ensures SaaS platforms remain reliable, performant, secure, and scalable. Responsibilities include building cloud infrastructure, implementing monitoring and alerting, automating operational runbooks and deployments, managing Infrastructure as Code, applying AI-powered observability and remediation, supporting Kubernetes and cloud networking, and leading incident triage and root-cause analysis during 24/7 on-call rotations.
Top Skills: AIAiopsAksAnsibleAppdynamicsAWSAzureAzure DevopsBashC# .NetCi/CdCloud NetworkingCloudopsCosmos DbDatadogDynatraceEksFirewallsHarnessIdera Sql Diagnostic ManagerInfrastructure As CodeJavaJenkinsKubernetesLinuxLoad BalancingNew RelicPowershellPythonRedgate Sql MonitorSolarwinds Database Performance AnalyzerSQLTerraformWindows

What you need to know about the Chicago Tech Scene

With vibrant neighborhoods, great food and more affordable housing than either coast, Chicago might be the most liveable major tech hub. It is the birthplace of modern commodities and futures trading, a national hub for logistics and commerce, and home to the American Medical Association and the American Bar Association. This diverse blend of industry influences has helped Chicago emerge as a major player in verticals like fintech, biotechnology, legal tech, e-commerce and logistics technology. It’s also a major hiring center for tech companies on both coasts.

Key Facts About Chicago Tech

  • Number of Tech Workers: 245,800; 5.2% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: McDonald’s, John Deere, Boeing, Morningstar
  • Key Industries: Artificial intelligence, biotechnology, fintech, software, logistics technology
  • Funding Landscape: $2.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Pritzker Group Venture Capital, Arch Venture Partners, MATH Venture Partners, Jump Capital, Hyde Park Venture Partners
  • Research Centers and Universities: Northwestern University, University of Chicago, University of Illinois Urbana-Champaign, Illinois Institute of Technology, Argonne National Laboratory, Fermi National Accelerator Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account