Delinea Logo

Delinea

Manager, Site Reliability Engineering

Posted 2 Days Ago
Be an Early Applicant
Remote
Hiring Remotely in U.S.
160K-180K Annually
Senior level
Remote
Hiring Remotely in U.S.
160K-180K Annually
Senior level
Lead a hands-on Site Reliability Engineering and DevOps team supporting production SaaS environments across Azure and AWS. Own availability, observability, automation, incident response, on-call operations, disaster recovery, and reliability improvements. Manage employees and contractors, coordinate Sev1 and Sev2 incidents, improve monitoring and deployment practices, report operational health, and expand support for FedRAMP High and other regulated environments.
The summary above was generated by AI

About Delinea:
Delinea is a pioneer in securing human and machine identities through intelligent, centralized authorization, empowering organizations to seamlessly govern their interactions across the modern enterprise. Leveraging AI-powered intelligence, Delinea’s leading cloud-native Identity Security Platform applies context throughout the entire identity lifecycle – across cloud and traditional infrastructure, data, SaaS applications, and AI. It is the only platform that enables you to discover all identities – including workforce, IT administrator, developers, and machines – assign appropriate access levels, detect irregularities, and respond to threats in real-time. With deployment in weeks, not months, 90% fewer resources to manage than the nearest competitor, and a 99.995% uptime, Delinea delivers robust security and operational efficiency without compromise. Learn more about Delinea on Delinea.com, LinkedIn, X, and YouTube.

Join our passionate, global team at Delinea and help us make the world a safer and more secure place. Our success is driven by world-class product leadership, outstanding engineers, and strategic investment from TPG. We value diversity, innovation, and a culture of respect and fairness. If you're ready to push boundaries and challenge the status quo in security, we want to hear from you.
 

Apply today to help us achieve our mission.

Summary:

Delinea is looking for a hands-on Manager of Site Reliability Engineering to lead the SRE and DevOps engineers supporting the Delinea products. This is a working manager role. You will be expected to lead people and lead work: writing and reviewing automation, digging into AKS and Azure telemetry, commanding Sev1 and Sev2 incidents, and improving the observability and deployment practices your team depends on.

The initial scope is the Platform SRE and DevOps team. Over time, the role is expected to expand to cover our FedRAMP High environment and the broader Commercial Platform footprint, so comfort operating in a regulated environment and a willingness to grow scope are essential.

You will lead a blended team of full-time engineers and contractors distributed across multiple time zones. Meeting your team inside their working hours is an expectation of the role. On-call participation is required.

What You Will Do:

  • Lead hands-on. Spend a meaningful portion of your week in the environment: reviewing pull requests, validating pipeline changes, tuning monitors and dashboards, running queries in Datadog, and troubleshooting production issues alongside your engineers. This role does not sit above the work.

  • Own availability and performance of the Delinea Platform production environments across Azure and AWS, including AKS workloads, ingress and networking, data services, messaging, and CDN or WAF layers.

  • Manage a blended team. Hire, onboard, coach, and develop full-time SRE engineers. Direct and manage contractor resources, including scoping work, setting quality expectations, and reviewing deliverables.

  • Lead a distributed team. Run one-on-ones, standups, and planning sessions at times that work for engineers in other geographies.

  • Participate in on-call. Carry the pager as part of the rotation, act as incident commander for Sev1 and Sev2 events, drive engagement of the right responders, and own communication cadence with support, engineering, and leadership until resolution.

  • Command incident response end to end. Own detection, triage, mitigation, customer-facing status communication, and post-incident review. Ensure RCAs are written to a customer-ready standard, preventative actions have owners and target dates, and those actions are driven to closure.

  • Raise the observability bar. Improve detection coverage so that issues are found by our monitoring rather than by a customer ticket. Own SLI and SLO definition, alert quality and noise reduction, synthetic coverage, APM instrumentation, log hygiene, and dashboard standards.

  • Support FedRAMP and regulated operations. Grow into supporting our FedRAMP High environment, including change control discipline, evidence collection, boundary awareness, and the operational differences between government and commercial environments.

  • Reduce toil through automation. Set the expectation that repeat manual work becomes code. Prioritize automation backlog alongside project and reliability work.

  • Report on operational health. Produce and present incident metrics, trends, and reliability commitments to leadership, and translate them into a concrete improvement plan.

What You Will Need:

  • 6+ years in Site Reliability Engineering, DevOps, or Cloud Operations, with demonstrated ownership of production SaaS systems.

  • 2+ years of direct people leadership, including performance management, hiring, and coaching. Experience managing contractors or an outsourced delivery team is desired.

  • Current, hands-on production experience with the Delinea technology stack, including Azure Kubernetes Service, core Azure services (SQL, Redis, Service Bus, Blob Storage), AWS services (SES, EC2, RDS), WAF, Azure DevOps pipelines, Datadog, and Atlassian Jira Service Management.

  • Hands-on experience across both Azure and AWS is required. You should be able to administer, troubleshoot, and reason about cost and security posture in each.

  • Deep observability expertise. Demonstrated ownership of an observability framework at scale: metrics, logs, traces, synthetics, SLOs, and alerting strategy. Hands-on proficiency with Datadog or an equivalent platform, including APM trace analysis and log-based troubleshooting.

  • Proven incident command. You have run major incidents as the incident commander, coordinated multiple responders under pressure, communicated to customers and executives during active impact, and authored the RCA afterward.

  • Strong cloud networking and security fundamentals: load balancing, DNS, TLS and certificate lifecycle, firewalls, VPN, routing, and identity and access management.

  • Automation and scripting ability in PowerShell, Python, Bash, or similar, plus practical infrastructure-as-code experience (Terraform, ARM, or Bicep).

  • Practical experience with multi-region, multi-tenant SaaS architectures, including backup, redundancy, and disaster recovery approaches.

  • Excellent written communication. You will write and approve customer-facing status updates and incident summaries under time pressure.

  • Willingness and availability to work across time zones and to participate in an on-call rotation.

We Would Love to See

  • Direct experience operating in a FedRAMP or other regulated environment (Azure Government, IL4/IL5, SOC 2, ISO 27001).

  • Experience standing up or maturing an incident management program, including sev definitions, escalation paths, on-call structure, and post-incident review process.

  • Experience with public status page operations and customer notification practices.

  • Experience with Atlassian Jira Service Management, Confluence, and Azure DevOps as the operational toolchain.

  • Track record of reducing customer-detected incidents through improved monitoring coverage.

  • Cost optimization experience across Azure and AWS on a meaningful scale.

For this Job, Delinea is not considering candidates that need any type of US work authorization now or in the future. This includes, but is not limited to: F1-OPT, F1-CPT, H-1B, TN, L-1, J1, etc.

Why work at Delinea?

  • We're passionate problem-solvers helping the world's largest organizations protect what matters most: their human and machine identities.

  • We invest in people who are smart, self-motivated, and collaborative.

  • What we offer in return is meaningful work, a culture of innovation and great career progression.

At Delinea, our core values are STRONG and guide our behaviors and success:

  • Spirited - We bring energy and passion to everything we do

  • Trust - We act with integrity and deliver on our commitments

  • Respect - We listen, value different perspectives, and work as one team

  • Ownership - We take initiative and follow through

  • Nimble - We adapt quickly in a fast-changing environment

  • Global - We embrace diverse people and ideas to drive better outcomes

We believe weaving these core values into our day-to-day actions, and our process for hiring, evaluating, and promoting employees, helps us cultivate a work environment that embraces collaboration and camaraderie.

We take care of our employees. We offer competitive salaries, a meaningful bonus program, and excellent benefits, including healthcare insurance, as well as pension/retirement matching, comprehensive life insurance, an employee assistance program, time off plans, and paid company holidays.

Delinea is an Equal Opportunity and Affirmative Action employer and prohibits discrimination and harassment of any type with regard to race, color, religion, age, sex, national origin, disability status, genetics, protected veteran status, sexual orientation, gender identity or expression, or any other characteristic protected by federal, state or local laws.

Upon conditional offer of employment, candidates are required to complete comprehensive criminal background check, verification of education, and verification of employment, per employment policy. In addition, all publicly posted social media sites may be reviewed.

 

 

 

 


Similar Jobs

5 Days Ago
Remote
United States
150K-170K Annually
Senior level
150K-170K Annually
Senior level
Insurance
Leads and develops a DevOps engineering team responsible for CI/CD, infrastructure automation, cloud enablement, release engineering, developer tooling, and operational support. Drives AWS modernization, DevSecOps, Infrastructure as Code, observability, secure deployment standards, and continuous improvement. Partners with engineering, security, infrastructure, risk, and audit teams to improve reliability, compliance, release safety, and developer experience. Provides hands-on technical guidance, manages service quality and team execution, and participates in incident reviews, change governance, and operational improvement.
Top Skills: AgileAutomated TestingAWSCi/CdContainerizationDevsecopsInfrastructure As CodeItilKanbanObservabilitySecrets Management
One Month Ago
In-Office or Remote
Senior level
Senior level
Mobile • Software
Lead and grow a distributed SRE team to ensure Radar's production infrastructure is highly available, scalable, and automated. Own observability, incident management, multi-region Kubernetes deployments on AWS (EKS), Terraform infrastructure, and MongoDB sharded clusters. Partner with product, platform, security and data teams, participate in on-call rotation, drive reliability and cost-saving initiatives, and incorporate customer feedback into infrastructure work.
Top Skills: AirflowSparkAWSCircleCICloudflareCloudwatchEksGrafanaKubernetesMongodb AtlasPagerdutyPingdomRustScalaTerraformTypescript
9 Days Ago
Remote
United States
143K-304K Annually
Senior level
143K-304K Annually
Senior level
Software • Quantum Computing • Metaverse • Infrastructure as a Service (IaaS)
Leads Site Reliability Engineering teams responsible for reliability, availability, incident response, disaster recovery, automation, telemetry, security, and compliance for Microsoft Substrate services in regulated environments. Establishes SLOs and SLIs, drives operational excellence, participates in on-call rotations, leads post-incident improvements, develops senior engineers, and influences cross-organizational engineering strategy.
Top Skills: Ai-Assisted EngineeringAutomationCloud ComputingDisaster RecoveryDistributed SystemsMicrosoft CloudMicrosoft SubstrateSite Reliability EngineeringSlisSlosTelemetry

What you need to know about the Chicago Tech Scene

With vibrant neighborhoods, great food and more affordable housing than either coast, Chicago might be the most liveable major tech hub. It is the birthplace of modern commodities and futures trading, a national hub for logistics and commerce, and home to the American Medical Association and the American Bar Association. This diverse blend of industry influences has helped Chicago emerge as a major player in verticals like fintech, biotechnology, legal tech, e-commerce and logistics technology. It’s also a major hiring center for tech companies on both coasts.

Key Facts About Chicago Tech

  • Number of Tech Workers: 245,800; 5.2% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: McDonald’s, John Deere, Boeing, Morningstar
  • Key Industries: Artificial intelligence, biotechnology, fintech, software, logistics technology
  • Funding Landscape: $2.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Pritzker Group Venture Capital, Arch Venture Partners, MATH Venture Partners, Jump Capital, Hyde Park Venture Partners
  • Research Centers and Universities: Northwestern University, University of Chicago, University of Illinois Urbana-Champaign, Illinois Institute of Technology, Argonne National Laboratory, Fermi National Accelerator Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account