Optum Logo

Optum

Principal Site Reliability Engineer - Remote

Posted An Hour Ago
Be an Early Applicant
In-Office or Remote
Hiring Remotely in Eden Prairie, MN
135K-231K Annually
Expert/Leader
In-Office or Remote
Hiring Remotely in Eden Prairie, MN
135K-231K Annually
Expert/Leader
Leads AI-assisted site reliability engineering across Azure and AWS. Designs observability, incident response, automation, resiliency testing, disaster recovery, chaos engineering, and recovery-validation capabilities. Establishes OpenTelemetry, SLI, SLO, error-budget, and reliability-scorecard standards; improves alert quality and operational insights; creates human-in-the-loop mitigation workflows; and mentors engineers while driving cross-functional reliability improvements.
The summary above was generated by AI
Requisition Number: 2371367
Optum Tech is a global leader in health care innovation. Our teams develop cutting-edge solutions that help people live healthier lives and help make the health system work better for everyone. From advanced data analytics and AI to cybersecurity, we use innovative approaches to solve some of health care's most complex challenges. Your contributions here have the potential to change lives. Ready to build the next breakthrough? Join us to start Caring. Connecting. Growing together.
Are you passionate about reimagining operations through AI? Optum Financial is seeking a Principal Site Reliability Engineer to lead the evolution of our reliability platform by combining modern SRE practices with AI-assisted operations. You'll design systems that help engineers detect issues faster, automate response workflows, improve resiliency, and transform operational data into actionable insights. As a technical leader, you'll influence reliability standards, mentor engineers, and drive next-generation observability and automation strategies across Azure and AWS environments.
You'll enjoy the flexibility to work remotely * from anywhere within the U.S. as you take on some tough challenges.
For all hires in the Minneapolis or Washington, D.C. area, you will be required to work in the office a minimum of four days per week.
Primary Responsibilities:
  • Build AI-assisted SRE capabilities that accelerate incident detection, triage, mitigation, and recovery
  • Connect observability, deployment, runbook, ownership, and incident data into actionable operational context
  • Design human-in-the-loop workflows for safe mitigation, approvals, recovery verification, and auditability
  • Standardize OpenTelemetry, SLIs, SLOs, error budgets, and reliability scorecards across critical services
  • Improve alert quality by reducing noise, clarifying customer impact, and identifying likely causes faster
  • Lead resiliency testing, DR exercises, chaos engineering, and automated recovery validation
  • Mentor engineers and lead cross-functional reliability improvements across Optum Financial

You'll be rewarded and recognized for your performance in an environment that will challenge you and give you clear direction on what it takes to succeed in your role as well as provide development for other roles you may be interested in.
Required Qualifications:
  • 10+ years of experience in software engineering, platform engineering, DevOps, or SRE roles
  • 3+ years of experience in a principal, staff, lead, or senior technical leadership role
  • 5+ years of experience with cloud platforms and container orchestration, preferably Azure or AWS
  • 3+ years of experience with observability tools such as OpenTelemetry, Prometheus, Grafana, Datadog, or similar platforms
  • 1+ years of experience designing production automation, tooling, or AI-assisted workflows for incident response or operational decision-making

Preferred Qualifications:
  • Bachelor's degree in Computer Science, Information Technology, Engineering, or related field
  • Experience with LLM-based systems, AI agents, RAG, tool orchestration, evaluations, or guardrails
  • Experience with resiliency engineering, disaster recovery, chaos engineering, or recovery validation
  • Experience with infrastructure as code and automation tools such as Terraform, Pulumi, Ansible, Helm, or Kubernetes operators
  • Solid background in incident command, runbooks, postmortems, production readiness, and reliability governance

*All employees working remotely will be required to adhere to UnitedHealth Group's Telecommuter Policy
Pay is based on several factors including but not limited to local labor markets, education, work experience, certifications, etc. In addition to your salary, we offer benefits such as, a comprehensive benefits package, incentive and recognition programs, equity stock purchase and 401k contribution (all benefits are subject to eligibility requirements). No matter where or when you begin a career with us, you'll find a far-reaching choice of benefits and incentives. The salary for this role will range from $134,600 - $230,800 annually based on full-time employment. We comply with all minimum wage laws as applicable.
Application Deadline: This will be posted for a minimum of 2 business days or until a sufficient candidate pool has been collected. Job posting may come down early due to volume of applicants.
At UnitedHealth Group, our mission is to help people live healthier lives and make the health system work better for everyone. We believe everyone-of every race, gender, sexuality, age, location and income-deserves the opportunity to live their healthiest life. Today, however, there are still far too many barriers to good health which are disproportionately experienced by people of color, historically marginalized groups and those with lower incomes. We are committed to mitigating our impact on the environment and enabling and delivering equitable care that addresses health disparities and improves health outcomes - an enterprise priority reflected in our mission.
UnitedHealth Group is an Equal Employment Opportunity employer under applicable law and qualified applicants will receive consideration for employment without regard to race, national origin, religion, age, color, sex, sexual orientation, gender identity, disability, or protected veteran status, or any other characteristic protected by local, state, or federal laws, rules, or regulations.
UnitedHealth Group is a drug - free workplace. Candidates are required to pass a drug test before beginning employment.

Similar Jobs at Optum

19 Days Ago
In-Office or Remote
Expert/Leader
Expert/Leader
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Define and scale SRE standards across teams, implement SLOs/SLIs/error budgets, build observability and resiliency patterns, drive automation and AIOps, improve reliability for large-scale Azure cloud systems, and influence engineering and platform teams.
Top Skills: Ai/MlAiopsAutomationAzureError BudgetsIncident ManagementLogsObservability (MetricsOpentelemetrySlisSlosTracing)
An Hour Ago
In-Office or Remote
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Develop and deploy AI and machine learning solutions, analyze complex enterprise datasets, build predictive and optimization models, conduct statistical experiments, and translate findings into business recommendations. The role includes applying deep learning, NLP, generative AI, and ethical AI practices; creating scalable decision-support frameworks; evaluating emerging technologies; collaborating across teams; and mentoring junior data scientists and analysts.
Top Skills: Agentic AiCloud Analytics PlatformsDeep LearningGenerative AiLarge Language ModelsMachine LearningNatural Language ProcessingPythonPyTorchRScikit-LearnSQLTensorFlowXgboost
An Hour Ago
Remote or Hybrid
92K-164K Annually
Senior level
92K-164K Annually
Senior level
Artificial Intelligence • Big Data • Healthtech • Information Technology • Machine Learning • Software • Analytics
Lead release engineering for Java applications on Azure, designing CI/CD pipelines with GitHub Actions and Azure DevOps. Manage Azure infrastructure using Terraform, including AKS, networking, storage, identity, and monitoring. Build AI agents and automation for release engineering, incident triage, and operational workflows. Implement blue/green, canary, and rolling deployments; troubleshoot application, pipeline, infrastructure, and automation issues; and collaborate with development, SRE, security, and platform teams on governance, reliability, and reusable automation.
Top Skills: Agent OrchestrationAi AgentsArtifact RepositoriesAzure App ServicesAzure DevopsAzure Kubernetes Service (Aks)DockerGitGithub ActionsGradleJavaKubernetesMavenMicroservicesAzureSpring BootTerraform

What you need to know about the Chicago Tech Scene

With vibrant neighborhoods, great food and more affordable housing than either coast, Chicago might be the most liveable major tech hub. It is the birthplace of modern commodities and futures trading, a national hub for logistics and commerce, and home to the American Medical Association and the American Bar Association. This diverse blend of industry influences has helped Chicago emerge as a major player in verticals like fintech, biotechnology, legal tech, e-commerce and logistics technology. It’s also a major hiring center for tech companies on both coasts.

Key Facts About Chicago Tech

  • Number of Tech Workers: 245,800; 5.2% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: McDonald’s, John Deere, Boeing, Morningstar
  • Key Industries: Artificial intelligence, biotechnology, fintech, software, logistics technology
  • Funding Landscape: $2.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Pritzker Group Venture Capital, Arch Venture Partners, MATH Venture Partners, Jump Capital, Hyde Park Venture Partners
  • Research Centers and Universities: Northwestern University, University of Chicago, University of Illinois Urbana-Champaign, Illinois Institute of Technology, Argonne National Laboratory, Fermi National Accelerator Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account