Relativity Logo

Relativity

Product Manager - AI Evaluations Tooling

Posted 6 Days Ago
Remote
Hiring Remotely in Illinois, USA
115K-173K Annually
Mid level
Remote
Hiring Remotely in Illinois, USA
115K-173K Annually
Mid level
Own the vision, strategy, and roadmap for an AI evaluations platform supporting offline and online evaluation, production monitoring, agentic testing, SME workflows, and quality dashboards. Partner with applied scientists, engineers, product teams, legal experts, and executives to define evaluation standards, calibrate LLM judges, assess build-versus-buy decisions, and help teams measure and improve AI performance.
The summary above was generated by AI

Posting Type

Remote/Hybrid

Job Overview

Relativity's AI portfolio spans 20+ products and we are moving quickly into agentic capabilities. As that portfolio grows, so does the opportunity to give every team building AI at Relativity a shared, standardized way to understand and improve their own AI performance. The Evaluations Tooling team exists to build that foundation: a platform that standardizes how we evaluate AI, accelerates how quickly teams can ship it with confidence, and turns AI quality into shared evidence that applied scientists, application teams, product managers, and legal experts can all act on.
We are looking for a Product Manager to own the product strategy and roadmap for Relativity's evaluation platform. Your mission is to enable every team at Relativity building AI products to understand, measure, and improve their own AI performance. That means productizing what applied scientists have already validated into a platform that application teams, PMs, and legal subject-matter experts can use directly; building a world-class agentic testing system that lets us ship new agentic capabilities quickly and confidently; and using production monitoring and team-facing dashboards to move the organization from reacting to quality problems to proactively getting ahead of them.
This is a platform PM role with unusually direct leverage. A successful candidate is energized by deeply technical customers, comfortable making build-vs-buy calls in a fast-moving vendor landscape, and able to hold a categorical bar ("every AI product ships behind elite evals") while sequencing pragmatically to get there.

Job Description and Requirements

Role Responsibilities  

  • Own the product vision, strategy, and roadmap for the Evaluations Platform across its pillars: offline evaluation pipeline, online evaluation and production monitoring, eval discovery and reuse, SME authoring and approval, and agentic evaluation.  
  • Serve every team building AI at Relativity. Run continuous discovery with applied scientists, application engineering teams, and product PMs to understand how each of them experiences AI quality today and design the platform so teams can measure and improve their own performance without routing every question through an applied scientist.  
  • Lead the orchestration of a world-class agentic testing system: trajectory evaluation, tool-call correctness, intermediate-state rubrics, long-running judge orchestration, and simulation of multi-step flows. Make it fast and safe to ship new agentic capabilities and make the platform a competitive advantage in how quickly Relativity can iterate on agents.  
  • Work closely with legal subject-matter experts to define the tasks our AI is evaluated against, what a correct outcome looks like, where the hard cases are, and how judgment criteria should be expressed so they can be encoded, scored, and reused across products. Build the SME authoring and approval workflow around how experts work.  
  • Own the team-facing quality dashboards. Every AI product team should be able to open a view and understand their own performance, where they stand against their quality bar, and what changed since last release, with no engineering help required. 
  • Lead build-vs-buy evaluations for tracing infrastructure and commercial eval tooling and run a recurring review as the vendor landscape shifts.  
  • Partner with engineering and applied science on judge calibration so LLM-as-judge is trustworthy enough to gate deployments. 
  • Clearly articulate technical tradeoffs to stakeholders ranging from applied scientists to attorneys to executive leadership. 

 Preferred qualifications:

  • Experience with LLM-based products and a working understanding of how they are evaluated (rubrics, LLM-as-judge, offline vs. online evaluation). Direct experience with agentic systems is a plus, not a requirement.  
  • Experience as a platform, developer-tools, or data/ML PM serving technical customers, or a demonstrated ability to learn a technical domain quickly and earn credibility with engineers and scientists.  
  • Comfort bringing non-engineering domain experts into technical workflows and translating their judgment into something a system can act on.  
  • Strong analytical instincts: sets metrics, tests assumptions, and changes course when the evidence says to. 

 

Minimum qualifications:

  • 4+ years of product management experience at a technology company, with at least 2 years on technical, platform, infrastructure, or data/ML products. 
  • Solid understanding of the software development lifecycle and agile practices. 
  • Experience conducting user discovery with technical audiences and translating findings into roadmap decisions. 
  • Excellent written and verbal communication skills, with the ability to explain complex technical concepts to both technical and non-technical audiences. 

. 

Relativity is committed to competitive, fair, and equitable compensation practices.

This position is eligible for total compensation which includes a competitive base salary, an annual performance bonus, and long-term incentives.

The expected salary range for this role is between following values:

$115,000 and $173,000

The final offered salary will be based on several factors, including but not limited to the candidate's depth of experience, skill set, qualifications, and internal pay equity. Hiring at the top end of the range would not be typical, to allow for future meaningful salary growth in this position. 

Required Skills:

Agile Methodology, Cross-Functional Teamwork, Market Research, Market Strategy, Product Lifecycle, Product Lifecycle Management (PLM), Product Management, Product Strategies, Stakeholder Management, Team Leadership
HQ

Relativity Chicago, Illinois, USA Office

We’re a community of passionate, life-long learners tackling challenging problems. We care about each other and about our community.

Similar Jobs

2 Minutes Ago
Remote
USA
Mid level
Mid level
Insurance • Financial Services
Supports key insurance accounts by addressing production, recruiting, onboarding, appointment, activation, and portal issues. Tracks and communicates performance metrics, coaches agents on sales fundamentals, mentors recruiting leaders, and delivers webinars and in-person seminars. Partners with licensing, operations, technology, and compliance teams to resolve field issues, document solutions, and improve agent readiness and production. This is a remote role requiring travel for training, seminars, and meetings.
Top Skills: Agent PortalsCRMExcelMS OfficeOutlookPowerPointReporting Tools
22 Minutes Ago
Remote
Illinois, USA
17-17 Hourly
Mid level
17-17 Hourly
Mid level
Fintech • Real Estate • Sales • Financial Services
Originates mortgage loans by building referral networks, generating business, analyzing clients’ financial information, recommending suitable loan products, and guiding borrowers through application to closing. The role requires mortgage origination experience, NMLS and state licenses, strong communication, independent judgment, and client-service skills. It also involves community outreach, relationship management, lead generation, resolving client concerns, and staying current on mortgage products and qualification requirements.
34 Minutes Ago
Easy Apply
Remote
United States
Easy Apply
146K-225K Annually
Entry level
146K-225K Annually
Entry level
Big Data • Fintech • Mobile • Payments • Financial Services
Design and implement backend APIs, microservices, and checkout platform components for Affirm’s financial products. Own system health, improve performance and reliability, collaborate across engineering teams, and contribute to technical standards, tooling, and processes. The role requires developing, testing, and shipping high-quality software at scale while making customer-centric technical decisions.
Top Skills: AWSGitKotlinMySQLPythonRedisRpc

What you need to know about the Chicago Tech Scene

With vibrant neighborhoods, great food and more affordable housing than either coast, Chicago might be the most liveable major tech hub. It is the birthplace of modern commodities and futures trading, a national hub for logistics and commerce, and home to the American Medical Association and the American Bar Association. This diverse blend of industry influences has helped Chicago emerge as a major player in verticals like fintech, biotechnology, legal tech, e-commerce and logistics technology. It’s also a major hiring center for tech companies on both coasts.

Key Facts About Chicago Tech

  • Number of Tech Workers: 245,800; 5.2% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: McDonald’s, John Deere, Boeing, Morningstar
  • Key Industries: Artificial intelligence, biotechnology, fintech, software, logistics technology
  • Funding Landscape: $2.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Pritzker Group Venture Capital, Arch Venture Partners, MATH Venture Partners, Jump Capital, Hyde Park Venture Partners
  • Research Centers and Universities: Northwestern University, University of Chicago, University of Illinois Urbana-Champaign, Illinois Institute of Technology, Argonne National Laboratory, Fermi National Accelerator Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account