Arcadia is the most trusted healthcare platform powering outcomes. We transform complex healthcare data into trusted intelligence, helping providers, payers, and life sciences organizations act with clarity, make confident decisions, and achieve measurable clinical, operational, and financial outcomes.
Built on a comprehensive data foundation spanning tens of millions of patient lives, Arcadia combines advanced analytics and responsible AI to surface meaningful insights, coordinate action, and improve performance at scale. Our approach to AI and automation is governed and transparent — designed to strengthen human expertise, not replace judgment or obscure responsibility.
Hundreds of organizations rely on Arcadia to improve cost, quality, and outcomes. Backed by Nordic Capital, we continue to invest in our platform, AI capabilities, and people as we pursue our purpose: helping healthcare deliver better outcomes for every person, every community, and every generation.
Why This Role Is Important to Arcadia
Arcadia’s data and analytics platform is used by hundreds of health systems, ACOs, payers, and life sciences organizations, touching tens of millions of patient lives. This role owns how our agentic capabilities perform at that same scale: accurate, transparent about their own confidence, and safe for the clinicians, care teams, and patients who depend on them.
As a staff-level individual contributor, you will own the product-layer decisions that shape agent behavior, including prompting, retrieval and context, memory and state, evaluation, and escalation, while partnering with Product and Engineering on the systems that support them. Your work will help Arcadia make evidence-based launch decisions and scale responsible AI that is steerable, trustworthy, and ready for real healthcare workflows.
- You have established a production-grounded baseline for priority agentic workflows, with documented failure modes, severity-weighted evaluation rubrics, and a clear measurement plan
- You have mapped the current retrieval, context, memory, and escalation patterns and identified the highest-value opportunities to improve reliability, calibration, and cost
- You have earned trust across Product and Engineering by turning production evidence into clear, actionable recommendations
In 6 months
- Production-representative evaluation suites and regression checks inform model-change decisions for priority agentic workflows
- You have delivered measurable improvements in accuracy, reliability, steerability, latency, or cost for one or more priority workflows
- Human-review and escalation behavior has been validated under adversarial and edge-case conditions, with decision criteria and ownership boundaries clearly documented
In 12 months
- Arcadia has a repeatable product-layer AI performance practice that moves from production failure to diagnosis, experiment, evaluation, and release decision
- High-severity regressions are caught earlier, and agent behavior is more transparent, calibrated, and trustworthy at scale
- Model cards, intended-use guidance, limitations, and performance documentation are current and useful to product and customer-facing teams
What You'll Be Doing
- Design and iterate on agent behavior across real, live workflows, including long-horizon, multi-turn agentic tasks
- Design retrieval and context architecture so the right source data reaches a model in the right structure and agents remain grounded in real data rather than filling gaps with assumptions
- Design memory and state handling across multi-turn and multi-agent flows, determining what is carried forward, summarized, or dropped and why
- Create context and prompt templates that combine few-shot examples, structured formatting, and reasoning scaffolding for consistent agent behavior
- Improve performance through prompting, tool-use strategy, and context construction, validated through direct experimentation rather than guesswork
- Build and run evaluations against real production conditions to measure performance, regressions, failure modes, and edge cases
- Author evaluation rubrics, quality heuristics, and thresholds that weight failures by severity and cost, not just frequency, and monitor those measures against production behavior
- Design and validate escalation paths that route agents to human review based on confidence and uncertainty while preserving safety and consistency under adversarial and edge-case conditions
- Design for cost-aware performance alongside latency, reliability, and accuracy through efficient context construction and tool-call economy
- Evaluate and sign off on model changes by baselining current behavior, running comparative evaluations, and making the go/no-go call before a change reaches a customer
- Maintain product-level AI documentation, including model cards, intended use, limitations, and known failure modes, so customer-facing teams work from actual agent behavior
- Partner closely with Product and product managers to ensure agents are not just capable, but steerable, trustworthy, and ready to scale
What You'll Bring
- We value equivalent practical experience that demonstrates the depth required for this staff-level role
- 8+ years of production software engineering experience, including 3+ years of hands-on ownership of ML, LLM, or agentic systems in production, with direct experience in healthcare, finance, or another regulated industry
- Demonstrated ability to diagnose why an agent failed, correctly attribute the fix to instruction, retrieval, context, or memory design, and weigh failures by severity and cost rather than frequency alone
- Hands-on experience with RAG architecture, production-grounded evaluation frameworks, and fallback or human-in-the-loop logic for automated systems
- Working familiarity with AWS AI/ML services, including Bedrock and SageMaker, sufficient to build and evaluate effectively in Arcadia’s environment
- Evidence-led judgment and the credibility to push back on launch decisions, paired with a builder’s instinct to run the experiment and move from a production failure to a fix
Would Love for You to Have
- Experience applying AI to healthcare data or workflows where safety, transparency, and calibrated uncertainty directly affect care teams or patients
- Experience with long-horizon, multi-turn or multi-agent workflows and product-level AI documentation such as model cards
What You'll Get
- The opportunity to define how agent performance, safety, and readiness are measured for production healthcare workflows
- Meaningful ownership across prompts, context, memory, evaluations, and escalation patterns at product scale
- A cross-functional role translating production evidence into AI improvements used across Arcadia’s platform
- A mission-driven company working to improve how patients receive care
- A flexible, remote-friendly culture with personality and heart
- Employee-driven programs and initiatives for personal and professional development
- Membership in the talented, energized, diverse, and purpose-driven Arcadian community
Similar Jobs at Arcadia
What you need to know about the Chicago Tech Scene
Key Facts About Chicago Tech
- Number of Tech Workers: 245,800; 5.2% of overall workforce (2024 CompTIA survey)
- Major Tech Employers: McDonald’s, John Deere, Boeing, Morningstar
- Key Industries: Artificial intelligence, biotechnology, fintech, software, logistics technology
- Funding Landscape: $2.5 billion in venture capital funding in 2024 (Pitchbook)
- Notable Investors: Pritzker Group Venture Capital, Arch Venture Partners, MATH Venture Partners, Jump Capital, Hyde Park Venture Partners
- Research Centers and Universities: Northwestern University, University of Chicago, University of Illinois Urbana-Champaign, Illinois Institute of Technology, Argonne National Laboratory, Fermi National Accelerator Laboratory

