Raydar Logo

Raydar

Forward Deployed Machine Learning Engineer

Posted 13 Hours Ago
Remote
Hiring Remotely in United States
170K-270K Annually
Mid level
Remote
Hiring Remotely in United States
170K-270K Annually
Mid level
Build production machine learning benchmarks and evaluation systems for foundation models. Own backend infrastructure including data pipelines, execution environments, storage, and orchestration, as well as sandboxed environments for agentic evaluations. Partner with researchers, enterprise customers, and company leadership to scope and deliver technical solutions. Identify scalable evaluation patterns, communicate clearly, and operate effectively in an ambiguous, high-ownership startup environment.
The summary above was generated by AI

About the company

Our client is a fast-growing AI data infrastructure company building a secure marketplace and data lab for high-quality model-training data. The company helps data holders license sensitive, real-world datasets to vetted AI teams while protecting governance, privacy, intellectual property, and security. It has raised $65 million, including a $55 million Series A backed by leading venture firms, employs approximately 80 people, and is scaling quickly after reaching its full-year growth target early.

The role and why it matters

This is the first Machine Learning Engineer dedicated to the company's Benchmarks and Evaluations vertical. You will partner directly with the general manager, researchers, and early enterprise customers to establish the technical foundation for evaluating foundation models across domains and modalities. The role combines hands-on ML evaluation, backend infrastructure, and customer-facing delivery in a high-ownership environment.

What you'll do

• Define, design, and build production benchmarks and evaluations with customers and internal researchers.

• Build and own backend infrastructure, including data pipelines, execution environments, storage, and orchestration.

• Create sandboxed environments for agentic evaluations involving tools, code execution, and multi-step tasks.

• Own the engineering portion of customer engagements from technical scoping through production delivery.

• Identify repeatable evaluation patterns and infrastructure gaps that can become scalable products.

• Move quickly through ambiguity while maintaining strong technical judgment and clear written communication.


Requirements

What we're looking for

• 4 or more years of engineering experience, including hands-on machine learning model evaluation work.

• Experience deploying end-to-end ML evaluation or benchmark systems to production against demanding customer timelines.

• Strong proficiency with ML evaluation frameworks and benchmark design, including approaches such as LLM-as-judge.

• Ownership of backend and infrastructure systems such as large-scale data pipelines, execution environments, storage, and orchestration.

• Customer-facing engineering experience managing enterprise stakeholders.

• A degree in computer science, physics, or a related technical field.

• Current unrestricted U.S. work authorization without visa sponsorship.

Bonus points

• Experience building evaluations or human-data pipelines for large language models.

• Experience working directly with AI researchers or foundation-model labs.

• Published work or meaningful open-source contributions in ML evaluations or benchmarks.

• Early-stage B2B startup, forward-deployed engineering, or high-ownership generalist experience.


Benefits

Compensation and benefits

• $170K-$270K base salary.

• Competitive equity.

• Comprehensive benefits provided by a well-capitalized, high-growth company.

• High autonomy and the opportunity to build a new technical vertical from the ground up.

Location and work model

• Full-time and remote within the United States.

• Strong independent ownership, customer responsiveness, and cross-functional collaboration are expected.

Similar Jobs

8 Days Ago
Remote
United States
175K-200K Annually
Senior level
175K-200K Annually
Senior level
Other
Build and deploy production generative AI and machine learning systems for utility clients. Own engagements end to end, including discovery, architecture, retrieval, model selection, evaluation, deployment, monitoring, operator-facing applications, and post-go-live support. Integrate with client data and business systems, establish AI safety and governance controls, manage scope and risk, and communicate with technical teams and executives. Contribute reusable accelerators, evaluation frameworks, and reference architectures across client engagements.
Top Skills: SparkAWSAzureCi/CdDashDatabricksDockerGenerative AiGitGoogle Cloud PlatformGradioGraphql ApisHybrid RetrievalJavaKnowledge GraphsLarge Language ModelsMachine LearningPythonReactRest ApisRetrieval-Augmented GenerationScalaStreamlitTypescriptVector Databases
27 Days Ago
Remote
United States of America
Entry level
Entry level
Information Technology • Software • Consulting
Build and deploy AI-powered applications for customers, combining Python and TypeScript application engineering with LLMs, generative AI, agent orchestration, APIs, distributed systems, and cloud platforms. Partner with stakeholders to understand requirements, design scalable architectures, rapidly prototype solutions, integrate AI capabilities, and evolve proofs of concept into production systems. The role also involves customer-facing consulting, technical decision-making, and collaboration throughout the solution lifecycle.
Top Skills: Agent OrchestrationAi ObservabilityAi/MlAPIsAWSAzureDistributed SystemsGCPGenerative AiLarge Language Models (Llms)MlopsModel DeploymentPythonRagTypescriptVector Databases
One Month Ago
Remote
United States
117K-187K Annually
Senior level
117K-187K Annually
Senior level
Edtech
Embed with strategic higher-education partners to design, build, and ship production AI and platform integrations (LLM/RAG, LTI) across Cengage products. Own end-to-end deployments, write production code, ensure compliance and accessibility, evaluate AI safety and outcomes, mentor FDEs, and translate field learnings into product roadmap contributions.
Top Skills: AgsApi GatewayAuroraAWSCaliperCloudwatchCoppaDeep LinkingEcsEksFerpaGraphQLJavaScriptLambdaLlm ApisLti 1.3Lti AdvantageMindtapNrpsPrompt EngineeringPythonRag (Retrieval-Augmented Generation)RdsRestS3SQLTool-Calling AgentsTypescriptWebassignXapi

What you need to know about the Chicago Tech Scene

With vibrant neighborhoods, great food and more affordable housing than either coast, Chicago might be the most liveable major tech hub. It is the birthplace of modern commodities and futures trading, a national hub for logistics and commerce, and home to the American Medical Association and the American Bar Association. This diverse blend of industry influences has helped Chicago emerge as a major player in verticals like fintech, biotechnology, legal tech, e-commerce and logistics technology. It’s also a major hiring center for tech companies on both coasts.

Key Facts About Chicago Tech

  • Number of Tech Workers: 245,800; 5.2% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: McDonald’s, John Deere, Boeing, Morningstar
  • Key Industries: Artificial intelligence, biotechnology, fintech, software, logistics technology
  • Funding Landscape: $2.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Pritzker Group Venture Capital, Arch Venture Partners, MATH Venture Partners, Jump Capital, Hyde Park Venture Partners
  • Research Centers and Universities: Northwestern University, University of Chicago, University of Illinois Urbana-Champaign, Illinois Institute of Technology, Argonne National Laboratory, Fermi National Accelerator Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account