Hyperbolic Logo

Hyperbolic

Forward Deployed Infrastructure Engineer - Eastern US

Reposted 21 Days Ago
Remote
Hiring Remotely in USA
Mid level
Remote
Hiring Remotely in USA
Mid level
The role involves benchmarking infrastructure performance, designing tests, debugging customer trials, and maintaining benchmarking infrastructure while ensuring clear documentation and communication.
The summary above was generated by AI
Who We Are

Hyperbolic Labs is on a mission to democratize AI by breaking down the barriers to computing power with our Open-Access AI Cloud. By making better use of idle computing resources across the globe, we offer an innovative GPU marketplace and AI inference service that promise affordability and accessibility for all. As pioneers at the intersection of AI and open-source technology, we believe in an open future where AI innovation is limited only by imagination, not by access to resources. We're looking for forward-thinking individuals who share our passion for making AI universally accessible, secure, and affordable. Join us in building a platform that empowers innovators everywhere to turn their visionary AI projects into reality.

NOTE: This role is focused on EASTERN TIME ZONE

About the Role

Our reserved customers run large multinode GPU clusters, and when those clusters misbehave the problem is rarely simple. You are the engineer embedded with those customers: you stand their cluster up, you hand it over, and you stay with it.

You do not own tickets. You own environments. Technical Support Engineers own the ticket lifecycle and pull you in when an issue needs real depth: multinode collective performance, hardware faults, fabric problems, or a provider who needs to be told what is wrong with their hardware.

One thing we will be straight about, because it shapes the job. We aggregate capacity from suppliers rather than owning most of the hardware ourselves. That means a real part of this role is technical liaison work: proving where a fault actually lives, taking it to the provider with evidence, and coordinating the fix on the customer's behalf. The engineers who enjoy this role are the ones who find that interesting rather than frustrating.

Who You Are

  • Cluster stand-up and handoff. Build, validate, and benchmark new customer clusters, then hand them over with documentation the customer's own engineers can work from.

  • Deep escalations. Multinode and NCCL performance debugging, GPU and hardware faults (XID and ECC errors, lspci, dmesg), driver and fabric issues, container and scheduler problems.

  • Provider escalation and coordination. Maintenance windows, RMAs, hung nodes, and disputed fault attribution. You bring the evidence that makes the provider act, and you keep the customer informed while it happens.

  • Embedded ownership of named accounts. You are the engineer your customers know by name. You learn their workload, not just their infrastructure, and you tell them what to change before they hit the wall.

  • Proactive monitoring. Own the monitoring and alerting we put in front of customer clusters (Grafana, Prometheus) so we find faults before the customer reports them.

  • Tooling and pushing work down. Automate the repeat work and turn your own escalations into runbooks the L1 tier can run. Anything you fix three times should stop reaching you.

  • Deep Linux experience and total comfort in the CLI, including in someone else's broken environment.

  • Hands-on multinode GPU experience: NCCL, InfiniBand or RoCE, collective performance debugging, topology and placement.

  • Hardware fault triage on GPU nodes: XID and ECC errors, dmesg, lspci, nvidia-smi, thermal and power faults.

  • Experience provisioning and operating GPU clusters with Kubernetes, Slurm, or both.

  • Genuinely comfortable customer-facing, including delivering bad news and saying “this is ours” or “this is the provider’s” with confidence.

  • Sound judgment on when to keep digging and when to escalate.

Preferred Qualifications

  • Experience with parallel filesystems (Weka, Lustre, GPFS) and high-performance storage.

  • Grafana, Prometheus, or similar observability stacks in production.

  • Prior forward deployed engineer, solutions architect, or technical account manager experience.

  • Infrastructure as code (Terraform, Ansible) and CI for cluster provisioning.

  • Exposure to model training or inference workloads from the practitioner side.

Hyperbolic is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.

Similar Jobs

28 Seconds Ago
Remote or Hybrid
USA
130K-385K Annually
Expert/Leader
130K-385K Annually
Expert/Leader
Hardware • Healthtech • Software • Analytics
Own the full enterprise SaaS sales cycle, from prospecting through executive negotiations and close, while expanding existing accounts. Build pipeline, manage forecasts, develop executive relationships, and partner with Product, Marketing, Customer Success, and Leadership to refine go-to-market strategy. Contribute to sales enablement, account strategy, sales methodology, and the growth of a scalable commercial organization.
Top Skills: SaaS
6 Minutes Ago
Remote or Hybrid
221K-387K Annually
Senior level
221K-387K Annually
Senior level
Artificial Intelligence • Cloud • HR Tech • Information Technology • Productivity • Software • Automation
Lead product vision and execution for an MVP autonomous agent platform, prototyping and shipping agentic capabilities. Own roadmap for orchestration, memory, action execution, and secure governance. Drive adoption, define metrics, partner with engineering, GTM, and open-source ecosystem, and translate frontier AI research into scalable platform primitives.
Top Skills: Agent OrchestrationAgentic AiLlm OrchestrationLlmsLow-Code/No-CodeMemory ArchitecturesModel Fine-TuningOpen-Source Ai FrameworksPrompt EngineeringRpaServicenow Ai
6 Minutes Ago
Remote or Hybrid
United States
84K-124K Annually
Entry level
84K-124K Annually
Entry level
Cloud • Fintech • Information Technology • Machine Learning • Software
Manage a portfolio of accounting and bookkeeping partners to drive retention, platform adoption, revenue growth, and subscription expansion. Conduct business reviews, create success plans, identify churn risks, and position Xero capabilities including payroll, bill pay, analytics, and AI tools. Use Salesforce CRM for account health, forecasting, and pipeline management while collaborating with consultants, sales operations, marketing, finance, and customer experience teams.
Top Skills: Ai-Powered SoftwareAnalytics ToolsPayroll SoftwareSalesforce CRMXero Platform

What you need to know about the Chicago Tech Scene

With vibrant neighborhoods, great food and more affordable housing than either coast, Chicago might be the most liveable major tech hub. It is the birthplace of modern commodities and futures trading, a national hub for logistics and commerce, and home to the American Medical Association and the American Bar Association. This diverse blend of industry influences has helped Chicago emerge as a major player in verticals like fintech, biotechnology, legal tech, e-commerce and logistics technology. It’s also a major hiring center for tech companies on both coasts.

Key Facts About Chicago Tech

  • Number of Tech Workers: 245,800; 5.2% of overall workforce (2024 CompTIA survey)
  • Major Tech Employers: McDonald’s, John Deere, Boeing, Morningstar
  • Key Industries: Artificial intelligence, biotechnology, fintech, software, logistics technology
  • Funding Landscape: $2.5 billion in venture capital funding in 2024 (Pitchbook)
  • Notable Investors: Pritzker Group Venture Capital, Arch Venture Partners, MATH Venture Partners, Jump Capital, Hyde Park Venture Partners
  • Research Centers and Universities: Northwestern University, University of Chicago, University of Illinois Urbana-Champaign, Illinois Institute of Technology, Argonne National Laboratory, Fermi National Accelerator Laboratory

Sign up now Access later

Create Free Account

Please log in or sign up to report this job.

Create Free Account