Cox Exponential Jobs

Founding Engineer, AI Infra

Cox Exponential

Founding Engineer, AI Infra

Posted 24 Days Ago

Remote or Hybrid

Hiring Remotely in CA, USA

Senior level

Remote or Hybrid

Hiring Remotely in CA, USA

Senior level

Design, build, and operate end-to-end training and inference infrastructure for large language and multimodal models. Improve efficiency (memory, parallelism, kernel optimizations), ensure robust scalable training and RL pipelines, optimize low-latency/high-throughput serving (quantization, caching, speculative decoding), manage multi-GPU and multi-cloud orchestration, and productionize new algorithms with strong observability and reproducibility.

The summary above was generated by AI

About Goaly

At Goaly, our mission is to make custom AI affordable for every business. Our founding team comes from the front lines of top AI labs and tech giants (Meta MSL, TikTok AI, Google DeepMind, xAI, Microsoft Research, etc.), where we built large-scale training infrastructure powering trillion-parameter models and scaled GenAI models to a global user base. Now, we are building something we wish we had before: a platform that makes training and adapting custom AI affordable for all modern companies, not just Big Tech. Our north star is ambitious: for a domain-specific task, reach 90% of SOTA performance at less than 10% of the cost. To get a taste of what we are doing, see our first tech blog.

About the Role

You will sit at the intersection of systems engineering and applied ML, building specialized infrastructure that keeps large language and multimodal models fast, reliable, and cost-effective. You will partner with research, product, and infra teams to ship production-ready platforms for training and serving AI at scale.

Key Responsibilities

Efficiency & performance: Improve LLM training and inference efficiency through better memory utilization, optimized parallelism, and kernel-level innovations (e.g. FlashAttention, CUDA/Triton).
Training & RL robustness: Build scalable, stable training and RL pipelines with strong reproducibility, observability, and debuggability.
Serving & inference optimization: Design and tune high-throughput, low-latency model serving systems, including quantization, caching, and speculative decoding.
Scalability & infrastructure: Own end-to-end training and inference infrastructure — from data ingestion and checkpointing to multi-GPU and multi-cloud orchestration.
Production enablement: Work closely with researchers and product engineers to turn new algorithms into reliable, production-ready systems.

Requirements

5+ years building or operating ML infrastructure at scale, ideally supporting large language or multimodal models.
Deep understanding of GPU architecture, distributed training frameworks (PyTorch, DeepSpeed, Megatron, Ray), and parallelism strategies.
Hands-on experience running inference stacks (vLLM / SGLang, TGI, Triton) and optimizing them via low-level profiling.
Strong software engineering fundamentals in Python and one of C++/Rust/Go, with clean, reliable code shipped to production.
Working knowledge of modern data pipelines, feature stores, and vector databases used in production AI systems.
Comfort automating infrastructure with Kubernetes, Terraform/Pulumi, and observability stacks (Prometheus, Grafana, OpenTelemetry).

Bonus Points

Experience deploying open-source LLMs (Llama 3, Qwen, DeepSeek) or training custom foundation models.
Contributions to ML systems tooling (compilers, kernels, inference runtimes) or open-source infrastructure projects.
Background in reinforcement learning, evaluation harnesses, or alignment tooling that hardens production AI systems.

Similar Jobs

Apryse

Growth Marketing Manager

18 Hours Ago

Remote

90K-120K Annually

Mid level

90K-120K Annually

Mid level

Productivity • Software • App development • Automation

Run pipeline, lifecycle, and demand programs to drive multi-seat B2B SaaS conversions. Build and execute full-funnel campaigns, manage HubSpot workflows and reporting, partner with sales on account targeting, and run customer advocacy, review-generation, and content initiatives to grow pipeline and bookings.

Top Skills: Ai ToolsAutomation PlatformsCanvaCapterraFigmaG2HubspotMartech

Inspiren

VP of Quality

18 Hours Ago

Easy Apply

In-Office or Remote

United States

Easy Apply

260K-300K Annually

Expert/Leader

260K-300K Annually

Expert/Leader

Artificial Intelligence • Hardware • Healthtech • Software

The VP of Quality leads the development and maintenance of the Quality Management System (QMS), ensures compliance with ISO 13485, collaborates with engineering on product quality, and develops a high-performing quality team.

Top Skills: CapaFmeaIec 62304Iso 13485Plm Software

Block

Machine Learning Engineer

18 Hours Ago

In-Office or Remote

CA, USA

277K-415K Annually

Expert/Leader

277K-415K Annually

Expert/Leader

Blockchain • eCommerce • Fintech • Payments • Software • Financial Services • Cryptocurrency

Design, build, and operate production ML decision systems to detect and prevent payment fraud, account takeover, scams, and other abuse. Integrate diverse signals into low-latency serving and batch scoring, own feature pipelines and model lifecycle, develop AI-assisted triage and feedback loops, and partner cross-functionally to balance fraud reduction with legitimate customer access.

Top Skills: Cloud InfrastructureData LakehouseData WarehouseEmbeddingsFeature StoreJavaKafkaKotlinKubernetesLightgbmModel ServingMonitoringObservabilityPythonPyTorchSQLTensorFlowWorkflow OrchestrationXgboost

What you need to know about the Chicago Tech Scene

With vibrant neighborhoods, great food and more affordable housing than either coast, Chicago might be the most liveable major tech hub. It is the birthplace of modern commodities and futures trading, a national hub for logistics and commerce, and home to the American Medical Association and the American Bar Association. This diverse blend of industry influences has helped Chicago emerge as a major player in verticals like fintech, biotechnology, legal tech, e-commerce and logistics technology. It’s also a major hiring center for tech companies on both coasts.

Key Facts About Chicago Tech

Number of Tech Workers: 245,800; 5.2% of overall workforce (2024 CompTIA survey)
Major Tech Employers: McDonald’s, John Deere, Boeing, Morningstar
Key Industries: Artificial intelligence, biotechnology, fintech, software, logistics technology
Funding Landscape: $2.5 billion in venture capital funding in 2024 (Pitchbook)
Notable Investors: Pritzker Group Venture Capital, Arch Venture Partners, MATH Venture Partners, Jump Capital, Hyde Park Venture Partners
Research Centers and Universities: Northwestern University, University of Chicago, University of Illinois Urbana-Champaign, Illinois Institute of Technology, Argonne National Laboratory, Fermi National Accelerator Laboratory