Job

Staff Software Engineer — LLM Serving, Routing & Model Infrastructure

Trainety Curated Opportunities

Location
United States
Industry
Legal
Organization size
Individual
Updated
September 19, 2026

Description

Every Harvey AI request eventually passes through infrastructure that has to decide where it should run, whether the selected model is healthy, what happens if a provider fails, and whether the latency and cost are acceptable.


This position owns that layer.


The Model Infrastructure team is building the systems that connect Harvey’s applications with multiple frontier-model providers while keeping production workloads highly available. The engineer will work on the Unified Model Controller and related infrastructure that can intelligently route traffic according to reliability, quality, latency, compliance requirements, and cost.


Failure handling is a central engineering problem. Model APIs can degrade, providers can have partial outages, capacity can tighten, and latency can change rapidly. Harvey needs automated health monitoring, failover, traffic management, provisioning, and recovery systems so those issues do not immediately become user-facing failures.


Observability goes beyond ordinary service metrics. The platform tracks model health, token usage, latency, cost attribution, capacity, and reliability across providers. That information supports both production operations and decisions about when a new model should replace or complement an existing one.


Engineers will also integrate emerging providers and model APIs, allowing Harvey to adopt new models without rebuilding the surrounding product stack. Current relevant ecosystems include OpenAI, Anthropic, Azure OpenAI, Fireworks, Baseten, and open-source models.


This role favors someone with deep distributed-systems experience rather than someone focused only on prompt engineering. Go, Java, Python, Rust, or C++ experience can be relevant, along with Kubernetes, networking, cloud infrastructure, service meshes, observability, and high-availability systems.


Background in inference gateways, GPU platforms, capacity management, model serving, Kafka/Spark-style data infrastructure, or SRE can transfer particularly well.


The technical challenge is essentially to make rapidly changing AI models behave like a dependable production utility for the rest of Harvey’s engineering organization.


Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.


https://www.harvey.ai/company/careers/3ae2aecf-16c4-4e6e-a8fd-0db33ebeb16c

Expertise

  • LLM Serving
  • Model Routing
  • Distributed Systems
  • Kubernetes
  • Observability
  • Traffic Engineering
  • Cloud Infrastructure
  • Curated Opportunity

More from Trainety Curated Opportunities

Explore more opportunities

Continue browsing available Jobs and Projects on Trainety.