Job

Supercomputing Engineer — GPU Infrastructure for Frontier AI

Trainety Curated Opportunities

Location
Remote / Worldwide
Industry
Technology & Internet
Organization size
Individual
Updated
September 9, 2026

Description

Thousands of GPUs are useful only if the infrastructure around them can provision resources predictably, move data fast enough, recover from failures, and keep long-running training jobs alive.


Magic's Supercomputing Platform & Infrastructure team owns that foundation.

The engineer will build and operate clusters supporting both training and inference, with Terraform-based infrastructure-as-code used to keep large environments reproducible and understandable. Kubernetes coordinates workloads across the GPU fleet, while networking, storage, operating systems, drivers, cloud services, and hardware all become part of the reliability surface.

Typical problems can range from designing modular infrastructure modules and improving cluster deployment consistency to diagnosing cross-layer performance failures. Long-context AI workloads generate sustained pressure on storage and networking, so the team also works on high-throughput data movement and automation around fault detection and recovery.This is not a conventional corporate cloud-infrastructure position. The systems directly support frontier model research, which means workload requirements can change quickly as training methods, architectures, and inference patterns evolve.

Candidates should bring strong production infrastructure experience and substantial hands-on knowledge of Terraform. Experience operating GPU clusters, Kubernetes, high-performance distributed platforms, major cloud providers, networking, storage, monitoring, and production incident response is especially valuable.

As Magic grows its compute footprint, this hire can influence broader supercomputing architecture rather than simply maintaining an established platform.Curated opportunity. Please verify details and apply via the original link below. No Signals are required for this project/job.

https://magic.dev/careers/45d25f90-3be6-417d-810f-0d95f7704961

Expertise

  • GPU Infrastructure
  • Kubernetes
  • Terraform
  • Distributed Systems
  • Cloud Infrastructure
  • Networking
  • Storage Systems
  • Observability

More from Trainety Curated Opportunities

Explore more opportunities

Continue browsing available Jobs and Projects on Trainety.