Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.
Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.
Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or Neuron SDK.
Profile and improve inference performance in vLLM, SGLang, and custom runtimes through kernel fusion, scheduling, and memory optimizations.
Ship code upstream to open-source AI infrastructure projects with tests and documentation, working on a well-scoped project from design to production.
Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are a remote-first team with a focus on high-performance AI computing and offer a flexible, collaborative work environment.
Develop and optimize low-level kernels, runtime components, and system software for high-performance AI inference workloads.
Improve inference engine performance across GPU platforms by identifying bottlenecks and implementing advanced optimization techniques.
Profile, debug, and resolve system-level and hardware-level performance issues across CPU and GPU environments.
This position is listed on behalf of a partner company building cutting-edge AI infrastructure for large-scale inference platforms. They operate in a highly technical, international, and innovation-driven environment where engineering excellence and ownership are valued.
Architect and maintain production high-traffic LLM serving systems.
Optimize throughput, latency, and cost for leading open-source LLMs.
Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.
We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.