Similar Jobs

See all

About the Role:

  • We are seeking a highly motivated Research Engineer Intern to work on Trainium, GPU kernels, and LLM systems optimization over a 12–16 week internship.
  • You will own a well-scoped project at the intersection of AI Systems, Compiler and Runtime Optimization, and Distributed Training Inference, taking it from design to working, profiled code on real hardware.

Responsibilities:

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA.
  • Profile and improve inference performance in vLLM, SGLang, and our custom runtimes — kernel fusion, scheduling, KV-cache and memory optimizations.
  • Build benchmarks, chase down performance regressions, and turn profiler traces into concrete speedups.

Qualifications:

  • Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a related field with solid programming skills in Python and familiarity with C++.
  • Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy) from coursework, research, or projects.
  • Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels — class projects and personal projects count.

Yotta Labs

Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are a remote-first team with a focus on high-performance AI computing and offer a flexible, collaborative work environment.

Apply for This Position