Similar Jobs
See allLLM Inference Engineer
NEAR AI
US
PyTorch
CUDA
System Engineer (Token Factory)
Jobgether
Europe
C++
CUDA
Linux
Machine Learning Engineer — Inference Optimization
Partner Company
Canada
PyTorch
CUDA
GPU Optimization
AI Infrastructure Engineer
Pragmatike
EMEA
Python
Kubernetes
VLLM
Engineering Manager
Cohere
Global
Kubernetes
About the Role:
- We are seeking a highly motivated Research Engineer Intern to work on Trainium, GPU kernels, and LLM systems optimization over a 12–16 week internship.
- You will own a well-scoped project at the intersection of AI Systems, Compiler and Runtime Optimization, and Distributed Training Inference, taking it from design to working, profiled code on real hardware.
Responsibilities:
- Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or the Neuron SDK with PyTorch/XLA.
- Profile and improve inference performance in vLLM, SGLang, and our custom runtimes — kernel fusion, scheduling, KV-cache and memory optimizations.
- Build benchmarks, chase down performance regressions, and turn profiler traces into concrete speedups.
Qualifications:
- Currently pursuing a BS, MS, or PhD in Computer Science, Computer Engineering, or a related field with solid programming skills in Python and familiarity with C++.
- Understanding of GPU/accelerator architecture fundamentals (memory hierarchy, parallelism, occupancy) from coursework, research, or projects.
- Experience writing CUDA, Triton, ROCm/HIP, or Neuron kernels — class projects and personal projects count.
Yotta Labs
Yotta Labs is building the next generation multi-silicon AI cloud and runtime platform to power the world’s most demanding AI workloads. They are a remote-first team with a focus on high-performance AI computing and offer a flexible, collaborative work environment.