Source Job

Global

  • Implement and optimize compute kernels for Attention, GEMM, MoE, and quantization on NVIDIA, AMD, or AWS Trainium using CUDA, Triton, ROCm/HIP, or Neuron SDK.
  • Profile and improve inference performance in vLLM, SGLang, and custom runtimes through kernel fusion, scheduling, and memory optimizations.
  • Ship code upstream to open-source AI infrastructure projects with tests and documentation, working on a well-scoped project from design to production.

CUDA Python C++ PyTorch

13 jobs similar to Research Engineer Intern - AI Systems

Jobs ranked by similarity.

US

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.
  • Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.

We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.

Europe

  • Develop and optimize low-level kernels, runtime components, and system software for high-performance AI inference workloads.
  • Improve inference engine performance across GPU platforms by identifying bottlenecks and implementing advanced optimization techniques.
  • Profile, debug, and resolve system-level and hardware-level performance issues across CPU and GPU environments.

This position is listed on behalf of a partner company building cutting-edge AI infrastructure for large-scale inference platforms. They operate in a highly technical, international, and innovation-driven environment where engineering excellence and ownership are valued.

Canada

  • Optimize machine learning inference systems for latency, throughput, and cost-efficiency.
  • Profile and troubleshoot GPU/CPU bottlenecks, implement advanced techniques like quantization and speculative decoding.
  • Collaborate with research and engineering teams to productionize new models and improve inference infrastructure.

The company is an AI-focused organization that develops advanced machine learning systems for production environments. It values technical excellence and experimentation, offering a flexible remote work environment.

EMEA

  • Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
  • Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
  • Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.

They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

US

  • Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
  • Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
  • Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.

Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.

Canada

  • Lead the design and operation of GPU infrastructure for AI workloads.
  • Manage Kubernetes-based environments and optimize for AI training and inference.
  • Define operational standards, implement monitoring, and collaborate with AI engineering teams.

ELEKS is a software engineering company that partners with enterprises to accelerate digital transformation. They have a global team of over 2,000 professionals and foster a culture of innovation and collaboration.

United States

  • Architect and build large-scale ML systems spanning data, training, evaluation, inference, and deployment.
  • Implement evaluation pipelines covering performance, robustness, safety, and bias.
  • Own production deployment including GPU optimization, memory efficiency, latency reduction, and scaling policies.

$190,000–$230,000/yr
US

  • Lead the design and development of our production inference platform, defining the technical roadmap for inference infrastructure, model serving, and runtime optimization.
  • Build and operate scalable, cost-effective systems for serving large language models in production, optimizing latency, throughput, GPU utilization, and memory efficiency.
  • Partner with ML engineers to productionize new models and inference techniques, establish benchmarking methodologies, and make key architectural decisions.

Syllo is on a mission to transform litigation with a unified platform that enables lawyers to safely harness AI. Since going to market, they have gained diverse enterprise customers including big law firms and corporations, and are quickly expanding.

North America Unlimited PTO

  • Serve as the primary technical point of contact for teams running large-scale training and inference workloads, owning onboarding end to end.
  • Diagnose and resolve complex failures in customer environments, from network fabric to ML frameworks, and build automation to prevent recurrence.
  • Profile and improve distributed training performance, lead incident response, and turn field insights into product improvements.

Andromeda Cluster provides scaled AI infrastructure to early-stage startups, founded by Nat Friedman and Daniel Gross. They work with leading AI labs, data centers, and cloud providers to deliver compute globally, building the liquidity layer for AI compute.

$140,000–$150,000/yr
Global

  • Design and execute performance benchmarks for AI training and inference workloads.
  • Profile and characterize GPU workloads to identify bottlenecks and optimization opportunities.
  • Systematically tune workload parameters to maximize throughput and establish performance baselines across GPU platforms.

Vultr provides high-performance, affordable cloud infrastructure for enterprises and AI innovators with 33 global data centers. It is the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, with a culture focused on growth and innovation.

$125,000–$250,000/yr
Global

  • Design, operate, and improve reliable infrastructure for AI training and inference workloads.
  • Build monitoring, alerting, runbooks, and incident-response practices for easier operations.
  • Partner with ML, research, and platform teams to translate workload needs into infrastructure improvements.

Boson AI builds production-grade AI systems that make communication with AI more natural, capable, and useful. The team is focused on infrastructure reliability, operating GPU clusters and networks for AI workloads.