Source Job

Europe

  • Develop and optimize low-level kernels, runtime components, and system software for high-performance AI inference workloads.
  • Improve inference engine performance across GPU platforms by identifying bottlenecks and implementing advanced optimization techniques.
  • Profile, debug, and resolve system-level and hardware-level performance issues across CPU and GPU environments.

C++ CUDA Linux

20 jobs similar to System Engineer (Token Factory)

Jobs ranked by similarity.

US

  • Architect and maintain production high-traffic LLM serving systems.
  • Optimize throughput, latency, and cost for leading open-source LLMs.
  • Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.

We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.

$190,000–$230,000/yr
US

  • Lead the design and development of our production inference platform, defining the technical roadmap for inference infrastructure, model serving, and runtime optimization.
  • Build and operate scalable, cost-effective systems for serving large language models in production, optimizing latency, throughput, GPU utilization, and memory efficiency.
  • Partner with ML engineers to productionize new models and inference techniques, establish benchmarking methodologies, and make key architectural decisions.

Syllo is on a mission to transform litigation with a unified platform that enables lawyers to safely harness AI. Since going to market, they have gained diverse enterprise customers including big law firms and corporations, and are quickly expanding.

$140,000–$150,000/yr
Global

  • Design and execute performance benchmarks for AI training and inference workloads.
  • Profile and characterize GPU workloads to identify bottlenecks and optimization opportunities.
  • Systematically tune workload parameters to maximize throughput and establish performance baselines across GPU platforms.

Vultr provides high-performance, affordable cloud infrastructure for enterprises and AI innovators with 33 global data centers. It is the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, with a culture focused on growth and innovation.

Europe

  • Design and develop advanced computing architectures for AI accelerators.
  • Analyze and optimize computational workloads for low power and high performance.
  • Collaborate with cross-functional teams to translate system requirements into architectural solutions.

Axelera AI develops next-generation AI platforms to advance humanity. With 220+ employees and $370 million raised, they foster a collaborative, innovative culture.

  • Optimize production LLM serving with vLLM and SGLang to maximize throughput and minimize latency through batching and quantization.
  • Profile training runs to find bottlenecks and resolve them with attention implementations like FlashAttention on H200 and GB200 hardware.
  • Deploy and operate multiple models on shared GPU clusters with autoscaling, bin-packing, and efficient handling of mixed workloads.

Egen is a fast-growing technology company with a data-first mindset, partnering with clients on Google Cloud and Salesforce to drive action through data and insights. We are a team of dedicated engineers who thrive on solving tough problems and continually innovate to achieve fast, effective results.

US

  • Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
  • Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
  • Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.

Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.

$100,000–$180,000/yr
US Unlimited PTO

  • Participate in sales meetings, provide architectural recommendations, and build proof-of-concept solutions for onboarding high-spending customers.
  • Troubleshoot and resolve complex technical issues using code analysis, scripting, and log analysis.
  • Create and maintain technical documentation and deliver training sessions, webinars, and demos.

Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We are a small, remote-first team that takes ownership seriously, moves fast, and ships work relied on by more than a million developers daily.

France

  • Contribute to high-performance C++ backend solutions supporting genomic analysis and precision healthcare.
  • Optimize large-scale systems and take end-to-end ownership of technical projects from design to deployment.
  • Collaborate with backend engineers, bioinformaticians, and product stakeholders to deliver robust, scalable solutions.

This position is listed on behalf of a partner company specializing in genomic analysis and precision healthcare technologies. The team is collaborative and distributed, working with multidisciplinary experts across software engineering, bioinformatics, and product development.

US

  • Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
  • Investigate performance, availability, and reliability issues across infrastructure and platform components.
  • Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.

Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.

UK 6w PTO

  • Design and build performance-critical components for a high-frequency trading platform.
  • Own the full lifecycle from design to production support, including incident response.
  • Mentor junior engineers and contribute to architectural decisions that shape the platform.

Reactive Markets is the 2026 OTC Trading Platform of the Year, handling over $50 billion in daily trading volumes across FX, Equities, and Cryptocurrency. Our engineering team is small, senior, and deeply invested in building cutting edge, reliable, high-performance systems.

Europe 7w PTO

  • Act as the right hand to the CTO, steering backend and ML teams while remaining hands-on in development.
  • Translate complex business challenges into scalable technical solutions and lead small teams toward fast execution.
  • Ensure alignment across engineering, product, and infrastructure in a fast-moving, globally distributed environment.

The company is a technology organization focused on AI product development and engineering. They operate in a globally distributed, fast-paced environment with small, focused teams and a startup mindset.

Europe

  • Build and scale massive distributed compute and storage systems for AI training.
  • Architect multi-cluster orchestration and optimize workload placement across global regions.
  • Design future-proof storage formats and implement metadata systems for exabyte-scale growth.

Mistral provides full-stack AI solutions, from frontier models to developer tools. They are a dynamic, collaborative team with a diverse workforce, passionate about innovation and low-ego teamwork.

Germany

  • Contribute to the design, development, and operation of compute node services for large-scale GPU-based infrastructure.
  • Work on virtualization layers, Kubernetes, and systems engineering to optimize performance and scalability.
  • Collaborate on integrating GPU, DPU, and high-performance hardware acceleration technologies.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They operate globally with a focus on efficient recruitment. The culture is collaborative and leverages technology to enhance hiring processes.

India

  • Develop, maintain, and optimize applications using C++ in a Linux-based environment.
  • Write clean, maintainable, and high-performance code aligned with software engineering best practices.
  • Test, debug, and deploy applications while ensuring system stability and reliability.

Jobgether uses AI-powered matching to streamline job applications for candidates. This posting is on behalf of a partner company, which manages all applications and next steps; the role is part of a collaborative and technology-driven environment.

Switzerland

  • Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
  • Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
  • Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.

Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.

EMEA

  • Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
  • Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
  • Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.

They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.

India

  • Research and implement state-of-the-art techniques to accelerate AI inference: quantization, sparsity, distillation, speculative decoding, and caching.
  • Partner closely with hardware and compiler teams to ensure algorithmic improvements translate to real gains on custom silicon.
  • Build profiling tools and comprehensive benchmarking frameworks to measure model quality and efficiency.

EnCharge AI is building the next generation AI platform using novel in-memory-computing architecture. The team consists of experienced AI researchers, silicon & systems engineers, and architects backed by leading investors.

$124,000–$329,200/yr
US

  • Design and scale highly available backend services and APIs supporting AI-powered developer tools.
  • Develop distributed systems optimizing reliability, latency, cost, and performance at global scale.
  • Provide technical leadership through mentorship, code reviews, and collaboration across engineering teams.

Our partner is building next-generation AI-assisted software development experiences. This remote-first role contributes to a globally impactful AI platform with a collaborative culture focused on innovation, inclusion, and technical excellence.

US

  • Improve core inference services including networking, speech processing, audio transcoding, and latency optimization.
  • Develop processes for measuring, building, and optimizing services to maximize system performance.
  • Debug complex system issues involving networking, scheduling, and high performance computing interactions.

Deepgram is the leading platform for the Voice AI economy, providing real-time APIs for speech-to-text, text-to-speech, and voice agents. With over 200,000 developers and 1,300+ organizations, including major partners like Twilio and Cloudflare, Deepgram emphasizes an AI-first mindset and rapid innovation.

Global 6w PTO 26w maternity 26w paternity

  • Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
  • Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
  • Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.