Architect and maintain production high-traffic LLM serving systems.
Optimize throughput, latency, and cost for leading open-source LLMs.
Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.
We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.
Lead the design and development of our production inference platform, defining the technical roadmap for inference infrastructure, model serving, and runtime optimization.
Build and operate scalable, cost-effective systems for serving large language models in production, optimizing latency, throughput, GPU utilization, and memory efficiency.
Partner with ML engineers to productionize new models and inference techniques, establish benchmarking methodologies, and make key architectural decisions.
Syllo is on a mission to transform litigation with a unified platform that enables lawyers to safely harness AI. Since going to market, they have gained diverse enterprise customers including big law firms and corporations, and are quickly expanding.
Design and execute performance benchmarks for AI training and inference workloads.
Profile and characterize GPU workloads to identify bottlenecks and optimization opportunities.
Systematically tune workload parameters to maximize throughput and establish performance baselines across GPU platforms.
Vultr provides high-performance, affordable cloud infrastructure for enterprises and AI innovators with 33 global data centers. It is the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, with a culture focused on growth and innovation.
Design and develop advanced computing architectures for AI accelerators.
Analyze and optimize computational workloads for low power and high performance.
Collaborate with cross-functional teams to translate system requirements into architectural solutions.
Axelera AI develops next-generation AI platforms to advance humanity. With 220+ employees and $370 million raised, they foster a collaborative, innovative culture.
Optimize production LLM serving with vLLM and SGLang to maximize throughput and minimize latency through batching and quantization.
Profile training runs to find bottlenecks and resolve them with attention implementations like FlashAttention on H200 and GB200 hardware.
Deploy and operate multiple models on shared GPU clusters with autoscaling, bin-packing, and efficient handling of mixed workloads.
Egen is a fast-growing technology company with a data-first mindset, partnering with clients on Google Cloud and Salesforce to drive action through data and insights. We are a team of dedicated engineers who thrive on solving tough problems and continually innovate to achieve fast, effective results.
Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.
Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.
Participate in sales meetings, provide architectural recommendations, and build proof-of-concept solutions for onboarding high-spending customers.
Troubleshoot and resolve complex technical issues using code analysis, scripting, and log analysis.
Create and maintain technical documentation and deliver training sessions, webinars, and demos.
Runpod is the AI Developer Cloud, providing a platform for over one million developers to experiment, train, fine-tune, deploy, and scale AI. We are a small, remote-first team that takes ownership seriously, moves fast, and ships work relied on by more than a million developers daily.
Contribute to high-performance C++ backend solutions supporting genomic analysis and precision healthcare.
Optimize large-scale systems and take end-to-end ownership of technical projects from design to deployment.
Collaborate with backend engineers, bioinformaticians, and product stakeholders to deliver robust, scalable solutions.
This position is listed on behalf of a partner company specializing in genomic analysis and precision healthcare technologies. The team is collaborative and distributed, working with multidisciplinary experts across software engineering, bioinformatics, and product development.
Monitor, operate, and support production AI infrastructure platforms and resolve incidents.
Investigate performance, availability, and reliability issues across infrastructure and platform components.
Collaborate with engineering teams, hardware vendors, and data center personnel to resolve technical issues.
Mirantis is the Kubernetes-native AI infrastructure company, enabling organizations to build and operate scalable, secure, and sovereign infrastructure for modern AI, machine learning, and data-intensive applications. It is a Silicon Valley leader with passionate, talented colleagues, offering a competitive compensation package and strong benefits.
Design and build performance-critical components for a high-frequency trading platform.
Own the full lifecycle from design to production support, including incident response.
Mentor junior engineers and contribute to architectural decisions that shape the platform.
Reactive Markets is the 2026 OTC Trading Platform of the Year, handling over $50 billion in daily trading volumes across FX, Equities, and Cryptocurrency. Our engineering team is small, senior, and deeply invested in building cutting edge, reliable, high-performance systems.
Act as the right hand to the CTO, steering backend and ML teams while remaining hands-on in development.
Translate complex business challenges into scalable technical solutions and lead small teams toward fast execution.
Ensure alignment across engineering, product, and infrastructure in a fast-moving, globally distributed environment.
The company is a technology organization focused on AI product development and engineering. They operate in a globally distributed, fast-paced environment with small, focused teams and a startup mindset.
Build and scale massive distributed compute and storage systems for AI training.
Architect multi-cluster orchestration and optimize workload placement across global regions.
Design future-proof storage formats and implement metadata systems for exabyte-scale growth.
Mistral provides full-stack AI solutions, from frontier models to developer tools. They are a dynamic, collaborative team with a diverse workforce, passionate about innovation and low-ego teamwork.
Contribute to the design, development, and operation of compute node services for large-scale GPU-based infrastructure.
Work on virtualization layers, Kubernetes, and systems engineering to optimize performance and scalability.
Collaborate on integrating GPU, DPU, and high-performance hardware acceleration technologies.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They operate globally with a focus on efficient recruitment. The culture is collaborative and leverages technology to enhance hiring processes.
Develop, maintain, and optimize applications using C++ in a Linux-based environment.
Write clean, maintainable, and high-performance code aligned with software engineering best practices.
Test, debug, and deploy applications while ensuring system stability and reliability.
Jobgether uses AI-powered matching to streamline job applications for candidates. This posting is on behalf of a partner company, which manages all applications and next steps; the role is part of a collaborative and technology-driven environment.
Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.
Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.
They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.
Research and implement state-of-the-art techniques to accelerate AI inference: quantization, sparsity, distillation, speculative decoding, and caching.
Partner closely with hardware and compiler teams to ensure algorithmic improvements translate to real gains on custom silicon.
Build profiling tools and comprehensive benchmarking frameworks to measure model quality and efficiency.
EnCharge AI is building the next generation AI platform using novel in-memory-computing architecture. The team consists of experienced AI researchers, silicon & systems engineers, and architects backed by leading investors.
Design and scale highly available backend services and APIs supporting AI-powered developer tools.
Develop distributed systems optimizing reliability, latency, cost, and performance at global scale.
Provide technical leadership through mentorship, code reviews, and collaboration across engineering teams.
Our partner is building next-generation AI-assisted software development experiences. This remote-first role contributes to a globally impactful AI platform with a collaborative culture focused on innovation, inclusion, and technical excellence.
Improve core inference services including networking, speech processing, audio transcoding, and latency optimization.
Develop processes for measuring, building, and optimizing services to maximize system performance.
Debug complex system issues involving networking, scheduling, and high performance computing interactions.
Deepgram is the leading platform for the Voice AI economy, providing real-time APIs for speech-to-text, text-to-speech, and voice agents. With over 200,000 developers and 1,300+ organizations, including major partners like Twilio and Cloudflare, Deepgram emphasizes an AI-first mindset and rapid innovation.
Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.