Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.
Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.
Architect and maintain production high-traffic LLM serving systems.
Optimize throughput, latency, and cost for leading open-source LLMs.
Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.
We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.
Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.
Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.
Optimize production LLM serving with vLLM and SGLang to maximize throughput and minimize latency through batching and quantization.
Profile training runs to find bottlenecks and resolve them with attention implementations like FlashAttention on H200 and GB200 hardware.
Deploy and operate multiple models on shared GPU clusters with autoscaling, bin-packing, and efficient handling of mixed workloads.
Egen is a fast-growing technology company with a data-first mindset, partnering with clients on Google Cloud and Salesforce to drive action through data and insights. We are a team of dedicated engineers who thrive on solving tough problems and continually innovate to achieve fast, effective results.
Develop and optimize low-level kernels, runtime components, and system software for high-performance AI inference workloads.
Improve inference engine performance across GPU platforms by identifying bottlenecks and implementing advanced optimization techniques.
Profile, debug, and resolve system-level and hardware-level performance issues across CPU and GPU environments.
This position is listed on behalf of a partner company building cutting-edge AI infrastructure for large-scale inference platforms. They operate in a highly technical, international, and innovation-driven environment where engineering excellence and ownership are valued.
Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.
Research and implement state-of-the-art techniques to accelerate AI inference: quantization, sparsity, distillation, speculative decoding, and caching.
Partner closely with hardware and compiler teams to ensure algorithmic improvements translate to real gains on custom silicon.
Build profiling tools and comprehensive benchmarking frameworks to measure model quality and efficiency.
EnCharge AI is building the next generation AI platform using novel in-memory-computing architecture. The team consists of experienced AI researchers, silicon & systems engineers, and architects backed by leading investors.
Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.
They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.
Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
Architect scalable inference systems and collaborate closely with application engineering.
This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.
Design and implement advanced knowledge distillation pipelines, including teacher-student approaches and multi-teacher architectures.
Run large-scale machine learning experiments to optimize model quality, latency, efficiency, and cost trade-offs.
Collaborate with research teams to transform emerging distillation techniques into reliable production-ready implementations.
Our partner is an innovative company focused on advancing the efficiency and scalability of next-generation machine learning systems. They offer a remote-friendly work environment with an async-first culture and a small, senior team combining research expertise and engineering excellence.
Build ML infrastructure for low-latency model deployment, distributed inference pipelines, and real-time telemetry.
Scale ranking systems by moving models from experimentation to production, optimizing latency and cost trade-offs.
Implement model CI/CD for automated versioning, canary releases, hot-swappable container rollouts, and zero-downtime rollbacks.
Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with infrastructure to run them in production at sub-10ms latency and enterprise scale. They are a small, highly technical, early-stage team focused on turning recent AI advances into production-grade systems.
Design and execute performance benchmarks for AI training and inference workloads.
Profile and characterize GPU workloads to identify bottlenecks and optimization opportunities.
Systematically tune workload parameters to maximize throughput and establish performance baselines across GPU platforms.
Vultr provides high-performance, affordable cloud infrastructure for enterprises and AI innovators with 33 global data centers. It is the world's largest privately-held cloud infrastructure company, valued at $3.5 billion, with a culture focused on growth and innovation.
Lead the design and development of our production inference platform, defining the technical roadmap for inference infrastructure, model serving, and runtime optimization.
Build and operate scalable, cost-effective systems for serving large language models in production, optimizing latency, throughput, GPU utilization, and memory efficiency.
Partner with ML engineers to productionize new models and inference techniques, establish benchmarking methodologies, and make key architectural decisions.
Syllo is on a mission to transform litigation with a unified platform that enables lawyers to safely harness AI. Since going to market, they have gained diverse enterprise customers including big law firms and corporations, and are quickly expanding.
Design and maintain scalable ML infrastructure including data pipelines, training workflows, and model deployment systems.
Own end-to-end ML lifecycle operations, ensuring reliable delivery of models into production at scale.
Implement monitoring, telemetry, and feedback loops for ML models running across large-scale device fleets.
Our partner company develops ML systems for connected hardware products used by customers worldwide. They operate in a fast-paced, product-driven environment with a collaborative and technically ambitious culture focused on real-world ML impact.
Design and scale production ML systems for LLM-based applications.
Build training and evaluation pipelines for continuous model improvement.
Fine-tune foundation models using modern adaptation techniques such as LoRA, QLoRA, SFT and DPO.
A1 is a new AI venture building the next generation of AI-native productivity applications, starting with an email agent that uses autonomous AI. Backed by an initial $100M investment, the company is a small, high-talent founding engineering team focused on solving challenging AI infrastructure problems.
Build and improve ML components across data, training, evaluation, and inference.
Implement evaluation and testing to understand model behavior.
Debug model issues, performance problems, and production incidents.
This company builds core ML components for large-scale production systems. They emphasize real-world learning, iteration, and collaboration with senior engineers.
Design, develop, and deploy production-grade AI-powered backend systems.
Integrate large language models and machine learning models into scalable architectures.
Optimize system performance and implement strong testing practices.
Our partner company is building advanced AI-powered systems to create meaningful customer value. The team operates in a high-autonomy, fast-moving environment focused on production-ready AI solutions.
Lead a team to build, scale, and optimize the ML infrastructure powering drug discovery.
Collaborate with ML engineering, data science, and research teams to deliver scalable solutions.
Mentor and coach team members in MLOps, distributed computing, and infrastructure engineering.
Recursion is a clinical-stage TechBio company decoding biology to develop medicines. With a focus on AI and machine learning, the company fosters a culture of bold integrity and cross-functional collaboration.
Lead and mentor a team of engineers specializing in GPU infrastructure, fostering a culture of technical excellence and continuous improvement.
Define and execute the technical roadmap for GPU cluster deployment, optimization, and scaling.
Partner with AI researchers to understand emerging infrastructure needs and translate them into robust solutions.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation models and end-to-end products for enterprises. It is a global technology company with a team of researchers, engineers, and designers passionate about their craft.