Optimize machine learning inference systems for latency, throughput, and cost-efficiency.
Profile and troubleshoot GPU/CPU bottlenecks, implement advanced techniques like quantization and speculative decoding.
Collaborate with research and engineering teams to productionize new models and improve inference infrastructure.
The company is an AI-focused organization that develops advanced machine learning systems for production environments. It values technical excellence and experimentation, offering a flexible remote work environment.
Own end-to-end ML system execution including data pipelines, training workflows, evaluation systems, inference architecture, and deployment.
Fine-tune and adapt models using state-of-the-art methods such as LoRA, QLoRA, SFT, DPO, and distillation.
Architect scalable inference systems, balance latency, cost, and reliability, and deploy production-grade ML solutions.
Gina's Tech Jobs is a recruiting and staffing company that helps firms hire technical talent. They are a small agency focused on IT roles, fostering a high-trust, collaborative environment.
Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
Architect scalable inference systems and collaborate closely with application engineering.
This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.
Optimize production LLM serving with vLLM and SGLang to maximize throughput and minimize latency through batching and quantization.
Profile training runs to find bottlenecks and resolve them with attention implementations like FlashAttention on H200 and GB200 hardware.
Deploy and operate multiple models on shared GPU clusters with autoscaling, bin-packing, and efficient handling of mixed workloads.
Egen is a fast-growing technology company with a data-first mindset, partnering with clients on Google Cloud and Salesforce to drive action through data and insights. We are a team of dedicated engineers who thrive on solving tough problems and continually innovate to achieve fast, effective results.
Architect and maintain production high-traffic LLM serving systems.
Optimize throughput, latency, and cost for leading open-source LLMs.
Debug and optimize major inference engines like SGLang, vLLM, or TensorRT using PyTorch and CUDA.
We are building decentralized and confidential machine learning infrastructure to enable user-owned AI. Our team is focused on highly scalable and efficient infrastructure for open-source AI at a global scale, with a culture that values innovation and performance.
Design and maintain scalable ML infrastructure including data pipelines, training workflows, and model deployment systems.
Own end-to-end ML lifecycle operations, ensuring reliable delivery of models into production at scale.
Implement monitoring, telemetry, and feedback loops for ML models running across large-scale device fleets.
Our partner company develops ML systems for connected hardware products used by customers worldwide. They operate in a fast-paced, product-driven environment with a collaborative and technically ambitious culture focused on real-world ML impact.
Build and improve ML components across data, training, evaluation, and inference.
Implement evaluation and testing to understand model behavior.
Debug model issues, performance problems, and production incidents.
This company builds core ML components for large-scale production systems. They emphasize real-world learning, iteration, and collaboration with senior engineers.
Design and build production-grade ML inference infrastructure using frameworks like vLLM and Triton.
Optimize GPU utilization, memory efficiency, and model artifact storage for cost-effective performance.
Collaborate with infrastructure and AI teams to establish engineering best practices and scalable platform architecture.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It uses technology to review applications and share top candidate shortlists with employers, operating in a remote-first environment.
Lead the design and operation of production machine learning systems for batch and online use cases with a focus on reliability and scalability.
Build and improve ML lifecycle infrastructure including training pipelines, inference workflows, monitoring, and automation.
Partner with cross-functional teams to translate business problems into ML solutions and guide prototypes to robust production systems.
Included Health is a healthcare company delivering integrated virtual care and navigation, aiming to raise the standard of healthcare for everyone. They are a remote-first organization offering comprehensive benefits and fostering a culture of inclusion.
Build ML infrastructure for low-latency model deployment, distributed inference pipelines, and real-time telemetry.
Scale ranking systems by moving models from experimentation to production, optimizing latency and cost trade-offs.
Implement model CI/CD for automated versioning, canary releases, hot-swappable container rollouts, and zero-downtime rollbacks.
Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with infrastructure to run them in production at sub-10ms latency and enterprise scale. They are a small, highly technical, early-stage team focused on turning recent AI advances into production-grade systems.
Build and operate production-grade model serving infrastructure using vLLM, TGI, or Triton frameworks.
Design and implement auto-scaling, multi-model architectures, and intelligent request routing for ML inference.
Optimize GPU utilization, memory efficiency, and observability to ensure low-latency, cost-effective systems.
They are a distributed cloud infrastructure startup building AI-native cloud services with GPU-powered compute. The company is well-funded, fast-scaling, and operates in a remote-first environment with a focus on sustainability and decentralization.
Develop scalable ML pipelines across the full lifecycle and champion responsible AI for content understanding models and signals in production.
Provide technical leadership and mentorship to ML engineers and software engineers, setting technical standards and conducting design reviews.
Build evaluation and quality monitoring systems for content understanding signals using state-of-the-art LLM-as-judge practices.
Reddit is a community of communities built on shared interests, passion, and trust, home to the most open and authentic conversations on the internet. With 100,000+ active communities and approximately 126 million daily active unique visitors, it is one of the internet's largest sources of information.
Design, build, and ship ML models that power content generation and quality eval scoring for Canva's generated element and template library.
Own the full ML lifecycle — from data pipelines and training through to deployment, monitoring, and iteration.
Partner with Content Engine, CORE AI Research, AI Media, and Discovery teams to align ML work with the broader content strategy.
Canva is redefining how the world experiences design with its intuitive design platform. We serve hundreds of millions of users globally and foster a culture of flexibility, inclusion, and innovation.
Design and implement advanced knowledge distillation pipelines, including teacher-student approaches and multi-teacher architectures.
Run large-scale machine learning experiments to optimize model quality, latency, efficiency, and cost trade-offs.
Collaborate with research teams to transform emerging distillation techniques into reliable production-ready implementations.
Our partner is an innovative company focused on advancing the efficiency and scalability of next-generation machine learning systems. They offer a remote-friendly work environment with an async-first culture and a small, senior team combining research expertise and engineering excellence.
Design and scale production ML systems for LLM-based applications.
Build training and evaluation pipelines for continuous model improvement.
Fine-tune foundation models using modern adaptation techniques such as LoRA, QLoRA, SFT and DPO.
A1 is a new AI venture building the next generation of AI-native productivity applications, starting with an email agent that uses autonomous AI. Backed by an initial $100M investment, the company is a small, high-talent founding engineering team focused on solving challenging AI infrastructure problems.
Build, ship, and own product features end-to-end using cutting-edge AI/ML techniques.
Apply classical ML and LLM-based approaches like RAG, prompt engineering, and fine-tuning to enhance the audit and risk platform.
Collaborate with cross-functional teams in an Agile environment to deliver scalable, production-quality code.
Optro is a leading audit, risk, ESG, and InfoSec platform trusted by over 50% of the Fortune 500. The company has been named one of the 500 fastest-growing tech companies in North America for seven consecutive years, fostering a culture of innovation and collaboration.
Research and implement state-of-the-art techniques to accelerate AI inference: quantization, sparsity, distillation, speculative decoding, and caching.
Partner closely with hardware and compiler teams to ensure algorithmic improvements translate to real gains on custom silicon.
Build profiling tools and comprehensive benchmarking frameworks to measure model quality and efficiency.
EnCharge AI is building the next generation AI platform using novel in-memory-computing architecture. The team consists of experienced AI researchers, silicon & systems engineers, and architects backed by leading investors.
Lead a team to build, scale, and optimize the ML infrastructure powering drug discovery.
Collaborate with ML engineering, data science, and research teams to deliver scalable solutions.
Mentor and coach team members in MLOps, distributed computing, and infrastructure engineering.
Recursion is a clinical-stage TechBio company decoding biology to develop medicines. With a focus on AI and machine learning, the company fosters a culture of bold integrity and cross-functional collaboration.
Build and maintain scalable machine learning solutions in production.
Train and validate deep learning and statistical models for real-world applications.
Partner with product managers and engineers to define requirements and drive ML roadmap.
Twilio is a cloud communications platform that empowers businesses to build personalized customer experiences through APIs. With thousands of employees worldwide, the company champions a remote-first culture focused on inclusion and innovation.