Lead the design and operation of production machine learning systems for batch and online use cases with a focus on reliability and scalability.
Build and improve ML lifecycle infrastructure including training pipelines, inference workflows, monitoring, and automation.
Partner with cross-functional teams to translate business problems into ML solutions and guide prototypes to robust production systems.
Included Health is a healthcare company delivering integrated virtual care and navigation, aiming to raise the standard of healthcare for everyone. They are a remote-first organization offering comprehensive benefits and fostering a culture of inclusion.
Design, deploy, and maintain scalable ML infrastructure for model training, batch processing, and real-time inference.
Build and manage cloud-based infrastructure with AWS and Snowflake using Infrastructure-as-Code practices.
Develop CI/CD pipelines, automation frameworks, and monitoring for ML systems to improve reliability and governance.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They operate with a team-oriented culture and offer remote work flexibility, focusing on efficient, unbiased recruitment.
Build ML infrastructure for low-latency model deployment, distributed inference pipelines, and real-time telemetry.
Scale ranking systems by moving models from experimentation to production, optimizing latency and cost trade-offs.
Implement model CI/CD for automated versioning, canary releases, hot-swappable container rollouts, and zero-downtime rollbacks.
Sequen provides an integrated platform that pairs cutting-edge frontier ranking models with infrastructure to run them in production at sub-10ms latency and enterprise scale. They are a small, highly technical, early-stage team focused on turning recent AI advances into production-grade systems.
Lead the development and improvement of MLOps platform capabilities supporting machine learning workflows across Oura.
Drive design and unification of workflows and tooling for reliable training, orchestration, and deployment of ML systems.
Partner with data scientists and engineers to improve the end-to-end ML lifecycle, including governance and maintainability.
Oura empowers people to own their inner potential through its award-winning Oura Ring and app, providing daily health insights. As a quickly growing company, we focus on team well-being and helping people live healthier, happier lives.
Design, deploy, and optimize production-grade machine learning systems for the full ML lifecycle.
Build scalable MLOps platforms, CI/CD workflows, and model serving infrastructure.
Collaborate with engineering teams to improve platform scalability, security, and operational best practices.
They build scalable MLOps infrastructure for enterprise AI solutions. They foster a collaborative, remote-first culture focused on innovation and professional growth.
Own ML models across their full lifecycle from data pipelines to deployment and monitoring, ensuring reliable performance.
Run and improve the ML platform including GitOps CI/CD, monitoring serving endpoints, and defining SLOs.
Collaborate with risk, operational, and product teams to turn ML into business value across the organization.
Alma provides installment and deferred payment solutions to help merchants boost sales and customer loyalty, without encouraging bad debt. With over 25,000 merchants, 10 million consumers, 380+ employees, and over €100M ARR, they are a Next40 member scaling rapidly across Europe.
Design, build, and deploy production-grade machine learning and AI systems for customer-facing analytics and automation.
Develop and operationalize end-to-end ML workflows from data preparation to model monitoring.
Collaborate with Product and Engineering teams to identify high-impact use cases and deliver AI-powered data products.
Boulevard provides a client experience platform for appointment-based self-care businesses, empowering customers to give clients magical moments. The company values diversity, experimentation, and simplicity, and celebrates diverse backgrounds.
Design, build, and deploy machine learning models for cybersecurity use cases like threat detection and risk scoring.
Own the full model lifecycle from data preparation to production deployment, working closely with engineering and product teams.
Build preprocessing and feature engineering pipelines, and monitor model performance with continuous improvement.
SpyCloud transforms recaptured darknet data to disrupt cybercrime. With over 250 employees, it is home to cybersecurity experts protecting businesses and consumers from stolen identity data.
Design and implement scalable real-time data integration and Change Data Capture (CDC) solutions using the Striim platform.
Build proof-of-concepts, reference architectures, and deployment patterns for enterprise implementations.
Collaborate with Engineering, Product, and GTM teams to validate architectural designs and improve platform capabilities.
Striim is a unified data integration and streaming platform that connects clouds, data, and applications with real-time analytics for enterprise customers. It is a Silicon Valley startup with a culture fostering entrepreneurship and growth, operating as one team with unlimited potential and dignity.
Design, build, and operate large-scale ML solutions for ad delivery performance.
Collaborate with cross-functional teams to deploy models into production.
Mentor peers and contribute to ML best practices and AI-native development.
The company is a technology firm focused on AI-driven advertising solutions. It operates in a remote-first environment with a collaborative and inclusive culture.
Design, build, deploy, and optimize machine learning models that process large volumes of complex, unstructured data.
Develop and maintain scalable ML pipelines capable of supporting millions of documents and diverse customer requirements.
Lead technical initiatives from early experimentation through production implementation and ongoing improvement.
The company develops intelligent automation solutions using AI and machine learning for large-scale document processing. It offers a highly autonomous engineering environment with a focus on continuous improvement and innovation.
Design, train, and ship ML systems for governance and security like anomaly detection and trust scoring.
Build data pipelines, model serving, evaluation frameworks, and feedback loops.
Set technical direction, own architecture, and help recruit and mentor as the team grows.
Docker provides developer tooling trusted by over 20 million monthly users and billions of container pulls. They are a globally distributed, remote-first team building tools for software delivery.
Own productization of Alt's pricing and underwriting models from research through production, keeping them accurate and fast at scale.
Optimize pricing models to reduce infrastructure costs while improving accuracy, especially for high-value assets.
Lead the full ML lifecycle from model training and feature generation to production deployment and monitoring.
Alt unlocks the value of alternative assets, starting with the $5B trading-card market, offering a platform for collectors to buy, sell, vault, and finance their cards. Backed by leaders at Stripe, Coinbase, Seven Seven Six, and pro athletes, the company is at an inflection point with its pricing intelligence infrastructure as a competitive moat.
You will experiment with emerging technologies and contribute to building new models and systems.
You will implement prototypes in Python and focus on delivering solutions to production.
You will partner with the platform engineering team to streamline MLOps workflows and maintain high code quality.
Verve creates a more efficient and privacy-focused way to buy and monetize advertising by fusing data, media, and technology. With 30 offices globally, they serve top advertisers and publishers and foster a collaborative, fun culture.
Lead a team of platform engineers to build and operate ML training and serving infrastructure, including GPU and low-latency serving.
Combine strong people leadership with technical judgment in ML infrastructure, partnering with senior ICs and cross-functional teams.
Drive delivery, operational health, and evaluate modern ML tooling to support company-wide ML priorities.
Affirm is reinventing credit to make it more honest and friendly, offering buy now pay later solutions without hidden fees. It is a remote-first company with a strong engineering culture, prioritizing people and providing competitive benefits.
Lead a team to build, scale, and optimize the ML infrastructure powering drug discovery.
Collaborate with ML engineering, data science, and research teams to deliver scalable solutions.
Mentor and coach team members in MLOps, distributed computing, and infrastructure engineering.
Recursion is a clinical-stage TechBio company decoding biology to develop medicines. With a focus on AI and machine learning, the company fosters a culture of bold integrity and cross-functional collaboration.
Independently own high-value optimization initiatives across training, inference, or launch-readiness for important Ads ML workloads.
Diagnose bottlenecks in real production systems using profiling, benchmarking, and observability.
Build performance tooling, optimization playbooks, and efficiency primitives that benefit multiple teams.
Reddit is a community of communities built on shared interests and authentic conversations. With 100,000+ active communities and approximately 126 million daily active unique visitors, Reddit has a flexible workforce and values collaboration.
Own end-to-end Machine Learning (ML) system execution including data pipelines, training, and deployment.
Fine-tune and adapt models using state-of-the-art methods like LoRA and DPO.
Architect scalable inference systems and collaborate closely with application engineering.
This company develops advanced production-grade machine learning systems. The team is small and high-trust, with a culture of ownership and pragmatism.
Build a cloud-native big data platform handling audience data for millions of attendees and billions of interactions.
Design and own ML infrastructure including feature stores, training pipelines, and model serving.
Own the full data pipeline from ingestion to business impact, ensuring reliability and performance.
Hive is a marketing platform that helps event marketers personalize and automate campaigns to sell out shows and engage fans. The company is a fully remote team located across Canada, fostering a work environment with strong work-life balance and a focus on impactful outcomes.
Develop, deploy, and maintain machine learning models in production environments using AWS SageMaker or similar platforms.
Perform exploratory data analysis, feature engineering, and build data pipelines to support scalable ML workflows.
Monitor production models, address performance issues, and collaborate with engineering and business stakeholders to deliver data-driven solutions.
Google is a global technology company that develops products and services to organize information and make it universally accessible and useful. The company is a large multinational with a culture focused on innovation, collaboration, and data-driven decision-making.