Own the full lifecycle of production machine learning models from handoff to deployment and monitoring.
Ensure model reliability through drift detection, performance monitoring, and incident response.
Extend the shared ML platform to make it stronger and cheaper for future models.
CarOnSale is the AI-powered platform for B2B used car trading in Europe, connecting over 40,000 buyers from 20+ countries. We are a team of engineers and data scientists building the operating system for the industry, with a culture of direct ownership and short decision paths.
Design, implement, and operate benchmark execution and evaluation harnesses for AI models and agentic workflows.
Develop evaluation methodologies that combine automated metrics with structured human subject matter expert judgment.
Produce defensible evaluation reports comparing candidate capabilities with current mission workflows, including documented limitations and failure modes.
OpenTeams helps enterprises and governments build AI they control, govern, and evolve themselves. Founded by NumPy and SciPy creator Travis Oliphant, the company is built by people with deep roots in the open-source ecosystem.
Conduct independent research in Generative AI, LLMs, NLP, and multimodal AI to design experiments and evaluate models.
Develop and implement LLM evaluation frameworks, analyze model performance, and identify data gaps for improvement.
Apply strong statistical and data science skills to clean, analyze, and interpret complex datasets for AI/ML research.
Innodata is a global data engineering company that enables the responsible advancement of artificial intelligence by providing data, evaluation frameworks, and human expertise. With a 36+ year legacy, the company is committed to delivering the highest quality data and outstanding outcomes for its customers.
Develop advanced ML models and agentic workflows to accelerate model development.
Use AI-assisted tools like Claude and Cursor to investigate model behavior and automate analysis.
Set technical direction, mentor engineers, and raise the bar for modeling rigor.
Reddit is a community of communities, built on shared interests, passion, and trust, and is home to the most open and authentic conversations on the internet. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information.
Own the ML strategy for dialogue systems, lead a team of 3 ML engineers, and drive LLM post-training and model adaptation.
Build evaluation layers, cut dialogue failure modes, and keep inference efficient on latency and cost.
Stay hands-on with prototypes, debugging agent traces, and reviewing team work.
Social Discovery Group (SDG) is a group of social discovery companies that solve problems of loneliness and isolation through social entertainment platforms. The company has an international team of digital nomads and has been recognized as a Great Place to Work winner and a top company for remote jobs.
Design end-to-end LLM evaluation plans for business scenarios such as dialogue and financial trading.
Build evaluation metric systems and rubrics to quantify model performance and drive improvement.
Lead construction of evaluation datasets, define dimensions, and ensure high-quality annotation standards.
Binance is a leading global blockchain ecosystem behind the world's largest cryptocurrency exchange by trading volume and registered users. Trusted by 300+ million people in 100+ countries, we offer a diverse range of digital-asset products and services.
Lead the development and optimization of machine learning models to detect and prevent AI security risks like prompt injection and jailbreaks.
Build reproducible training and evaluation pipelines on Reddit's ML platform, partnering with platform engineers to improve performance and reliability.
Set the technical vision and multi-quarter modeling roadmap, mentoring engineers and establishing best practices for responsible ML development.
Reddit is a community of communities, built on shared interests and authentic conversations, with 100,000+ active communities and 130 million daily active visitors. It is one of the internet's largest sources of information, fostering a culture of openness and trust.
Lead strategic data initiatives including advanced statistical analysis, machine learning modeling, and generative AI solutions.
Design and implement experimentation strategies (A/B testing, causal inference) to validate hypotheses and measure product impact.
Collaborate with multidisciplinary teams to translate business problems into actionable data solutions and communicate insights clearly.
CI&T is a global technology company that combines the disruptive power of artificial intelligence with human expertise to help large companies navigate technology and business changes. With 8,000 employees across more than 25 countries, they collaborate to build solutions with real impact.
Lead the development of analytical, statistical, and machine learning solutions for client use cases.
Translate business questions into structured analytical approaches, modeling strategies, and measurable outputs.
Apply AI methods including LLM workflows to support classification, pattern detection, and decision support.
Valtech is an experience innovation company that partners with leading brands to unlock new value in the digital world. With a global, borderless framework and a culture that fosters creativity, diversity, and autonomy, we empower our people to thrive and grow.
Drive innovation across the machine learning ecosystem and architect advanced ML solutions.
Mentor junior ML engineers and lead complex ML initiatives at scale.
Own models in production including deployment, monitoring, drift detection, and retraining.
Xsolla is a global commerce company providing tools and services to help video game developers fund, distribute, market, and monetize their games. Headquartered in Los Angeles, California, Xsolla has helped over 1,500 game developers grow their businesses worldwide.
Design, build, and deploy LLM-powered product features, including lab summaries and conversational agents.
Build backend services integrating LLMs and ML models, primarily using Python with exposure to Elixir.
Implement evaluation, monitoring, and CI/CD workflows for AI features, ensuring reliability and clinical relevance.
Fullscript is a health technology platform that helps practitioners deliver better care through clinical insights, lab interpretations, and patient analytics. With over 125,000 practitioners and 10 million patients, the company emphasizes a people-first culture, teamwork, and continuous learning in a remote-first environment.
Design and build AI agents and automation to solve real problems across engineering, product, and delivery.
Partner with stakeholders to identify high-leverage opportunities and deliver end-to-end solutions.
Stay current with LLM and agentic frameworks to drive innovation in healthcare technology.
HealtheDGE provides AI-powered operational infrastructure for health insurance companies, helping them modernize operations. The company is experiencing strong market momentum and invests in its people, offering a collaborative culture focused on innovation.
Design and ship ML systems for document understanding, extraction, classification, and scoring to power Copilot and Autopilot features.
Own the full lifecycle of ML solutions from problem framing to production monitoring, using user feedback and error analysis to drive continuous improvement.
Partner with product, engineering, and accounting experts to integrate ML naturally into user workflows and define quality metrics that reflect real value.
Pennylane is a fast-growing fintech building the financial operating system for French SMEs and accounting firms, with a vision to expand across Europe. The company has grown to 1,000 employees, raised €400 million from Sequoia, and maintains a remote-friendly culture with a 4.6/5 Glassdoor rating.
Design and implement LLM solutions including RAG, agents, and prompt orchestration.
Integrate AI provider APIs (OpenAI, Anthropic, Google) into production systems with focus on cost and latency.
Evaluate architectures and define metrics for model quality and production monitoring.
Redbee is an Argentine technology company with over 14 years of experience redefining digital products in the financial industry. They are passionate about co-creating innovative solutions and focus on delivering value quickly and with quality.
Own and extend the offline evaluation suite for AI products, building datasets and metrics.
Build online quality dashboards and close the production feedback loop by mining failure patterns.
Translate numbers into clear decisions for Product and domain experts.
Finom is a European tech startup developing an all-in-one financial B2B platform integrating banking, accounting, and invoicing for entrepreneurs. With over €115 million in Series C funding and a team dedicated to innovation, they foster a start-up culture that values bold ideas and swift implementation.
Own a multimodal ML work-stream from problem definition through deployment, translating clinical needs into clear ML objectives and building systems with transformers, self-supervised learning, and segmentation.
Design rigorous evaluations beyond offline metrics, partner with engineering to productionize models, and investigate failure modes like laterality errors and hallucination.
Communicate research findings clearly, mentor less experienced researchers, and contribute to the research roadmap by identifying promising approaches.
Rad AI is an AI-driven healthcare company revolutionizing radiology by saving time, reducing burnout, and improving patient care. With over $140M in funding, a valuation of $528M, and partnerships with thousands of radiologists, Rad AI is a fast-growing team of mission-driven innovators.
Design, build, and improve machine learning training and inference pipelines for AI-driven music experiences.
Apply machine learning and prompt engineering across complex ML pipelines to support large language model features.
Create evaluation frameworks with LLM-as-judge pipelines to measure quality and enable rapid iteration.
Spotify is a digital music service that provides access to millions of songs. With over 700 million monthly active users, Spotify fosters a culture of innovation and collaboration, prioritizing artist-first principles in its AI music lab.
Design and maintain reliable, low-latency ML APIs to integrate Safety AI model outputs into cloud applications.
Build scalable data pipelines for continuous model iteration, backtesting, and online evaluation.
Optimize model artifacts for production and monitor rollout health, ensuring predictable failure modes.
Samsara builds a Connected Operations Cloud that helps physical operations use IoT data to improve safety, efficiency, and sustainability. Samsara is a recently public company with an employee-led remote culture and a long-term focus.
Build, evaluate, and improve ML ranking models for private, team, and brand content search.
Drive relevance and recall improvements, focusing on enterprise user outcomes and emerging surfaces like agent search.
Investigate and resolve quality or relevance regressions in production search while communicating trade-offs clearly.
Canva is redefining how the world experiences design. With headquarters in Sydney and a second campus in Melbourne, the company fosters a flexible, inclusive culture and empowers teams to do their best work.
Architect, build, and optimize high-performance production LLM systems while maintaining a strong personal technical presence on the team.
Spearhead strategic technological changes and champion code refactoring efforts to keep the core codebase cutting-edge and performant.
Lead technical story breakdowns, architectural design, and mentor engineers across the department.
Appian provides AI automation for mission-critical work, automating complex processes in large enterprises and governments. With over 25 years of experience, the company is known for its reliability and scale, and fosters an inclusive culture with employee-led affinity groups.