Act as the technical authority on clinical data structures, standardizing and harmonizing data pipelines with global partners.
Write production-grade Python code to automate data cleaning, enforce quality control, and design data dictionaries for AI models.
Map clinical data to industry-standard ontologies (e.g., SNOMED, ICD) with a focus on oncology and immunology datasets.
Bioptimus is building the first universal AI foundation model for biology to accelerate innovation in biomedicine. With over $75M in funding, it is a fast-growing startup headquartered in Paris with a world-class team of scientists and engineers.
Capture customer manufacturing processes and translate them into structured data.
Build and maintain data infrastructure, dashboards, and reports using SQL and Python.
Analyze data to provide product insights and recommendations for robotics capabilities.
Sunrise Robotics builds intelligent robots to enhance manufacturing, reducing waste and cost. They are a remote-first team focused on creating autonomous agents for the future.
Design, build, and operate data pipelines processing terabytes of transactional data daily using Airflow, BigQuery, and GCP services.
Own end-to-end data models and transformations powering merchant analytics, operational reporting, and ML features.
Improve data quality, lineage, and observability through testing, alerting, and validation frameworks.
Narvar is building the data infrastructure behind the post-purchase experiences of hundreds of millions of consumers, powering analytics, ML, and merchant-facing products for over 1,500 brand partners. The company serves 125+ million consumers worldwide across 38 countries and 55 languages, fostering a culture of innovation, collaboration, and inclusivity.
Contribute to the architecture and improvement of our data acquisition and processing platform, increasing reliability, throughput, and observability.
Use and develop web crawling technologies to capture and catalog data on the internet.
Build, operate, and evolve large-scale distributed systems that collect, process, and deliver data from across the web.
People Data Labs provides people and company data, integrating thousands of compliantly sourced datasets into a single source of truth. They are a remote-first company focused on building innovative data solutions, with a collaborative team that values ownership and learning from failure.
Design, build, and maintain robust data ingestion and transformation pipelines for knowledge graph analytics platforms.
Profile source data and implement complex transformation and normalization logic for graph loading.
Collaborate with data scientists and engineers to align data processing with graph modeling and entity resolution requirements.
Redhorse Corporation delivers data insights and technology solutions to customers with missions critical to U.S. national interests. They are a solution-driven company that values thoughtful, skilled professionals who thrive as trusted partners building technology-agnostic solutions.
Review, create, and refine chemistry content to ensure scientific accuracy for AI training datasets.
Develop clear explanations of chemical principles, reactions, and laboratory processes across varying complexity levels.
Provide structured feedback and collaborate with cross-functional teams to improve chemistry content and guidelines.
Our partner is a company at the intersection of scientific expertise and AI development. This is a remote, collaborative, and detail-oriented environment with a focus on precision and consistency.
Design, build, and own data pipelines moving data from application and third-party sources into relational databases.
Build and ship production AI systems, including data infrastructure and evaluation systems for feature reliability.
Set standards for data work through clear writing, early context sharing, and team efficiency.
Dscout builds a flexible UX research platform trusted by top brands in finance, healthcare, and tech. They are a remote-first team of passionate professionals that prioritize learning, diversity, and inclusion.
Design, develop, and maintain scalable data pipelines that transform complex customer data into actionable business value.
Work across the full data lifecycle, from ingestion and transformation to activation and analytics.
Use modern cloud technologies and AI-assisted development workflows to create reliable, high-performing data solutions.
Jobgether uses an AI-powered matching process to help candidates find roles and connects top-fitting applicants directly with hiring companies. The company fosters a remote, collaborative environment with a focus on efficient recruitment.
Own end-to-end design and reliability of large-scale data acquisition systems using AI and LLMs for self-healing pipelines.
Build and maintain data serving layers, ETL/ELT pipelines, and reporting systems for real-time insights.
Collaborate with engineering and product leadership to shape AI-native data infrastructure.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies by using automated screening to ensure fair and objective application reviews. They operate in a fast-paced, high-growth environment combining startup innovation with enterprise stability.
Build an application that connects to legacy CRM systems, extracts data, and prepares it for migration into Dwelly's core platform.
Design workflows for cleaning, mapping, and validating complex CRM data, and build safe import flows with retries and error handling.
Own the platform end-to-end from architecture to production support, working closely with engineering, product, and operations teams.
Dwelly is a UK-based, AI-enabled lettings and property management platform that acquires estate agencies and integrates them into a tech-enabled digital platform. It is a fast-growing, product-focused company backed by top-tier investors with a team experienced in real estate, technology, and operations.
Define data architecture and platform strategy for enterprise data platforms.
Build and optimize scalable data pipelines supporting batch and real-time processing.
Mentor junior engineers and partner with leadership on data strategy.
Robots & Pencils is an applied AI engineering firm building AI co-workers for enterprise operations. Founded in 2009, with delivery centers across North America and Europe, our teams average 15+ years of experience and we value craft, speed, and ownership.
Own production data products that directly power business decisions, from marketing automation to core semantic models.
Design and analyze experiments with Product and Growth teams to define metrics and recommend next steps.
Solve ambiguous business problems by exploring data, identifying opportunities, and translating insights into impact.
Smallpdf makes document management simple for everyone, from PDF conversion to e-signing. With over 1.1 billion users globally and 17 million monthly users, it's one of Europe's largest product-led tech companies, fostering a culture of growth and collaboration.
Perform data imports from various sources using Talend ETL tool, Python scripts, and public/private APIs.
Build, improve, and optimize data services processes and integrations between web applications.
Work closely with the Professional Services team to address customer requests and become an expert in your area.
SmartRecruiters provides a recruitment platform that powers superhuman hiring with next-gen AI functionality. It serves 4,000 customers including Bosch, LinkedIn, and Visa, and is a values-driven, globally focused tech employer with strong financial backing, recognized for its culture and benefits.
Own the design and scaling of the database (Postgres/Supabase) that powers our matching, outreach, and agents.
Ensure data quality through validation, dedup, entity resolution, and monitoring.
Build ingestion and enrichment pipelines for funding rounds, market news, and company research.
We are a platform where founders come to raise capital, pairing them with investors from a global network and running warm outreach. We are a profitable, self-funded company with a small, senior, flat team that ships fast.
Design, build, and maintain scalable data pipelines for geospatial data.
Develop and drive data quality standards, including validation and monitoring.
Collaborate with engineers and scientists to understand requirements and deliver solutions.
Vibrant Planet provides a cloud-based, AI-driven platform for wildfire risk management. They are a startup backed by climate leaders, fostering a collaborative and innovative engineering culture.
Design and optimize large-scale ETL pipelines using Python, PySpark, SQL, DBT, and cloud-based data platforms.
Define technical vision and architecture for data integration solutions, ensuring scalability and reliability.
Lead technical initiatives, mentor engineers, and collaborate with cross-functional teams to deliver high-quality data solutions.
Jobgether uses AI-powered matching to connect candidates with job opportunities at partner companies. They focus on efficient, objective recruitment processes for a fast-growing remote-first environment.
Design, implement, and maintain ETL/ELT pipelines for clean, scalable data flows across multiple systems.
Own data warehouse architecture, including partitioning, access controls, and security in partnership with Security and Infrastructure teams.
Champion data quality and governance, building dashboards and analytics to deliver actionable insights for leadership and cross-functional teams.
CertifID protects life's largest transactions from fraud, helping title companies, law firms, lenders, and consumers safeguard billions of dollars from wire fraud daily. They are a growth-minded team passionate about securing sensitive data and transforming fraud prevention.
Manage and enrich data from internal and third-party sources, turning fragmented data into clean, usable views.
Build automations that put data to work for sales, growth, and account management teams.
Partner with vertical stakeholders to strategize outreach and consolidate contact data into a single reliable view.
A fast-moving Supply Technology team is scaling how it acquires, enriches, and acts on data across multiple supply verticals including hotels, homes, services, and experiences. This is a contract role on a team that values practical, business-facing implementation.
Ensure timely delivery of customer data and generate standard reports for all projects.
Onboard customers to the data portal and provide ongoing support.
Collaborate with the Data team for troubleshooting and build custom analysis pipelines.
Nomic makes biology easier to measure with its nELISA proteomic platform, combining DNA nanotechnology, flow cytometry, automation, and machine learning. We are a diverse team of engineers, scientists, and problem-solvers who recently closed a $42M Series B and can process over 2.5 million samples annually.