Build and maintain data pipelines for analytics, ML, and product applications.
Design scalable data infrastructure with a focus on quality and observability.
Collaborate with cross-functional teams to understand data needs and implement solutions.
Prolific builds human data infrastructure to power the next wave of AI innovation. They are a remote-first company focused on ethical data collection and mission-driven culture.
Own backend features end-to-end from discovery to rollout.
Design and evolve distributed systems with performance and scalability.
Build and maintain APIs and data pipelines for analytics and machine learning.
Traackr is a global SaaS company providing a data-driven influencer marketing platform. The company is remote-first with offices in multiple countries and a culture based on trust, diversity, value, ownership, and mutual success.
Architect and implement scalable ETL and data pipelines for real-time risk management and advanced analytics.
Design, develop, and optimize distributed data storage solutions for high performance and reliability at scale.
Drive schema evolution, data modeling, and pipeline orchestration with ownership of end-to-end data flow.
Oscilar builds the most advanced AI Risk Decisioning™ Platform for banks, fintechs, and digitally native organizations to manage fraud, credit, and compliance risk. The company is mission-driven with a remote-first culture and team members from Meta, Uber, Citi, and Confluent.
Work directly with external partners to understand their technical use cases, guide API integrations, and develop solutions that unlock value from Beacon's data platform.
Design and implement integrated solutions, including PoCs, dashboards, and production-ready code, to help partners succeed with analytics and data access.
Collaborate with Datastore engineers, product managers, and scientific teams to translate partner needs into robust data models, tools, and workflows.
Beacon Biosignals is revolutionizing precision medicine for the brain by providing the leading at-home EEG platform for clinical development of neurological, psychiatric, and sleep disorder therapeutics. With nearly 100,000 patients' EEG data and a cloud-native analytics platform, they enable quantitative biomarker discovery and implementation, fostering a diverse and empathetic team culture.
Design and build scalable backend services and APIs.
Develop integrations with healthcare and financial systems.
Build reliable, observable data ingestion and processing pipelines.
This healthcare AI startup builds products that help providers navigate reimbursement landscapes using machine learning and AI. It has assembled an exceptional technical team and is fast-growing.
Own the design and scaling of the database (Postgres/Supabase) that powers our matching, outreach, and agents.
Ensure data quality through validation, dedup, entity resolution, and monitoring.
Build ingestion and enrichment pipelines for funding rounds, market news, and company research.
We are a platform where founders come to raise capital, pairing them with investors from a global network and running warm outreach. We are a profitable, self-funded company with a small, senior, flat team that ships fast.
Build and maintain custom integrations and ETL processes between K-12 data and internal/external systems.
Design, implement, and optimize data pipelines to support school operations.
Collaborate with stakeholders to understand data needs and deliver robust, scalable solutions.
Newsela is a leading education technology company dedicated to meaningful classroom learning for every student. The company offers an inclusive culture with flexible PTO, comprehensive health benefits, and professional development allowances.
Design, build, and maintain backend services, REST APIs, databases, and big data pipelines that power customer-facing insights and analytics.
Implement and maintain near-real-time stream-based data processing pipelines in collaboration with batch-oriented data refresh workflows.
Scale data processing and insights generation pipelines to handle growing volumes of activity data while managing infrastructure costs.
Backstory helps companies understand the state of their revenue business by answering questions that span customer interactions, sales activity, pipeline health, and deal execution. Headquartered in San Francisco, CA, Backstory is backed by Y Combinator and top investors, and is listed in the top 20 percent of Inc 5000 companies.
Develop long-term technical vision and design scalable data systems.
Build and maintain production data pipelines using Python and integrate external APIs.
Mentor engineers and uphold standards for engineering excellence.
Correlation One is the largest provider of AI and data workforce development programs globally, having trained over 500,000 professionals across 11 countries. They work with Fortune 500 enterprises and government agencies to close skills gaps, and foster a culture of empowerment and diversity.
Design and implement scalable, low-latency real-time data streaming architectures and end-to-end data pipelines for high-volume systems.
Develop clean, maintainable, and efficient code using technologies such as Go, Python, TypeScript/Node.js, Scala, and SQL.
Build and optimize complex data processing workflows handling structured and unstructured data sources.
Our partner is looking for a Sr. Data Engineer - Data Analytics based in India. The company operates in a global, remote-first environment with a multicultural team spanning 45+ nationalities, focusing on work-life balance and employee well-being.
Design and build scalable data pipelines, clean room environments, and privacy-safe integrations for NBCUniversal’s data collaboration ecosystem.
Implement identity resolution logic and configure secure, role-based access controls across data platforms.
Optimize query performance and operational reliability, including monitoring, cost tracking, and incident response.
NBCUniversal is one of the world's leading media and entertainment companies, creating and distributing content across film, television, and streaming. A subsidiary of Comcast Corporation, it champions an inclusive culture and has a rich tradition of community service.
Design and implement a scalable enterprise data quality framework across data platforms and business domains.
Lead the implementation and operationalization of GX Core (Great Expectations) or similar data validation solutions.
Build automated validation checks for critical datasets, workflows, and operational processes.
The hiring company is a technology organization specializing in data engineering and quality solutions. They offer a remote work environment and emphasize collaboration across engineering teams.
Design and build scalable data pipelines using Python and SQL to ingest, transform, and curate data from internal and external sources.
Implement schema validation, data quality checks, and job monitoring to ensure trustworthy data outputs.
Collaborate with analytics engineers, architects, and product partners to define technical requirements and deliver data products.
Cohere Health provides a clinical intelligence platform that uses AI to connect health plans and providers, improving care quality and reducing costs. We are a growing company recognized as a top LinkedIn startup and backed by leading investors, fostering a supportive, growth-oriented culture.
Analyze and process large datasets to support business intelligence and product development.
Design and maintain data pipelines and infrastructure for efficient data flow.
Collaborate with cross-functional teams to implement data-driven solutions.
Tucows is a technology company that provides domain registration, internet services, and mobile connectivity solutions. It operates through its brands Ting, Wavelo, and Tucows Domains, and fosters a culture of inclusion with a focus on community and innovation.
Own data end to end, from raw source to analysis that drives decisions for growth, finance, and product.
Build and operate pipelines in BigQuery with Airflow, model in dbt, and deliver analysis in Hex.
Work AI-first, using modern tooling to move faster while staying accountable for production output.
Supabase provides a complete backend platform including Database, Auth, Storage, Edge Functions, Realtime, and Vector Search, built by developers for developers. With 280+ team members across 55+ countries and $500M raised, they are open-source-first and globally distributed.
Design, develop, and maintain scalable data pipelines that transform complex customer data into actionable business value.
Work across the full data lifecycle, from ingestion and transformation to activation and analytics.
Use modern cloud technologies and AI-assisted development workflows to create reliable, high-performing data solutions.
Jobgether uses an AI-powered matching process to help candidates find roles and connects top-fitting applicants directly with hiring companies. The company fosters a remote, collaborative environment with a focus on efficient recruitment.
Work cross-functionally with Product and experts to conceptualize, prototype, and build data solutions
Build and maintain data engineering systems and high-quality data models from multi-source healthcare datasets
Develop and test data pipelines and draft internal and external technical documentation
Turquoise Health is a Series C price transparency platform for finance leaders across healthcare. Backed by a16z, Oak HC/FT, and others, we are a remote-first, US-based team that values transparency, empathy, inclusivity, creativity, and ownership.
Design, build, and maintain scalable data and ML pipelines for analytics and AI systems.
Build and optimize workflows for structured and unstructured data, enabling semantic search and RAG use cases.
Manage and optimize vector databases and indexing strategies for efficient retrieval and AI-powered search.
This is a partner company seeking a Data & Machine Learning Engineer based in Brazil. They operate in a highly technical and global environment with strong emphasis on scalability, performance, and innovation.
Design, build, and maintain robust, scalable ELT/ETL data pipelines from various source systems into cloud data platforms.
Perform data modeling, including dimensional modeling, and build transformation layers using dbt to create analytics-ready datasets.
Support operational reliability, monitor data pipelines, and ensure SLAs for timeliness, freshness, and accuracy.
Troveo builds the data platform that AI labs and model builders need to train the next generation of models. Backed by top investors, we’re a small, high-impact team solving one of the biggest bottlenecks in AI development.
Design and deliver near real-time data solutions for the analytics platform.
Analyze business needs, optimize data models, and identify slow queries for performance improvement.
Write clean, scalable code using Scala, Python, and SQL while mentoring team members.
Aircall is an AI-powered customer communications platform used by 22,000+ companies worldwide. It is a unicorn startup with offices across multiple countries and a focus on innovation and collaboration.