Build, expand, and optimize data infrastructure to create the most accurate dataset of identities and their relationships.
Develop and operate secure, scalable, and reliable data ingestion and ETL/ELT pipelines that meet product requirements.
Design and maintain a data observability framework to ensure data meets strict quality and freshness standards.
SentiLink provides innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. The company is growing rapidly, has verified hundreds of millions of identities, and is backed by top investors like Craft Ventures and Andreessen Horowitz, with offices across the US and India.
Contribute to the design, development, and operation of core components of Abnormal’s data platform.
Build tools and services that make it easy for other teams to adopt and scale data systems.
Help automate infrastructure and operations to improve reliability, performance, and scalability.
Abnormal protects the humans behind the world's most critical organizations from AI-powered cybercrime. 4,500+ enterprises trust our behavioral AI platform.
Build the data control plane for governance and discovery.
Create paved roads for data products with SDKs and CI/CD.
Ensure every data product is observable by default.
Kargo creates powerful moments of connection between brands and consumers to build businesses. With 600+ employees, the company uses a creative science approach to deliver innovative ad experiences across CTV, eCommerce, social, and mobile.
Contribute to the architecture and improvement of our data acquisition and processing platform, increasing reliability, throughput, and observability.
Use and develop web crawling technologies to capture and catalog data on the internet.
Build, operate, and evolve large-scale distributed systems that collect, process, and deliver data from across the web.
People Data Labs provides people and company data, integrating thousands of compliantly sourced datasets into a single source of truth. They are a remote-first company focused on building innovative data solutions, with a collaborative team that values ownership and learning from failure.
Design, build, and maintain data infrastructure and pipelines spanning both batch and real-time workloads.
Build and maintain scalable ETL/ELT pipelines using Python, SQL, Spark, and orchestration frameworks.
Partner with data scientists, ML engineers, and product teams to deliver data products and establish data quality frameworks.
Gemini is a global crypto and Web3 platform founded by Cameron and Tyler Winklevoss in 2014, offering a wide range of crypto products to individuals and institutions in over 70 countries. As a publicly traded company, Gemini is poised to accelerate the vision of reshaping the global financial system with greater scale, reach, and impact.
You will develop and maintain end-to-end data pipelines and contribute to Samsara's Data Platform for advanced automation and analytics.
You will design, build, and optimize large-scale Spark and PySpark workflows for batch and streaming data processing.
You will build and maintain MCP servers and AI agents, and champion data engineering best practices across the team.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for actionable insights. As a recently public company, they foster a culture of autonomy, support, and rapid career development in a hyper-growth environment.
Design and build batch data pipelines that ingest, validate, and transform multi-billion-row datasets.
Model complex real-world data including dimensional models and temporal data.
Develop and operate workloads on lakehouse platforms like Databricks, Spark, and Delta.
Simulmedia builds an advanced TV and streaming advertising platform. They have a team of engineers, data scientists, and designers who are obsessed with building cutting-edge technology.
Contribute to developing the team and organization’s long term technical strategy.
Refine and maintain our messaging infrastructure to support asynchronous communication for hundreds of millions of users.
Own the infrastructure (managed and self-hosted) that supports produce and consume requests along with necessary tooling and automation.
Reddit is a community of communities built on shared interests, passion, and trust, home to the most open and authentic conversations on the internet. With 100,000+ active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet’s largest sources of information.
Deliver on business-critical outcomes by owning backend and data services end-to-end, from design to production operation.
Design, build, and operate scalable data pipelines and streaming ingestion for high-volume security telemetry.
Contribute to data modeling and warehouse/lakehouse architecture decisions that serve detection, analytics, and product features.
Blackpoint Cyber is the leading provider of world-class cybersecurity threat hunting, detection and remediation technology. Founded by former National Security Agency (NSA) cyber operations experts, the company is in hyper-growth mode, fueled by a recent $190m series C round.
Lead the modernization of the data platform by migrating legacy Hadoop, Spark, and Impala pipelines to a scalable Databricks architecture.
Accelerate migration efforts using AI coding assistants like Claude and Codex to convert SQL and modernize ETL workflows.
Design and optimize scalable data pipelines with Airflow and Databricks, ensuring reliability, cost-efficiency, and data quality.
LivePerson is a leader in enterprise conversational AI and digital transformation, powering nearly a billion conversational interactions monthly. Fast Company named them the #1 Most Innovative AI Company, and they foster a remote-first, innovative culture.
Design and develop distributed systems for a data platform handling petabytes of data.
Take ownership of components like data ingestion and decoding.
Write code in Kotlin and Go with a focus on good design and performance.
Dune is a collaborative multi-chain analytics platform that makes crypto data accessible. They are a team of ~60 employees working across Europe and eastern US timezones, backed by top investors.
Architect, design, build, deploy, and maintain Model Serving infrastructure using industry-standard AI tools.
Own projects that scale model serving and data processing services to handle 10x traffic.
Collaborate closely with MLE and Data Science teams to distill feedback and execute on strategy.
Abnormal AI protects the humans behind the world's most critical organizations from AI-powered cybercrime. Over 4,500 enterprises trust their behavioral AI platform, fostering a culture of innovation and security.
Design, develop, and maintain scalable ETL/ELT pipelines
Build and optimize data processing solutions using Python and Apache Spark
Develop and support real-time data streaming applications using Kafka
Sigma Software develops modern digital solutions for global businesses. They foster an international remote-first environment where engineers can grow, innovate, and influence technical decisions.
Define and drive the long-term data engineering strategy for the Subscriptions User Understanding domain.
Design scalable batch and streaming data platforms that power analytics, experimentation, machine learning, and subscriber experiences.
Partner with Product, Engineering, Data Science, Analytics, and Platform teams to turn complex business challenges into durable, well-designed data solutions.
Spotify is a global music streaming service with a mission to reach one billion users and $100 billion in revenue. The Subscriptions User Understanding team supports over 200 million Premium subscribers, fostering a culture of inclusivity and forward-thinking innovation.
Build a cloud-native big data platform handling audience data for millions of attendees and billions of interactions.
Design and own ML infrastructure including feature stores, training pipelines, and model serving.
Own the full data pipeline from ingestion to business impact, ensuring reliability and performance.
Hive is a marketing platform that helps event marketers personalize and automate campaigns to sell out shows and engage fans. The company is a fully remote team located across Canada, fostering a work environment with strong work-life balance and a focus on impactful outcomes.
Design and develop Privacy APIs and backend infrastructure for large-scale data and privacy workflows.
Own integration with third-party and internal data platforms (e.g., data warehouses, REST/GraphQL services, message queues).
Ensure system reliability, performance, and maintainability.
Skyflow secures the flow of data across datastores, models, and agents. Skyflow is trusted by Fortune 500 enterprises and headquartered in Palo Alto, California, founded in 2019.
Enable efficient data access by creating and maintaining data pipelines.
Collaborate with ML engineers to design and maintain automation for machine learning training, quality assessment, and model release.
Build data infrastructure for analytics, hypothesis testing, and company metrics.
Eneba is building an open, safe, and sustainable marketplace for gamers, supporting close to 20 million active users. We are a growing international team that values data-driven decision making and fosters a healthy data culture.
Design and develop automated ETL/ELT pipelines to ingest data into Google Cloud Platform.
Implement Medallion Architecture patterns and maintain governed views in BigQuery.
Balance new solution development with production incident resolution and support.
Our partner is a company seeking a senior data professional to design, build, and operate scalable data solutions in a large-scale corporate environment. They collaborate across multiple business areas and combine development with operational support.
Design, build, and maintain Python services for sports data ingestion, transformation, and distribution.
Integrate with third-party sports data providers and handle differences between provider models, formats, and update patterns.
Build reliable pipelines for near real-time and batch data processing.
Fliff builds social sports gaming experiences that allow users to compete for leaderboard positioning, achieve badges, and build status. The company is a multinational team with offices in the US and Bulgaria, known for being a close-knit, focused, and welcoming group.
Lead and nurture a team of high-impact engineers, providing technical and cultural guidance.
Drive projects aligned with goals and KPIs, ensuring quality and timely delivery of data infrastructure.
Engineer and optimize ETL pipelines and architect data models for petabyte-scale blockchain data processing.
TRM Labs provides AI-powered intelligence solutions that help public and private sector agencies investigate and disrupt crime. It is a Series C company with $220M in total funding, backed by Goldman Sachs, Bessemer, Y Combinator, Thoma Bravo, and others, operating as a distributed-first company with hubs in multiple cities.