Lead the development of Reddit's Ingestion Platform, designing and delivering reliable software for distributed data movement across streaming and batch workloads.
Own the architecture of the platform's control and data planes, including pipeline APIs, connectors, and sink integrations, expanding beyond Kafka-to-BigQuery to S3/GCS and Apache Iceberg.
Mentor engineers, drive migrations from legacy systems, and establish robust reliability, security, and operational practices for pipelines running on Kubernetes.
Reddit is a community of communities built on shared interests, passion, and trust, hosting authentic conversations across 100,000+ active communities and approximately 130 million daily active unique visitors. The company fosters an open, collaborative culture with a focus on reliability, performance, and efficiency.
Design and build the next generation big data compute platform for ETL, analytics, and machine learning at Airbnb.
Operate, manage, and improve the reliability, performance, observability, and cost efficiency of the data platform.
Write maintainable, self-documenting code, perform code reviews, and contribute to open source software.
Airbnb is a global home-sharing platform that connects travelers with unique accommodations and experiences. The company has grown to over 5 million hosts and 2 billion guest arrivals, fostering a culture of belonging and innovation.
Partner with cross-functional teams to design and deliver scalable backend systems for major product initiatives.
Own the full software lifecycle from technical design to rollout, using A/B experiments and data analysis to drive decisions.
Build and maintain high-performance APIs and distributed services using modern languages and tools.
Reddit is a community of communities, built on shared interests, passion, and trust. It is home to the most open and authentic conversations on the internet, with 100,000+ active communities and approximately 130 million daily active unique visitors.
Design and build distributed data systems handling large-scale ingestion and processing.
Drive architectural decisions and take end-to-end ownership of critical components.
Collaborate with product teams to translate ambiguous requirements into robust technical solutions.
Our partner builds a large-scale, multi-chain data platform that ingests, models, and delivers blockchain data to users and developers. They are a remote-first, distributed team with a strong engineering culture focused on ownership and collaboration.
Design and build reliable data pipelines using Spark, Kafka, Iceberg, and Airflow across batch, streaming, and real-time workloads.
Contribute to the evolution of the data lake and platform, including ingestion, processing, storage, and serving patterns.
Define and improve data quality, observability, reliability, and governance standards across data systems.
Webflow is an agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. The company is a growing, privately held organization that values grit, speed, and craft.
Design, build, and improve systems to reliably and efficiently process data at the scale of tens of terabytes per day.
Evaluate new technologies and set best-practice standards with Data Engineering and Data Platform Engineering.
Drive cross-functional projects with stakeholders across Mozilla and mentor teammates to develop engineering best practices.
Mozilla Corporation is a non-profit-backed technology company that makes pioneering brands like Firefox, the privacy-minded web browser, and works on diverse areas including AI, social media, security. The company is wholly owned by the non-profit Mozilla Foundation, with over 225 million monthly users, and focuses on making the internet better for people.
Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it.
Collaborate daily with researchers and engineers who are some of the best in the world at what they do.
Build and maintain the high-performance data layer that Modeling teams rely on for training and evaluation jobs.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation AI models and end-to-end products. They are a global team of researchers, engineers, and designers passionate about their craft, with offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul.
Architect the Data Platform by designing internal SDKs and self-service frameworks for distributed engineering teams.
Own platform performance, optimizing the Databricks ecosystem for cost-effectiveness and scalability.
Drive data contracts and governance, implementing schema validation and security standards across the global ecosystem.
G-P offers a SaaS-based Global Employment Platform that enables companies to expand into over 180 countries quickly and efficiently. The company fosters a diverse, remote-first culture where innovation thrives and every contribution is valued.
Design, build, and operate the ZMP real-time events pipeline for high-volume event data.
Provide technical leadership, own architecture, and mentor teammates through code and design reviews.
Collaborate with product, data, and consumer teams to turn SLAs into working systems.
Zeta Global is an AI-powered marketing cloud that leverages artificial intelligence and trillions of consumer signals to help marketers acquire, grow, and retain customers more efficiently. Founded in 2007 and headquartered in New York City with offices worldwide, the company fosters a culture of trust and belonging with a focus on diversity and inclusion.
Design and optimize scalable data platforms, pipelines, and governance frameworks.
Lead data strategy and collaborate with leadership, engineering, and AI teams.
Mentor engineers and drive platform modernization with an AI-forward mindset.
Robots & Pencils is an applied AI engineering firm building AI co-workers for enterprise operations. Founded in 2009 with delivery centers in Canada, US, Eastern Europe, and Latin America, the company is a nimble alternative to traditional system integrators with teams averaging 15+ years of experience.
Build and maintain reliable data pipelines and ETL/ELT workflows.
Develop and optimize data models for analytics and internal tools.
Support core data platform tools like Spark and AWS, and monitor pipeline quality and performance.
Sonatype is the software supply chain security company, providing end-to-end software supply chain security solutions. Over 2,000 organizations, including 70% of the Fortune 100, and 15 million software developers rely on Sonatype to optimize their software supply chains.
Own the technical strategy for the end-to-end Ads ML engineer lifecycle, focusing on feature development and training iteration.
Define architecture and technical standards for ML feature and training-data systems across batch/streaming computation, backfills, and quality.
Build platform abstractions and workflow automation to make ML development faster, safer, and more self-service.
Reddit is a community of communities built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, it is one of the internet's largest sources of information and has a culture of openness and authenticity.
Translate legacy Databricks notebook logic into modern ELT patterns using Python and PySpark, ensuring data contract preservation and reverse-view strategies.
Build scalable ingestion pipelines with YAML configurations and orchestrate automated DAG generation via workflow schedulers, applying quality assertions and migration validation.
Collaborate with business Data Stewards to align dependencies, negotiate refactoring scope, and validate migrated outputs for production systems.
CI&T helps large companies transform AI potential into real business impact with AI deployment, AI-native execution, and tech-integrated business solutions. With 30 years of tech transformation experience and over 8,000 CI&Ters in 25+ countries, we accelerate innovation through agentic SDLC, application modernization, Data & AI, martech, and business strategy.
Design, build, and maintain data pipelines using Airflow and Trino to ingest and process data from multiple systems.
Improve observability by implementing monitoring, validation, and alerting to ensure data accuracy and consistency.
Collaborate with cross-functional teams to define data requirements and deliver effective solutions.
SweedPos builds an all-in-one cannabis retail platform combining POS, eCommerce, Marketing, Analytics, and Inventory Management. With over 200 employees, they operate as a remote-first, product-driven startup focused on simplicity and innovation.
Enable efficient data access by creating and maintaining data pipelines.
Collaborate with ML engineers to design and maintain automation for machine learning training, quality assessment, and model release process.
Identify, design and implement improvements to optimize data delivery and automate manual processes.
Eneba is building an open, safe, and sustainable marketplace for gamers, supporting over 20 million active users. They are a fast-growing company that values data-driven decision-making and fosters a healthy data culture across the organization.
Build and maintain data ingestion and transformation pipelines using Dagster, dbt, and Snowflake through PR-driven development.
Manage production monitoring, backfills, and Airbyte connections to ensure reliable daily data flows.
Collaborate with analytics and product teams on dimensional modeling and data quality, and participate in code review and runbook writing.
PadSplit is at the forefront of solving the affordable housing crisis through its tech-driven platform. They have a growing remote-first team and offer a competitive benefits package including unlimited PTO and paid parental leave.
Design, build, and evolve the core data platform infrastructure including distributed query engines and orchestration.
Own our lakehouse infrastructure as code using Terraform and Ansible on Kubernetes.
Build and maintain low-latency streaming and batch ingestion pipelines, and scale the BI landscape.
Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, and more. We are a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world.
Design and implement core components of the streaming platform including Kafka-compatible APIs and storage systems.
Solve complex distributed systems problems involving performance optimization and reliability.
Debug production issues, collaborate with customers, and contribute to open-source projects.
StreamNative, founded by the creators of Apache Pulsar, is redefining real-time data streaming with its URSA engine. The company is a growing startup focused on building a cloud-native streaming platform with a culture of engineering excellence.
Design and deliver complex data platform components or migration solutions on AWS using Databricks.
Troubleshoot performance, scalability, and reliability issues in cloud-native data environments.
Communicate technical topics clearly to stakeholders and produce high-quality documentation.
Caylent is an AI-first cloud services company that helps organizations turn ambitious ideas into meaningful business impact. As a fully remote global company with employees in Canada, the United States, and Latin America, they celebrate diverse cultures and foster a community of technological curiosity.
Own and drive impactful distributed systems problems end-to-end, from inception through production launch.
Collaborate with a strong engineering team to prioritize and solve the most important problems for the company.
Raise the quality bar while keeping systems reliable and operationally lean, and mentor fellow engineers.
StarTree is a cloud-based software company that enables businesses to derive advanced insights from real-time and historical data using Apache Pinot. The company was founded by the core engineering team behind Apache Pinot, has secured Series B funding, and was named one of The Information's 50 Most Promising Startups.