Source Job

$212,000–$265,000/yr
US

  • Design and build the next generation big data compute platform for ETL, analytics, and machine learning at Airbnb.
  • Operate, manage, and improve the reliability, performance, observability, and cost efficiency of the data platform.
  • Write maintainable, self-documenting code, perform code reviews, and contribute to open source software.

Java Scala Spark Kubernetes

20 jobs similar to Staff Software Engineer, Data Warehouse

Jobs ranked by similarity.

Poland

  • Architect the Data Platform by designing internal SDKs and self-service frameworks for distributed engineering teams.
  • Own platform performance, optimizing the Databricks ecosystem for cost-effectiveness and scalability.
  • Drive data contracts and governance, implementing schema validation and security standards across the global ecosystem.

G-P offers a SaaS-based Global Employment Platform that enables companies to expand into over 180 countries quickly and efficiently. The company fosters a diverse, remote-first culture where innovation thrives and every contribution is valued.

$164,200–$229,900/yr
US

  • Refine and maintain data infrastructure for ML and analytics workflows on data from hundreds of millions of users.
  • Own the Data Movement Platform enabling batch and stream processing, investing in Spark, Flink, and Airflow technologies.
  • Build automated solutions to minimize toilsome work, providing a declarative, self-service experience for data users.

Reddit is a community of communities built on shared interests, passion, and trust, home to the most open and authentic conversations on the internet. With over 100,000 active communities and approximately 130 million daily active unique visitors, Reddit is one of the internet's largest sources of information, fostering a flexible and inclusive culture.

Colombia

  • Build and maintain reliable data pipelines and ETL/ELT workflows.
  • Develop and optimize data models for analytics and internal tools.
  • Support core data platform tools like Spark and AWS, and monitor pipeline quality and performance.

Sonatype is the software supply chain security company, providing end-to-end software supply chain security solutions. Over 2,000 organizations, including 70% of the Fortune 100, and 15 million software developers rely on Sonatype to optimize their software supply chains.

$140,000–$160,000/yr
Global Unlimited PTO

  • Design and ship data infrastructure at the core of modern data platforms, scalable pipelines, and reporting layers.
  • Work directly with the CEO and a small senior team, owning results from day one.
  • Build client-facing systems, translating business problems into architecture and delivering results.

Paradox Machines is a data and AI company that helps businesses make their data actually usable. They are a venture-backed company with a small, senior team that values ownership and impact.

$217,000–$303,900/yr
US

  • Lead the development of Reddit's Ingestion Platform, designing and delivering reliable software for distributed data movement across streaming and batch workloads.
  • Own the architecture of the platform's control and data planes, including pipeline APIs, connectors, and sink integrations, expanding beyond Kafka-to-BigQuery to S3/GCS and Apache Iceberg.
  • Mentor engineers, drive migrations from legacy systems, and establish robust reliability, security, and operational practices for pipelines running on Kubernetes.

Reddit is a community of communities built on shared interests, passion, and trust, hosting authentic conversations across 100,000+ active communities and approximately 130 million daily active unique visitors. The company fosters an open, collaborative culture with a focus on reliability, performance, and efficiency.

$95,460–$139,860/yr
Canada

  • Design, build, and improve systems to reliably and efficiently process data at the scale of tens of terabytes per day.
  • Evaluate new technologies and set best-practice standards with Data Engineering and Data Platform Engineering.
  • Drive cross-functional projects with stakeholders across Mozilla and mentor teammates to develop engineering best practices.

Mozilla Corporation is a non-profit-backed technology company that makes pioneering brands like Firefox, the privacy-minded web browser, and works on diverse areas including AI, social media, security. The company is wholly owned by the non-profit Mozilla Foundation, with over 225 million monthly users, and focuses on making the internet better for people.

$186,500–$255,000/yr
United States 18w maternity 12w paternity

  • Design and build reliable data pipelines using Spark, Kafka, Iceberg, and Airflow across batch, streaming, and real-time workloads.
  • Contribute to the evolution of the data lake and platform, including ingestion, processing, storage, and serving patterns.
  • Define and improve data quality, observability, reliability, and governance standards across data systems.

Webflow is an agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. The company is a growing, privately held organization that values grit, speed, and craft.

India

  • Own and drive impactful distributed systems problems end-to-end, from inception through production launch.
  • Collaborate with a strong engineering team to prioritize and solve the most important problems for the company.
  • Raise the quality bar while keeping systems reliable and operationally lean, and mentor fellow engineers.

StarTree is a cloud-based software company that enables businesses to derive advanced insights from real-time and historical data using Apache Pinot. The company was founded by the core engineering team behind Apache Pinot, has secured Series B funding, and was named one of The Information's 50 Most Promising Startups.

US 3w PTO

  • Support new and existing customers in their data engineering needs, guiding them to make optimal technical decisions.
  • Build and operationalize complex data solutions, including data governance, security, and quality.
  • Collaborate across teams to deliver data products and educate end users on analytic environments.

SunnyData is a high-growth consulting company specializing in data engineering and AI, dedicated to the Databricks platform. They foster a collaborative and innovative culture with a focus on customer impact and career growth.

Mexico

  • Design and deliver complex data platform components or migration solutions on AWS using Databricks.
  • Troubleshoot performance, scalability, and reliability issues in cloud-native data environments.
  • Communicate technical topics clearly to stakeholders and produce high-quality documentation.

Caylent is an AI-first cloud services company that helps organizations turn ambitious ideas into meaningful business impact. As a fully remote global company with employees in Canada, the United States, and Latin America, they celebrate diverse cultures and foster a community of technological curiosity.

$180,000–$220,000/yr
US

  • Lead system design and architecture for scaled data warehouses and data analytics applications.
  • Ensure technical health and quality by monitoring technical debt and driving reliability best practices.
  • Mentor and coach engineers on data storage, warehouse, processing, and SRE systems architecture.

Virtru provides a suite of data protection applications and an open platform based on the Trusted Data Format (TDF) open standard, enabling secure data sharing without sacrificing security. Backed by Iconiq Capital, Bessemer Venture Partners, and others, Virtru helps Fortune 500 companies and government agencies with a culture of curiosity and growth.

Global

  • Architect scalable data platforms for analytics, ML/AI products, and customer-facing data products.
  • Build reliable batch and near-real-time pipelines that are scalable and cost-efficient.
  • Productionize ML and AI products and support LLM/VLM workflows.

Vidmob is a creative data company that provides scoring software and analytics to help marketers improve creative effectiveness. It partners with the world's largest marketers and agencies and operates a robust human-reinforcement learning model for creativity.

United States

  • Own the technical strategy for the end-to-end Ads ML engineer lifecycle, focusing on feature development and training iteration.
  • Define architecture and technical standards for ML feature and training-data systems across batch/streaming computation, backfills, and quality.
  • Build platform abstractions and workflow automation to make ML development faster, safer, and more self-service.

Reddit is a community of communities built on shared interests, passion, and trust. With 100,000+ active communities and approximately 130 million daily active unique visitors, it is one of the internet's largest sources of information and has a culture of openness and authenticity.

$202,190–$237,870/yr
US

  • Design, build, and maintain business-critical data pipelines, including data ingestion and AWS Glue/Spark jobs feeding the warehouse and Tableau.
  • Build and operate core data infrastructure: the data warehouse, orchestration, and BI/reporting platform, deployed across multiple regions.
  • Drive work across producer and consumer teams, communicate clearly, and own outcomes end to end.

Horizon3 is a fast-growing cybersecurity company that provides autonomous pentesting and assessment operations via its NodeZero platform. They are a team of former special ops and engineers, committed to a culture of respect, collaboration, ownership, and results.

Americas

  • Design, build, and evolve the core data platform infrastructure including distributed query engines and orchestration.
  • Own our lakehouse infrastructure as code using Terraform and Ansible on Kubernetes.
  • Build and maintain low-latency streaming and batch ingestion pipelines, and scale the BI landscape.

Alpaca is a US-headquartered global leader in agent-first brokerage infrastructure for stocks, ETFs, options, crypto, and more. We are a dynamic team of 400+ globally distributed members who thrive working from our favorite places around the world.

India Unlimited PTO

  • Design and drive the execution of a scalable, cloud-native target-state architecture for the data platform.
  • Define architectural patterns for scalable automated data ingestion and deliver hands-on proof-of-concepts using SQL, PySpark, and Serverless Python.
  • Drive governance, security, and engineering excellence across all platform components, ensuring alignment with business and product strategies.

G-P is a SaaS-based Global Employment Platform that enables clients to expand into over 180 countries. Our diverse, remote-first teams foster innovation and value every contribution, empowering employees with flexibility and resources.

Global 6w PTO 26w maternity 26w paternity

  • Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it.
  • Collaborate daily with researchers and engineers who are some of the best in the world at what they do.
  • Build and maintain the high-performance data layer that Modeling teams rely on for training and evaluation jobs.

Cohere is a security-first enterprise AI company that builds cutting-edge foundation AI models and end-to-end products. They are a global team of researchers, engineers, and designers passionate about their craft, with offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul.

Global

  • Design and implement core components of the streaming platform including Kafka-compatible APIs and storage systems.
  • Solve complex distributed systems problems involving performance optimization and reliability.
  • Debug production issues, collaborate with customers, and contribute to open-source projects.

StreamNative, founded by the creators of Apache Pulsar, is redefining real-time data streaming with its URSA engine. The company is a growing startup focused on building a cloud-native streaming platform with a culture of engineering excellence.

Brazil

  • Translate legacy Databricks notebook logic into modern ELT patterns using Python and PySpark, ensuring data contract preservation and reverse-view strategies.
  • Build scalable ingestion pipelines with YAML configurations and orchestrate automated DAG generation via workflow schedulers, applying quality assertions and migration validation.
  • Collaborate with business Data Stewards to align dependencies, negotiate refactoring scope, and validate migrated outputs for production systems.

CI&T helps large companies transform AI potential into real business impact with AI deployment, AI-native execution, and tech-integrated business solutions. With 30 years of tech transformation experience and over 8,000 CI&Ters in 25+ countries, we accelerate innovation through agentic SDLC, application modernization, Data & AI, martech, and business strategy.