Own the Airflow codebase end-to-end, building reusable templates and enforcing standards.
Build and extend large-scale Spark pipelines on AWS Glue and support migration to Databricks.
Drive data model improvements around commercial pharma data with focus on structure and lineage.
Veeva Systems is a mission-driven organization and pioneer in industry cloud, helping life sciences companies bring therapies to patients faster. As one of the fastest-growing SaaS companies in history, it surpassed $3B in revenue and values doing the right thing, customer success, employee success, and speed.
Maintain and monitor Kafka, Hadoop, Presto, and RDBMS systems.
Ingest, validate, and process internal and third-party data flows.
Build Kafka consumers using Spark Streaming for near-real-time aggregation.
PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health journey signals. We are 300+ employees, growing, and a leading player in the US healthcare ad market.
Design and build scalable data pipelines using Apache Spark and cloud technologies.
Develop and maintain data integrations and transformation processes from multiple sources.
Collaborate with technical teams to ensure data quality and optimize data architecture.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They focus on using technology to streamline the hiring process and offer a remote work environment.
Support new and existing customers in their data engineering needs, guiding them to make optimal technical decisions.
Build and operationalize complex data solutions, including data governance, security, and quality.
Collaborate across teams to deliver data products and educate end users on analytic environments.
SunnyData is a high-growth consulting company specializing in data engineering and AI, dedicated to the Databricks platform. They foster a collaborative and innovative culture with a focus on customer impact and career growth.
You will develop and maintain end-to-end data pipelines and contribute to Samsara's Data Platform for advanced automation and analytics.
You will design, build, and optimize large-scale Spark and PySpark workflows for batch and streaming data processing.
You will build and maintain MCP servers and AI agents, and champion data engineering best practices across the team.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for actionable insights. As a recently public company, they foster a culture of autonomy, support, and rapid career development in a hyper-growth environment.
Guide strategic enterprise customers through cloud data engineering transformations, including performance testing and production-ready pipeline architecture.
Prove platform value by architecting solutions for big data, data warehousing, and lakehouse use cases.
Support technical sales through custom proofs of concept, workload sizing, and community workshops.
Databricks is the Data and AI company, providing a unified platform for data, analytics, and AI to over 20,000 organizations worldwide, including 70% of the Fortune 500. Headquartered in San Francisco with 30+ offices globally, it fosters a diverse and inclusive culture and offers comprehensive benefits.
Design and operate production-grade data pipelines for healthcare, identity, and campaign datasets using Python, SQL, Airflow, and Snowflake.
Implement maintainable data models and governed views for provider, claims, and audience data, ensuring privacy-by-design for PHI/PII.
Partner with product, analytics, and data science teams to translate business requirements into resilient technical solutions that power audience discovery and measurement.
Zeta Global is the AI-Powered Marketing Cloud that leverages advanced artificial intelligence and trillions of consumer signals to make marketing simpler. Founded in 2007 with offices worldwide, the company fosters a culture of trust and belonging, offering equity and wellness benefits to its employees.
Design, build, and maintain business-critical data pipelines, including data ingestion and AWS Glue/Spark jobs feeding the warehouse and Tableau.
Build and operate core data infrastructure: the data warehouse, orchestration, and BI/reporting platform, deployed across multiple regions.
Drive work across producer and consumer teams, communicate clearly, and own outcomes end to end.
Horizon3 is a fast-growing cybersecurity company that provides autonomous pentesting and assessment operations via its NodeZero platform. They are a team of former special ops and engineers, committed to a culture of respect, collaboration, ownership, and results.
Design and implement ETL pipelines using Apache Airflow, BigQuery, Python, and Spark to transform upstream data into curated data assets.
Provide technical leadership and best practices, mentoring other engineers and driving architecture decisions for high-performance systems.
Collaborate cross-functionally with Product Managers and end users to define key business questions and build relevant data sets.
InMarket is a leader in 360-degree consumer intelligence and real-time activation for top brands, offering a data-driven marketing platform. The company has a strong focus on technology and culture, with a commitment to diversity, equity, and inclusion, and offers competitive compensation and benefits.
Build and maintain batch data pipelines ingesting from Workday and legacy systems into a medallion lakehouse on Azure Databricks and Microsoft Fabric.
Develop transformation logic across Bronze/Silver/Gold layers using Spark, Python, and SQL, and implement data-quality checks.
Work within a federated architecture and Unity Catalog security model, contributing to documentation and code reviews.
CampusWorks is a consulting firm that helps higher education institutions leverage technology and managed services to drive transformation and student success. Founded in 1999, it is a large, virtual company with a culture that values work-life balance and employee appreciation.
Design, develop, and maintain scalable ETL/ELT pipelines and data integration processes.
Build and optimize cloud-based data architectures to support analytics and business intelligence initiatives.
Ensure data quality, consistency, governance, and reliability through validation, monitoring, and automated quality checks.
CI&T helps large enterprises transform AI potential into business impact with AI deployment, AI-native execution, and tech-integrated solutions. With 30 years of experience and 8,000 employees across 25 countries, they collaborate to build solutions with real impact.
Design, build, and maintain scalable data pipelines and ETL/ELT processes using Databricks and Spark.
Architect and optimize data models and storage solutions for analytics and operational use.
Implement observability, alerting, and data quality monitoring for critical pipelines.
Sonatype provides end-to-end software supply chain security solutions, protecting against malicious open source and managing SBOMs. With over 2,000 organizations and 15 million developers using its platform, it focuses on innovation and security in software development.
Define scalable architectures for enterprise data platforms across ingestion, storage, processing, governance, and consumption.
Establish technical standards for data modeling, security, reliability, and lifecycle management.
Mentor engineers and guide architectural direction while partnering with cross-functional teams.
The company is a technology organization focused on building enterprise data platforms. It fosters a culture of innovation and collaboration, with a focus on technical excellence.
Enable efficient data access by creating and maintaining data pipelines.
Collaborate with ML engineers to design and maintain automation for machine learning training, quality assessment, and model release.
Build data infrastructure for analytics, hypothesis testing, and company metrics.
Eneba is building an open, safe, and sustainable marketplace for gamers, supporting close to 20 million active users. We are a growing international team that values data-driven decision making and fosters a healthy data culture.
Architect enterprise-scale Databricks data platforms with guardrails, automation, and standards to enable safe self-service by business teams.
Design high-volume batch and real-time streaming pipelines using PySpark and Structured Streaming, integrating Unity Catalog with third-party metadata.
Communicate architectural decisions to technical and business stakeholders, produce documentation, and partner with client teams to translate use-case needs into platform capabilities.
Caylent is an AI-first cloud services company that helps organizations turn ambitious ideas into meaningful business impact using AWS, AI, and Anthropic's Claude platform. They are a fully remote global company with employees across Canada, the US, and Latin America, fostering a culture of technological curiosity and inclusion.
Design, build, and maintain scalable data pipelines and platform capabilities to power analytics, AI/ML, and healthcare products.
Implement data quality, governance, and security controls, ensuring healthcare data is handled securely and in compliance.
Provide technical leadership, mentor engineers, and collaborate across teams to translate requirements into scalable data solutions.
Experity is a mission-driven team transforming on-demand healthcare across the U.S., empowering urgent care clinics with industry-leading software. They foster a culture of care, growth, and celebration with day-one benefits, career development, and a supportive team environment.
Design and build batch data pipelines that ingest, validate, and transform multi-billion-row datasets.
Model complex real-world data including dimensional models and temporal data.
Develop and operate workloads on lakehouse platforms like Databricks, Spark, and Delta.
Simulmedia builds an advanced TV and streaming advertising platform. They have a team of engineers, data scientists, and designers who are obsessed with building cutting-edge technology.
Design and build data pipelines for ML use cases, implementing data versioning and orchestration.
Deploy data scientists' scripts and models from notebooks to production, managing Python dependencies and CI/CD.
Optimize data processing jobs for performance and cost on distributed systems and cloud data services.
NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior and revealing new pathways to growth. In 2023, NIQ combined with GfK, bringing together two industry leaders with unparalleled global reach, operating in 100+ markets.
Own the identity data model and coordinate cross-ecosystem mappings for billions of package observations.
Architect streaming and batch pipelines using Java, Spark, HBase, Databricks, and AWS primitives.
Lead design of data contracts for downstream teams, ensuring schema stability and versioning.
Sonatype is the software supply chain security company, providing end-to-end security solutions including proactive protection against malicious open source and enterprise-grade SBOM management. With over 2,000 organizations including 70% of the Fortune 100 and 15 million developers relying on its platform, Sonatype is a market leader in software supply chain security.
Develop and maintain scalable, resilient data pipelines using PySpark and distributed processing frameworks.
Modernize legacy processes and design data solutions on AWS, building data products and complex transformations.
Support data quality and optimization initiatives, collaborating with stakeholders to translate requirements into technical solutions.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It focuses on efficiency and fairness in recruitment, promoting diversity and inclusion in the workplace.