Design and build scalable data pipelines using Apache Spark and cloud technologies.
Develop and maintain data integrations and transformation processes from multiple sources.
Collaborate with technical teams to ensure data quality and optimize data architecture.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They focus on using technology to streamline the hiring process and offer a remote work environment.
Lead a team of data engineers to design and optimize scalable data solutions.
Architect and implement large-scale streaming and batch data processing using Kafka, Spark, and Hadoop.
Collaborate with cross-functional teams to define data strategies and ensure data governance and quality.
The company is a technology-focused organization that builds scalable data platforms. It offers a fully remote work environment with a collaborative culture and international teams.
Design, build, and maintain scalable data pipelines and ETL/ELT processes using Databricks and Spark.
Architect and optimize data models and storage solutions for analytics and operational use.
Implement observability, alerting, and data quality monitoring for critical pipelines.
Sonatype provides end-to-end software supply chain security solutions, protecting against malicious open source and managing SBOMs. With over 2,000 organizations and 15 million developers using its platform, it focuses on innovation and security in software development.
Develop and maintain scalable, resilient data pipelines using PySpark and distributed processing frameworks.
Modernize legacy processes and design data solutions on AWS, building data products and complex transformations.
Support data quality and optimization initiatives, collaborating with stakeholders to translate requirements into technical solutions.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. It focuses on efficiency and fairness in recruitment, promoting diversity and inclusion in the workplace.
Design and build batch data pipelines that ingest, validate, and transform multi-billion-row datasets.
Model complex real-world data including dimensional models and temporal data.
Develop and operate workloads on lakehouse platforms like Databricks, Spark, and Delta.
Simulmedia builds an advanced TV and streaming advertising platform. They have a team of engineers, data scientists, and designers who are obsessed with building cutting-edge technology.
Own the design, delivery, and operation of major data products including ETL pipelines and storage solutions.
Partner with engineering, product, and business stakeholders to define solutions and drive projects from design to production.
Mentor engineers, lead technical strategy, and participate in on-call rotation for production systems.
Lime is the largest global shared micromobility company, on a mission to make transportation shared, affordable, and carbon-free. It has powered over one billion rides across 30 countries and is a Time 100 Most Influential Company.
Design and implement scalable lakehouse architectures using Bronze, Silver, and Gold patterns with technologies like Kafka, Spark, and Airflow.
Build and manage high-performance streaming and batch data pipelines and open table format solutions such as Iceberg, Delta Lake, and Hudi.
Optimize distributed query environments, maintain platform reliability, and partner with data scientists and product teams to deliver impactful data capabilities.
The partner company is building scalable data platforms that power analytics, machine learning, and next-generation products. They offer a collaborative and supportive work environment focused on innovation and continuous learning.
You will develop and maintain end-to-end data pipelines and contribute to Samsara's Data Platform for advanced automation and analytics.
You will design, build, and optimize large-scale Spark and PySpark workflows for batch and streaming data processing.
You will build and maintain MCP servers and AI agents, and champion data engineering best practices across the team.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for actionable insights. As a recently public company, they foster a culture of autonomy, support, and rapid career development in a hyper-growth environment.
Architect and improve real-time data streaming systems processing millions of events per second with ultra-low latency.
Lead development of core data infrastructure enabling real-time decision-making and personalization.
Solve complex distributed systems challenges including fault tolerance, scalability, and performance optimization.
The hiring company is a technology firm focused on building real-time data platforms and AI-driven systems. The company culture emphasizes innovation, autonomy, and high engineering standards, though specific employee count is not provided.
Design, develop, and maintain scalable ETL/ELT pipelines
Build and optimize data processing solutions using Python and Apache Spark
Develop and support real-time data streaming applications using Kafka
Sigma Software develops modern digital solutions for global businesses. They foster an international remote-first environment where engineers can grow, innovate, and influence technical decisions.
Build and maintain batch data pipelines ingesting from Workday and legacy systems into a medallion lakehouse on Azure Databricks and Microsoft Fabric.
Develop transformation logic across Bronze/Silver/Gold layers using Spark, Python, and SQL, and implement data-quality checks.
Work within a federated architecture and Unity Catalog security model, contributing to documentation and code reviews.
CampusWorks is a consulting firm that helps higher education institutions leverage technology and managed services to drive transformation and student success. Founded in 1999, it is a large, virtual company with a culture that values work-life balance and employee appreciation.
Guide strategic enterprise customers through cloud data engineering transformations, including performance testing and production-ready pipeline architecture.
Prove platform value by architecting solutions for big data, data warehousing, and lakehouse use cases.
Support technical sales through custom proofs of concept, workload sizing, and community workshops.
Databricks is the Data and AI company, providing a unified platform for data, analytics, and AI to over 20,000 organizations worldwide, including 70% of the Fortune 500. Headquartered in San Francisco with 30+ offices globally, it fosters a diverse and inclusive culture and offers comprehensive benefits.
Build, expand, and optimize data infrastructure to create the most accurate dataset of identities and their relationships.
Develop and operate secure, scalable, and reliable data ingestion and ETL/ELT pipelines that meet product requirements.
Design and maintain a data observability framework to ensure data meets strict quality and freshness standards.
SentiLink provides innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. The company is growing rapidly, has verified hundreds of millions of identities, and is backed by top investors like Craft Ventures and Andreessen Horowitz, with offices across the US and India.
Design, develop, maintain, and optimize ETL and data transformation processes.
Develop and maintain Apache Spark-based data pipelines and contribute to data integration initiatives.
Ensure data quality, reliability, and performance across data processing solutions.
Talan is an international consulting group specializing in innovation and business transformation through technology. With over 7,200 consultants in 21 countries and a turnover of €850M, the company delivers impactful, future-ready solutions.
Build, maintain, and run CI/CD pipelines and infrastructure-as-code for the platform and services.
Provision and manage cloud-based Spark clusters and distributed data processing environments.
Investigate and resolve data pipeline issues while monitoring systems and managing cloud costs.
Smile Digital Health provides a FHIR-based data liberation platform for healthcare stakeholders to collect and exchange data. They were recognized as #19 on Deloitte's Technology Fast 50 Ranking for 2024 and operate in over 20 countries, fostering a culture of respect, inclusion, and diversity.
Design and build scalable batch data pipelines using AWS Glue, Lambda, and Step Functions.
Optimize analytical data models and query layers with Amazon Athena, Redshift, and ClickHouse.
Collaborate with backend engineers to integrate data workflows with microservices and event-driven architectures.
The company is a technology organization focused on building scalable data platforms and analytics solutions. They operate with a distributed international team and emphasize modern engineering practices.
Design, develop, and maintain scalable ETL/ELT pipelines and data integration processes.
Build and optimize cloud-based data architectures to support analytics and business intelligence initiatives.
Ensure data quality, consistency, governance, and reliability through validation, monitoring, and automated quality checks.
CI&T helps large enterprises transform AI potential into business impact with AI deployment, AI-native execution, and tech-integrated solutions. With 30 years of experience and 8,000 employees across 25 countries, they collaborate to build solutions with real impact.
Lead design and integration of new services for data and event processing.
Scale the Datastore’s core infrastructure and implement APIs for clinical studies.
Partner with product managers and teams to ensure data integrity and maintainability.
Beacon Biosignals is a precision medicine company focused on brain disorders, using EEG data and AI for biomarker discovery. With nearly 100,000 patient records and a cloud platform, their small team fosters empathy and scientific impact.
Build and maintain scalable data pipelines and enrichment workflows.
Integrate enriched data into broader systems and apply necessary transformations.
Contribute to code reviews and follow development workflows for quality and efficiency.
H1 provides a platform that uses data and AI to unlock medical insights and improve healthcare globally. It is a mission-driven company with flexible work and a focus on health equity.
Own the Airflow codebase end-to-end, building reusable templates and enforcing standards.
Build and extend large-scale Spark pipelines on AWS Glue and support migration to Databricks.
Drive data model improvements around commercial pharma data with focus on structure and lineage.
Veeva Systems is a mission-driven organization and pioneer in industry cloud, helping life sciences companies bring therapies to patients faster. As one of the fastest-growing SaaS companies in history, it surpassed $3B in revenue and values doing the right thing, customer success, employee success, and speed.