Design and build ingestion pipelines from enterprise source systems into the Databricks lakehouse using Delta Lake, owning Bronze-layer ingestion and building Silver-layer pipelines for cleansing and standardization.
Implement and maintain Databricks platform constructs for secure delivery, including catalogs, schemas, service principals, and job orchestration, and build CI/CD pipelines for data platform assets.
Apply data classification and segregation requirements within pipeline design, build data quality controls reflecting business meaning, and partner with teams to ensure reliable, governed data for downstream use.
Shield AI is a venture-backed defense-tech company founded in 2015 with the mission of protecting service members and civilians with intelligent systems, developing products like Hivemind autonomy software and V-BAT and X-BAT aircraft. The company has offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, and its technology actively supports operations worldwide.
You will develop and maintain end-to-end data pipelines and contribute to Samsara's Data Platform for advanced automation and analytics.
You will design, build, and optimize large-scale Spark and PySpark workflows for batch and streaming data processing.
You will build and maintain MCP servers and AI agents, and champion data engineering best practices across the team.
Samsara is the pioneer of the Connected Operations Cloud, enabling organizations to harness IoT data for actionable insights. As a recently public company, they foster a culture of autonomy, support, and rapid career development in a hyper-growth environment.
Design and build batch data pipelines that ingest, validate, and transform multi-billion-row datasets.
Model complex real-world data including dimensional models and temporal data.
Develop and operate workloads on lakehouse platforms like Databricks, Spark, and Delta.
Simulmedia builds an advanced TV and streaming advertising platform. They have a team of engineers, data scientists, and designers who are obsessed with building cutting-edge technology.
Design and maintain scalable batch and near real-time data pipelines in Databricks, integrating data from various sources.
Build unified customer profiles and support identity resolution, deduplication, enrichment, and Golden Record creation.
Enable CRM and marketing teams to access trusted, activation-ready customer data with minimal latency.
Massive Rocket is a rapidly scaling Braze and Snowflake agency that transforms how digital marketing, product, and engineering teams connect. They have grown quickly in five years and are aiming for $100M in revenue, with a culture of ownership, collaboration, and growth.
Design, build, and maintain scalable data pipelines and ETL/ELT processes using Databricks and Spark.
Architect and optimize data models and storage solutions for analytics and operational use.
Implement observability, alerting, and data quality monitoring for critical pipelines.
Sonatype provides end-to-end software supply chain security solutions, protecting against malicious open source and managing SBOMs. With over 2,000 organizations and 15 million developers using its platform, it focuses on innovation and security in software development.
Design and implement ETL pipelines using Apache Airflow, BigQuery, Python, and Spark to transform upstream data into curated data assets.
Provide technical leadership and best practices, mentoring other engineers and driving architecture decisions for high-performance systems.
Collaborate cross-functionally with Product Managers and end users to define key business questions and build relevant data sets.
InMarket is a leader in 360-degree consumer intelligence and real-time activation for top brands, offering a data-driven marketing platform. The company has a strong focus on technology and culture, with a commitment to diversity, equity, and inclusion, and offers competitive compensation and benefits.
Build and maintain BigQuery data models using Dataform, following medallion architecture patterns (Bronze/Silver/Gold).
Contribute to Looker dashboards and LookML models, working alongside senior engineers and analysts.
Build and maintain robust Python data pipelines with testing, linting, and CI/CD integration.
Kitman Labs is a performance intelligence company that transforms how the sports industry uses data to unlock athlete potential. With over 2000 teams in 50 leagues across 6 continents, the company has assembled a team of top data scientists, sports performance scientists, and product specialists, fostering an innovative and collaborative culture.
Maintain and monitor Kafka, Hadoop, Presto, and RDBMS systems.
Ingest, validate, and process internal and third-party data flows.
Build Kafka consumers using Spark Streaming for near-real-time aggregation.
PulsePoint sits at the intersection of healthcare and adtech, helping brands interpret health journey signals. We are 300+ employees, growing, and a leading player in the US healthcare ad market.
Support new and existing customers in their data engineering needs, guiding them to make optimal technical decisions.
Build and operationalize complex data solutions, including data governance, security, and quality.
Collaborate across teams to deliver data products and educate end users on analytic environments.
SunnyData is a high-growth consulting company specializing in data engineering and AI, dedicated to the Databricks platform. They foster a collaborative and innovative culture with a focus on customer impact and career growth.
Design and deliver complex data platform components or migration solutions on AWS using Databricks.
Troubleshoot performance, scalability, and reliability issues in cloud-native data environments.
Communicate technical topics clearly to stakeholders and produce high-quality documentation.
Caylent is an AI-first cloud services company that helps organizations turn ambitious ideas into meaningful business impact. As a fully remote global company with employees in Canada, the United States, and Latin America, they celebrate diverse cultures and foster a community of technological curiosity.
Build, expand, and optimize data infrastructure to create the most accurate dataset of identities and their relationships.
Develop and operate secure, scalable, and reliable data ingestion and ETL/ELT pipelines that meet product requirements.
Design and maintain a data observability framework to ensure data meets strict quality and freshness standards.
SentiLink provides innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. The company is growing rapidly, has verified hundreds of millions of identities, and is backed by top investors like Craft Ventures and Andreessen Horowitz, with offices across the US and India.
Design, develop, and maintain scalable, production-ready data pipelines and data products using Spark (Python/SQL) in a Databricks environment.
Lead the integration and transformation of complex data from diverse DoD and federal health systems into reliable, reusable data products.
Provide technical guidance and mentorship to other engineers, helping teams navigate complex technical challenges.
540 is a forward-thinking company that delivers innovative technology solutions for government missions. The team has a culture of breaking down barriers and solving mission-critical problems.
Become a trusted data and AI advisor, translating business questions into AI-ready data architectures.
Design and implement AI-optimized data platforms, including cloud data warehouses, lakehouses, and ETL/ELT pipelines.
Engineer modern ELT/ETL pipelines and data models using SQL, Python, and tools like Snowflake, Databricks, and dbt.
Aimpoint Digital is a fully remote data and analytics consultancy that partners with innovative software providers to solve complex business problems. The team is dynamic and collaborative, working independently on client engagements across industries.
Design and develop automated ETL/ELT pipelines to ingest data into Google Cloud Platform.
Implement Medallion Architecture patterns and maintain governed views in BigQuery.
Balance new solution development with production incident resolution and support.
Our partner is a company seeking a senior data professional to design, build, and operate scalable data solutions in a large-scale corporate environment. They collaborate across multiple business areas and combine development with operational support.
Design and build scalable data pipelines (batch & streaming) on GCP.
Develop and manage API-driven integrations (REST, JSON, event-based).
Work with healthcare data standards such as FHIR, HL7, and EDI.
Egen is a data-first consulting firm specializing in Google Cloud and Salesforce, helping clients drive action through data and insights. The company is fast-growing and entrepreneurial, with a culture focused on learning, problem-solving, and innovation.
Own data pipelines end to end - ingestion, transformation, warehousing, and delivery into BI.
Build and maintain dbt models with tests, documentation, and clear contracts.
Own product analytics: instrumentation, event tracking, usage models, and metrics for GTM and product teams.
Cortex is the Engineering Operations Platform providing visibility, governance, and golden paths for engineering organizations. We are a team of 80 passionate individuals excited about building a product that developers love.
Design, build, and maintain scalable data pipelines for ingestion and transformation.
Work with Python, SQL, Apache Airflow, and Google Cloud Platform.
Collaborate with cross-functional teams to deliver reliable and high-quality data solutions.
Our partner is a technology company focused on AI and data solutions. The team values collaboration, innovation, and continuous learning, and offers a remote work environment.
Own the Airflow codebase end-to-end, building reusable templates and enforcing standards.
Build and extend large-scale Spark pipelines on AWS Glue and support migration to Databricks.
Drive data model improvements around commercial pharma data with focus on structure and lineage.
Veeva Systems is a mission-driven organization and pioneer in industry cloud, helping life sciences companies bring therapies to patients faster. As one of the fastest-growing SaaS companies in history, it surpassed $3B in revenue and values doing the right thing, customer success, employee success, and speed.
Build out and configure a dedicated Government Databricks workspace within the existing environment
Design and implement data ingestion pipelines from core agency systems including financial, HR, CRM, and ITSM
Implement Databricks Unity Catalog for centralized data governance, metadata management, and end-to-end lineage tracking
Peraton is a next-generation national security company that drives missions of consequence globally. As a leading mission capability integrator and enterprise IT provider, the company serves essential government agencies and supports every branch of the U.S. armed forces.
Build and maintain batch data pipelines ingesting from Workday and legacy systems into a medallion lakehouse on Azure Databricks and Microsoft Fabric.
Develop transformation logic across Bronze/Silver/Gold layers using Spark, Python, and SQL, and implement data-quality checks.
Work within a federated architecture and Unity Catalog security model, contributing to documentation and code reviews.
CampusWorks is a consulting firm that helps higher education institutions leverage technology and managed services to drive transformation and student success. Founded in 1999, it is a large, virtual company with a culture that values work-life balance and employee appreciation.