Remote Data Jobs · Apache Spark

Job listings

Europe 5w PTO

  • Design and build scalable cloud-based data platforms and pipelines in AWS or Azure/Fabric with Databricks.
  • Enable data pipelines for AI, GenAI, and analytics, including batch and streaming processing.
  • Collaborate with data scientists and business teams to shape requirements and deliver production-grade data solutions.

Valtech is an experience innovation company that helps brands unlock new value through data, AI, creativity, and technology. They foster a values-driven culture with an emphasis on inclusion, diversity, and professional growth.

  • Lead the design and implementation of scalable data architectures for member communications, defining data models, contracts, and lineage.
  • Build and optimize production-grade data pipelines using Databricks, Apache Spark, PySpark, and SQL, ensuring reliability and performance.
  • Establish data-quality and governance practices, and partner with Privacy, Compliance, and Risk teams to manage sensitive member data securely.

Oportun is a mission-driven financial services company that provides responsible credit and savings to help members build a better financial future. Since inception, it has provided over $22.7 billion in credit and saved members more than $2.5 billion in interest and fees, reflecting a culture focused on impact and innovation.

  • Design and build production software and data systems for a premier banking client.
  • Work hands-on with cloud-native technologies and modern architecture on enterprise-scale data movement and transformation.
  • Collaborate in Agile teams applying strong Python, AWS, and relational database skills to real-time and batch processing.

Kunai builds full-stack technology solutions for banks, credit and payment networks, infrastructure providers, and their customers. Our exceptional team thrives in a culture of collaboration, creativity, and continuous learning, with a track record of success spanning over 20 years.

Global 1w PTO

  • Design and build Databricks lakehouse pipelines using PySpark, Spark SQL, and Delta Lake.
  • Collaborate with frontend, architecture, and clinical teams on APIs and data ingestion.
  • Ensure GxP, 21 CFR Part 11, HIPAA, and GDPR compliance while optimizing Spark and Delta Lake.

Muttdata builds innovative Data Products and Machine Learning solutions for complex business challenges. They are a fast-growing, remote-first startup with a collaborative, ownership-driven culture that values continuous learning.

  • Collaborates with researchers and cross-functional teams to design and deliver scalable data solutions on Azure Fabric.
  • Develops and maintains data pipelines, ETL/ELT processes, and ensures data quality and HIPAA compliance.
  • Evaluates emerging technologies and creates technical documentation for complex datasets and workflows.

Emory University is a leading research university that fosters excellence and attracts world-class talent to innovate today and prepare leaders for the future. We welcome candidates who can contribute to the excellence of our academic community, offering a collaborative and inclusive culture.

  • Design and implement data solutions using cloud-native technologies like AWS, Databricks, and Snowflake.
  • Develop and optimize complex SQL queries and manage data migration at enterprise scale.
  • Collaborate with Agile teams to drive software and data engineering for a premier banking client.

Kunai builds full-stack technology solutions for banks, credit and payment networks, infrastructure providers, and their customers. Our team thrives in a culture of collaboration, creativity, and continuous learning, with over 20 years of success.

Europe 8w PTO

  • Design, build, and improve a scalable data platform to provide data solutions for product teams.
  • Develop and manage automated data pipelines using cloud-based and on-premise technologies.
  • Ensure data quality, lineage, and observability across the stack while collaborating with product and business teams.

lemlist is a bootstrapped sales engagement platform that grew from $0 to $57M ARR in 8 years. It is a profitable B2B SaaS company trusted by 40,000+ sales teams worldwide.

  • Architect and implement scalable, secure, and reliable data pipelines using modern platforms like Spark, Databricks, and Airflow.
  • Develop ETL/ELT processes to ingest data from various structured and unstructured sources, and perform exploratory data analysis.
  • Collaborate with cross-functional teams to design data models, write clean Python code, and deliver end-to-end features.

Netomi is a leading agentic AI platform for enterprise customer experience, working with global brands like Delta and MetLife. Backed by WndrCo, Y Combinator, and Index Ventures, the company is a well-funded startup with a collaborative culture.

  • Architect the Data Platform by designing internal SDKs and self-service frameworks for distributed engineering teams.
  • Own platform performance, optimizing the Databricks ecosystem for cost-effectiveness and scalability.
  • Drive data contracts and governance, implementing schema validation and security standards across the global ecosystem.

G-P offers a SaaS-based Global Employment Platform that enables companies to expand into over 180 countries quickly and efficiently. The company fosters a diverse, remote-first culture where innovation thrives and every contribution is valued.

$89,000–$127,000/yr
UK Unlimited PTO 4w paternity

  • Deliver complex data engineering projects for EMEA and US customers, individually or leading small teams.
  • Support pre-sales, mentor junior engineers, and contribute to hiring best-in-class data talent.
  • Work remotely in the UK with occasional travel to Budapest for onboarding and quarterly visits.

Datapao is a data engineering and cloud migration consulting company that helps clients solve complex data puzzles using technologies like Apache Spark and Databricks. They are a rapidly growing tech company with a high-transparency culture, emphasizing open feedback and no politics.