Remote Data Jobs · PySpark

Job listings

  • Design, develop, and maintain scalable ETL/ELT pipelines and data integration processes.
  • Build and optimize cloud-based data architectures to support analytics and business intelligence initiatives.
  • Ensure data quality, consistency, governance, and reliability through validation, monitoring, and automated quality checks.

CI&T helps large enterprises transform AI potential into business impact with AI deployment, AI-native execution, and tech-integrated solutions. With 30 years of experience and 8,000 employees across 25 countries, they collaborate to build solutions with real impact.

  • Design, develop, and maintain data pipelines using Azure Databricks.
  • Build and optimize data transformations using PySpark and SQL in Databricks.
  • Implement and maintain Lakehouse architectures using Delta Lake.

Miratech helps visionaries change the world. We are a global IT services and consulting company with nearly 1000 full-time professionals and a culture of Relentless Performance.

  • Own Gold-layer design and delivery, including facts, dimensions, domain data marts, and governed KPI implementations on Databricks.
  • Build and maintain the enterprise semantic layer through curated views, governed semantic models, and metric definitions.
  • Partner directly with business stakeholders across domains to clarify KPI definitions and resolve competing definitions before implementation.

Shield AI is a venture-backed defense-tech company developing intelligent systems to protect service members and civilians. The company has offices across the U.S., Europe, the Middle East, and Asia-Pacific, and its technology supports operations worldwide.

$140,000–$210,000/yr

  • Design and build ingestion pipelines from enterprise source systems into the Databricks lakehouse using Delta Lake, owning Bronze-layer ingestion and building Silver-layer pipelines for cleansing and standardization.
  • Implement and maintain Databricks platform constructs for secure delivery, including catalogs, schemas, service principals, and job orchestration, and build CI/CD pipelines for data platform assets.
  • Apply data classification and segregation requirements within pipeline design, build data quality controls reflecting business meaning, and partner with teams to ensure reliable, governed data for downstream use.

Shield AI is a venture-backed defense-tech company founded in 2015 with the mission of protecting service members and civilians with intelligent systems, developing products like Hivemind autonomy software and V-BAT and X-BAT aircraft. The company has offices and facilities across the U.S., Europe, the Middle East, and Asia-Pacific, and its technology actively supports operations worldwide.