Remote Data Jobs · PySpark

Job listings

  • Build end-to-end data products using Databricks, Delta Lake, SQL, Python, and PySpark.
  • Design and maintain batch, streaming, and near-real-time pipelines with Kafka and event-driven architectures.
  • Develop REST APIs, backend services, and lightweight applications while ensuring data governance and scalability.

This partner company focuses on data and technology solutions, hiring a Senior Data Engineer to build production-grade data products. The company's size and culture are not explicitly detailed, but the role emphasizes collaboration, remote work, and modern tech stacks.

$140,375–$185,604/yr

  • Design advanced analytics and predictive models within Databricks using Python, SQL, and Spark to identify revenue leakage and recovery opportunities.
  • Develop interactive Databricks dashboards and visualizations for operational and executive visibility into revenue-cycle performance.
  • Build detection logic for missing charges, coding errors, denied claims, underpayments, and aged receivables, and create recoverability scoring models.

LMI is a digital solutions provider accelerating government impact with innovation and speed, serving defense, space, healthcare, and energy sectors. Headquartered in Tysons, Virginia, the company delivers commercial-grade platforms and mission-ready AI to federal agencies.

  • Design, build, and maintain scalable data pipelines and ETL/ELT processes supporting cloud-based data environments for the Department of Veterans Affairs.
  • Implement patient-matching and record-linkage logic across multiple identifiers to ensure accurate, deduplicated Veteran data across disparate source systems.
  • Harmonize datasets across VA facilities, resolving structural and format inconsistencies to produce analysis-ready data assets.

Aptive partners with federal agencies to achieve their missions through improved performance, streamlined operations and enhanced service delivery. Founded in 2012, the company employs over 300 people nationwide and focuses on human-centered services.

$73,000–$90,000/yr
US Unlimited PTO

  • Be the primary owner of high-impact data products, running end-to-end pipelines or building granular data feeds.
  • Work with massive datasets to answer key client questions and derive insights using proprietary methods.
  • Improve products and pipelines by developing new methodologies and features, and learn SQL, Python, and PySpark.

YipitData is a leading market research and analytics firm for the disruptive economy, analyzing billions of alternative data points daily to provide insights on ridesharing, e-commerce, and more. The company has over 90 Data team members, recently raised $475M from The Carlyle Group, and fosters a people-centric culture focused on mastery, ownership, and transparency.

  • Support and evolve existing SQL Server data platforms and ETL solutions.
  • Develop and maintain data pipelines using Azure Databricks, PySpark, and SQL.
  • Collaborate with BI teams to support Power BI and reporting platforms.

Miratech is a global IT services and consulting company that helps visionaries change the world. They retain nearly 1000 full-time professionals with a culture of relentless performance and over 99% project success rate.

  • Lead design and development of enterprise-grade data engineering solutions using Microsoft Fabric and Azure.
  • Design scalable data pipelines for batch and real-time processing, optimizing ETL/ELT processes with modern Azure tools.
  • Collaborate with architects, provide technical leadership, mentor data engineering teams, and ensure data quality, security, and governance.

HSO is an IT services company specializing in Microsoft technologies, including Fabric and Azure, for enterprise data and cloud transformation. They operate globally and emphasize a collaborative, skilled technology team.

  • Design and maintain scalable data pipelines and data models for internal and client-facing products.
  • Collaborate with product, design, and engineering teams to translate business needs into reliable data solutions.
  • Improve data quality, performance, and engineering practices while taking ownership of the data platform.

This company is a fast-growing global SaaS business operating in the healthcare technology sector. It offers a remote-first engineering environment with talented colleagues across multiple countries and cultures.

  • Provide thought leadership, mentor junior staff, and deliver design artifacts for each sprint.
  • Build and support enterprise data integrations with ETL workflows using Azure Databricks, Data Factory, and Synapse.
  • Collaborate with global teams to ensure scalable, performant, and secure data solutions aligned with enterprise architecture.

QVC Group is a Fortune 500 live social shopping company with six leading retail brands that redefine the shopping experience through video-driven commerce. They have team members across the U.S., U.K., Germany, Japan, Italy, Poland, and China, fostering an inspired, diverse, and collaborative culture.

  • Design and deliver complex data platform components or migration solutions on AWS using Databricks.
  • Troubleshoot performance, scalability, and reliability issues in cloud-native data environments.
  • Communicate technical topics clearly to stakeholders and produce high-quality documentation.

Caylent is an AI-first cloud services company that helps organizations turn ambitious ideas into meaningful business impact. As a fully remote global company with employees in Canada, the United States, and Latin America, they celebrate diverse cultures and foster a community of technological curiosity.

$130,000–$150,000/yr
US 4w PTO

  • Data Ingestion & Pipeline Reliability: Ensuring continuous, accurate data availability by monitoring and optimizing batch and real-time streaming pipelines.
  • Data Mart Delivery & Business Modeling: Delivering query-ready Gold-layer data marts by partnering with Data Analytics and Data Science teams.
  • Data Quality Frameworks & Precision: Maintaining flawless data integrity by writing and deploying automated data quality test suites.

Trupanion is a leading provider of medical insurance for cats and dogs in North America, helping pet owners budget and care for their pets. The company offers a collaborative, casual, and pet-friendly environment where everyone is encouraged to be themselves.