Deliver high-quality data sets by curating, consolidating, and manipulating large-scale data sources.
Build high-quality data pipelines and ETL processes on platforms like Snowflake and BigQuery.
Partner with analytics, data science, and machine learning teams to provide data solutions.
Tripadvisor connects people to experiences worth sharing, aiming to be the world's most trusted source for travel. With over 500 million reviews and 390 million monthly visitors, the company is data-driven and offers a unique, global work environment that captures the speed and innovation of a startup.
Analyze, organize, and process large volumes of data to support business and technology decision-making.
Develop and maintain data pipelines using Databricks, Python, PySpark, and SQL, ensuring efficiency and quality.
Build dashboards, reports, and executive presentations to translate data into actionable insights.
CI&T helps large companies transform AI potential into real business impact with AI deployment, AI-native execution, and tech-integrated business solutions. They have 8,000 employees in over 25 countries, collaborating to build solutions with real impact, with AI integrated into their daily work.
Design, implement, and maintain ETL/ELT pipelines for clean, scalable data flows across multiple systems.
Own data warehouse architecture, including partitioning, access controls, and security in partnership with Security and Infrastructure teams.
Champion data quality and governance, building dashboards and analytics to deliver actionable insights for leadership and cross-functional teams.
CertifID protects life's largest transactions from fraud, helping title companies, law firms, lenders, and consumers safeguard billions of dollars from wire fraud daily. They are a growth-minded team passionate about securing sensitive data and transforming fraud prevention.
Design, build, and maintain robust, scalable ELT/ETL data pipelines from various source systems into cloud data platforms.
Perform data modeling, including dimensional modeling, and build transformation layers using dbt to create analytics-ready datasets.
Support operational reliability, monitor data pipelines, and ensure SLAs for timeliness, freshness, and accuracy.
Troveo builds the data platform that AI labs and model builders need to train the next generation of models. Backed by top investors, we’re a small, high-impact team solving one of the biggest bottlenecks in AI development.
Design, build, and maintain scalable data pipelines for ingestion and transformation.
Work with Python, SQL, Apache Airflow, and Google Cloud Platform.
Collaborate with cross-functional teams to deliver reliable and high-quality data solutions.
Our partner is a technology company focused on AI and data solutions. The team values collaboration, innovation, and continuous learning, and offers a remote work environment.
Design, develop, and maintain automated data pipelines and analytics solutions.
Automate manual reporting processes and repetitive data workflows.
Optimize SQL queries, data models, and processing performance.
Blend is a premier AI services provider that co-creates meaningful impact through data science, AI, technology, and people. The company is dedicated to unlocking value and fostering innovation for clients, with a focus on blending human expertise with artificial intelligence.
Oversee the full lifecycle of data products within the Data Core team, from ideation to operation and continuous optimization.
Ensure the data infrastructure (Databricks, Unity Catalog, data lake) operates with excellence, reliability, and scalability.
Lead and develop a cross-functional team of data engineers, scientists, and analytics experts, balancing long-term vision with hands-on execution.
We help large companies transform AI potential into real business impact with AI deployment, AI-native execution, and tech-integrated business solutions. With 30 years of experience, we are 8,000 CI&Ters in over 25 countries, collaborating to build solutions with real impact.
Design, build, and maintain scalable batch and streaming data pipelines.
Develop reliable ETL/ELT workflows using Python, Spark, and modern orchestration tools.
Improve data quality, validation, monitoring, and observability across the platform.
The company is a Berlin-based, remote-first technology company building advanced market intelligence and software solutions for the automotive industry. It operates in a stable growth phase with an established product and a strong technical team.
Build, expand, and optimize data infrastructure to create the most accurate dataset of identities and their relationships.
Develop and operate secure, scalable, and reliable data ingestion and ETL/ELT pipelines that meet product requirements.
Design and maintain a data observability framework to ensure data meets strict quality and freshness standards.
SentiLink provides innovative identity and risk solutions, empowering institutions and individuals to transact with confidence. The company is growing rapidly, has verified hundreds of millions of identities, and is backed by top investors like Craft Ventures and Andreessen Horowitz, with offices across the US and India.
Work cross-functionally with Product and experts to conceptualize, prototype, and build data solutions
Build and maintain data engineering systems and high-quality data models from multi-source healthcare datasets
Develop and test data pipelines and draft internal and external technical documentation
Turquoise Health is a Series C price transparency platform for finance leaders across healthcare. Backed by a16z, Oak HC/FT, and others, we are a remote-first, US-based team that values transparency, empathy, inclusivity, creativity, and ownership.
Design and implement robust, scalable data ingestion and transformation pipelines using Databricks, PySpark, and distributed processing.
Implement Delta Lake principles focusing on CDC and schema evolution, and integrate data quality frameworks within CI/CD pipelines.
Develop and optimize complex SQL and Python scripts, handling diverse data sources and supporting data governance solutions.
Mobile Wave Solutions is a professional services company specializing in software development as a service. With a team of over 120 engineers, we deliver scalable, high-quality software that empowers our global clients to innovate and grow.
Design, develop, and maintain scalable data pipelines that transform complex customer data into actionable business value.
Work across the full data lifecycle, from ingestion and transformation to activation and analytics.
Use modern cloud technologies and AI-assisted development workflows to create reliable, high-performing data solutions.
Jobgether uses an AI-powered matching process to help candidates find roles and connects top-fitting applicants directly with hiring companies. The company fosters a remote, collaborative environment with a focus on efficient recruitment.
Design, build, and optimize scalable data platforms supporting analytics, AI/ML, and enterprise reporting.
Develop and maintain complex data pipelines using AWS Glue, Step Functions, and Databricks Workflows.
Collaborate with product, engineering, analytics, and ML teams to transform complex data challenges into reliable solutions.
Jobgether uses AI to match candidates with job openings for faster, fairer reviews. It is a technology platform focused on streamlining the hiring process for both candidates and employers.
Act as bridge between business and technology, translating complex requirements into robust technical solutions.
Own the complete lifecycle of data products from ingestion through modeling, validation, and delivery.
Implement rigorous data validation processes and ensure data accuracy and integrity for AI and analytics.
CI&T helps large enterprises transform AI potential into real business impact with AI deployment, AI-native execution, and tech-integrated solutions. With 30 years of experience and 8,000 employees across 25 countries, they accelerate innovation and collaborate to build impactful solutions.
Design, build, and maintain scalable data pipelines using Python, SQL, Snowflake, Dagster, dbt, and AWS.
Own end-to-end data engineering projects from ingestion through to analytics enablement.
Improve monitoring, alerting, and data quality across key pipelines.
Midnite is a next-generation sports betting and gaming platform built for a new wave of players. Over 400,000 players have joined, and the team operates with high ownership and fast iteration in a scale-up environment.
Develop and implement Snowflake-based governance solutions for data quality, security, and compliance.
Build automated privacy processes and support data cleanup initiatives to maintain integrity.
Collaborate with cross-functional teams to design scalable data architecture and pipeline solutions.
Blend is a premier AI services provider, co-creating meaningful impact through data science, AI, and technology. We harness world-class people and data-driven strategy to unlock value for clients, with a focus on innovation and fulfilling work.
Drive migration of legacy Hadoop/Spark/Impala data pipelines to a modern Databricks-centric stack.
Design and build scalable Airflow DAGs and Databricks Jobs for large-scale data pipelines.
Implement robust validation strategies to ensure data parity and quality between legacy and modern systems.
LivePerson is a leader in trusted enterprise conversational AI and digital transformation. Named the #1 Most Innovative AI Company by Fast Company, we power nearly a billion conversational interactions every month for top global brands.
Design, build, and maintain scalable data and ML pipelines for analytics and AI systems.
Build and optimize workflows for structured and unstructured data, enabling semantic search and RAG use cases.
Manage and optimize vector databases and indexing strategies for efficient retrieval and AI-powered search.
This is a partner company seeking a Data & Machine Learning Engineer based in Brazil. They operate in a highly technical and global environment with strong emphasis on scalability, performance, and innovation.