Design and build scalable inference pipelines running on large GPU clusters for synthetic data generation.
Conduct data ablations and experiments to enhance model performance and data quality.
Research and implement innovative synthetic data curation methods using Cohere's infrastructure.
Cohere is a leading security-first enterprise AI company that builds cutting-edge foundation AI models and end-to-end products for enterprises. They are a global team of researchers, engineers, and designers with offices in Toronto, London, New York, San Francisco, Montreal, Paris, Berlin, and Seoul.
Work directly on petabyte-scale storage infrastructure, and the networking and performance challenges that come with it.
Collaborate daily with researchers and engineers who are some of the best in the world at what they do.
Build and maintain the high-performance data layer that Modeling teams rely on for training and evaluation jobs.
Cohere is a security-first enterprise AI company that builds cutting-edge foundation AI models and end-to-end products. They are a global team of researchers, engineers, and designers passionate about their craft, with offices in Toronto, London, New York City, San Francisco, Montreal, Paris, Berlin, and Seoul.
Build and design large-scale distributed crawling bots and infrastructure to extract data from diverse sources including web pages, APIs, and PDFs.
Develop and maintain data pipelines using Python, Scrapy, Airflow, and Spark to ingest, clean, and normalize data for search indexing and downstream ML/NLP.
Collaborate with ML/NLP teams to integrate classification models and utilize LLMs for precise answers from OEM repair procedures.
Tekmetric provides all-in-one cloud-based software for auto repair shops, helping them run smarter and grow faster. Founded in 2017 and headquartered in Houston, the company has grown from a single shop's vision to an industry-leading solution, fostering a culture of transparency, integrity, innovation, and service-first mindset.
Design and build reliable data pipelines using Spark, Kafka, Iceberg, and Airflow across batch, streaming, and real-time workloads.
Contribute to the evolution of the data lake and platform, including ingestion, processing, storage, and serving patterns.
Define and improve data quality, observability, reliability, and governance standards across data systems.
Webflow is an agentic web marketing platform for modern marketing teams, helping organizations build, manage, and optimize high-performing web experiences. The company is a growing, privately held organization that values grit, speed, and craft.
Maintain, optimize, and troubleshoot database queries and related data systems to support efficient data access, processing, and reliability.
Assist in creating, maintaining, and improving data pipelines used to collect, process, transform, validate, and deliver large-scale datasets.
Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts or tools used to gather publicly available data.
We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models. We're lean, technical, and move fast.
Design, develop, and maintain scalable data pipelines and infrastructure.
Collaborate with cross-functional teams to ensure data integrity and usability.
Implement data quality and governance practices within data pipelines.
We help large enterprises transform with AI and tech-integrated solutions. With 8,000 employees across 25+ countries, we collaborate to build solutions with real impact.
Design, build, and maintain scalable data pipelines and ETL/ELT processes using Databricks and Spark.
Architect and optimize data models and storage solutions for analytics and operational use.
Implement observability, alerting, and data quality monitoring for critical pipelines.
Sonatype provides end-to-end software supply chain security solutions, protecting against malicious open source and managing SBOMs. With over 2,000 organizations and 15 million developers using its platform, it focuses on innovation and security in software development.
Own the Airflow codebase end-to-end, building reusable templates and enforcing standards.
Build and extend large-scale Spark pipelines on AWS Glue and support migration to Databricks.
Drive data model improvements around commercial pharma data with focus on structure and lineage.
Veeva Systems is a mission-driven organization and pioneer in industry cloud, helping life sciences companies bring therapies to patients faster. As one of the fastest-growing SaaS companies in history, it surpassed $3B in revenue and values doing the right thing, customer success, employee success, and speed.
Build and maintain reliable data pipelines and ETL/ELT workflows.
Develop and optimize data models for analytics and internal tools.
Support core data platform tools like Spark and AWS, and monitor pipeline quality and performance.
Sonatype is the software supply chain security company, providing end-to-end software supply chain security solutions. Over 2,000 organizations, including 70% of the Fortune 100, and 15 million software developers rely on Sonatype to optimize their software supply chains.
Design and build data pipelines for ML use cases, implementing data versioning and orchestration.
Deploy data scientists' scripts and models from notebooks to production, managing Python dependencies and CI/CD.
Optimize data processing jobs for performance and cost on distributed systems and cloud data services.
NIQ is the world's leading consumer intelligence company, delivering the most complete understanding of consumer buying behavior and revealing new pathways to growth. In 2023, NIQ combined with GfK, bringing together two industry leaders with unparalleled global reach, operating in 100+ markets.
Design, develop, and maintain scalable, production-ready data pipelines and data products using Spark (Python/SQL) in a Databricks environment.
Lead the integration and transformation of complex data from diverse DoD and federal health systems into reliable, reusable data products.
Provide technical guidance and mentorship to other engineers, helping teams navigate complex technical challenges.
540 is a forward-thinking company that delivers innovative technology solutions for government missions. The team has a culture of breaking down barriers and solving mission-critical problems.
Design and implement ETL pipelines using Apache Airflow, BigQuery, Python, and Spark to transform upstream data into curated data assets.
Provide technical leadership and best practices, mentoring other engineers and driving architecture decisions for high-performance systems.
Collaborate cross-functionally with Product Managers and end users to define key business questions and build relevant data sets.
InMarket is a leader in 360-degree consumer intelligence and real-time activation for top brands, offering a data-driven marketing platform. The company has a strong focus on technology and culture, with a commitment to diversity, equity, and inclusion, and offers competitive compensation and benefits.
Design and deliver complex data platform components or migration solutions on AWS using Databricks.
Troubleshoot performance, scalability, and reliability issues in cloud-native data environments.
Communicate technical topics clearly to stakeholders and produce high-quality documentation.
Caylent is an AI-first cloud services company that helps organizations turn ambitious ideas into meaningful business impact. As a fully remote global company with employees in Canada, the United States, and Latin America, they celebrate diverse cultures and foster a community of technological curiosity.
Design and maintain large-scale cloud data infrastructure for healthcare applications.
Build efficient ETL pipelines, self-service tools, and microservices using Azure, Snowflake, and Databricks.
Collaborate with product owners and team leads to define efficient data pipelines and schemas.
Sigma Software is an IT solutions company delivering innovative technology to global clients across multiple industries. They foster a supportive environment with cutting-edge technologies and meaningful projects.
Build and maintain data infrastructure processing petabytes of data across millions of users.
Write clean, well-tested code for data ingestion, transformation, and serving systems.
Own projects end-to-end and collaborate with data scientists, engineers, and product teams.
Discord is a voice, video, and text communication platform built for gaming communities, with millions of daily active users. The company fosters an inclusive, collaborative culture focused on deepening friendships around games and shared interests.
Design and operate production-grade data pipelines for healthcare, identity, and campaign datasets using Python, SQL, Airflow, and Snowflake.
Implement maintainable data models and governed views for provider, claims, and audience data, ensuring privacy-by-design for PHI/PII.
Partner with product, analytics, and data science teams to translate business requirements into resilient technical solutions that power audience discovery and measurement.
Zeta Global is the AI-Powered Marketing Cloud that leverages advanced artificial intelligence and trillions of consumer signals to make marketing simpler. Founded in 2007 with offices worldwide, the company fosters a culture of trust and belonging, offering equity and wellness benefits to its employees.
Structure, integrate, and maintain scalable data solutions for credit origination, monitoring, and risk analysis.
Build and evolve the credit analytics layer including datamarts, indicators, and reusable components.
Implement data quality, governance, and observability controls while ensuring compliance with LGPD.
Jobgether is an AI-powered recruitment platform that connects candidates with job opportunities using an objective matching process. It operates with a small team, focusing on efficiency, fairness, and data privacy.
Design, build, and operate the ZMP real-time events pipeline for high-volume event data.
Provide technical leadership, own architecture, and mentor teammates through code and design reviews.
Collaborate with product, data, and consumer teams to turn SLAs into working systems.
Zeta Global is an AI-powered marketing cloud that leverages artificial intelligence and trillions of consumer signals to help marketers acquire, grow, and retain customers more efficiently. Founded in 2007 and headquartered in New York City with offices worldwide, the company fosters a culture of trust and belonging with a focus on diversity and inclusion.
Design and build scalable data pipelines using Apache Spark and cloud technologies.
Develop and maintain data integrations and transformation processes from multiple sources.
Collaborate with technical teams to ensure data quality and optimize data architecture.
Jobgether is an AI-powered job matching platform that connects candidates with hiring companies. They focus on using technology to streamline the hiring process and offer a remote work environment.
Build the data control plane for governance and discovery.
Create paved roads for data products with SDKs and CI/CD.
Ensure every data product is observable by default.
Kargo creates powerful moments of connection between brands and consumers to build businesses. With 600+ employees, the company uses a creative science approach to deliver innovative ad experiences across CTV, eCommerce, social, and mobile.