We build infrastructure that delivers massive amounts of web data to the companies training the world’s most powerful AI models. We're lean, technical, and move fast.
Maintain, optimize, and troubleshoot database queries and related data systems to support efficient data access, processing, and reliability.
Assist in creating, maintaining, and improving data pipelines used to collect, process, transform, validate, and deliver large-scale datasets.
Support web scraping and data collection initiatives, including developing, testing, and maintaining scripts or tools used to gather publicly available data.