Similar Jobs
See allResponsibilities:
- Design, build, and maintain robust ETL/ELT processes to ingest, transform, and deliver data across a modern Data Lake architecture.
- Develop and optimize distributed data processing workflows using Python and PySpark to handle large-scale datasets efficiently.
- Implement and refine partitioning strategies for data lake storage frameworks to balance query performance with storage costs.
Requirements:
- Solid experience with ETL processes and data pipeline development on AWS, with strong proficiency in Python and PySpark.
- Thorough understanding of SQL, including complex queries (CTEs, window functions, aggregations) and translating workloads from legacy RDBMS.
- Hands-on experience with AWS Glue, Athena, Redshift, and Data Lake architectures.
Benefits:
- Premium Healthcare, Meal Voucher, and Maternity and Parental leaves.
- Mobile services subsidy, Sick pay, and Life insurance.
- CI&T University, Colombian Holidays, and Paid Vacations.
CI&T
CI&T helps large enterprises transform the potential of AI into real business impact. With 30 years of experience, we are 8,000 CI&Ters across more than 25 countries.