Similar Jobs

See all

Responsibilities:

  • Design, build, and maintain robust ETL/ELT processes to ingest, transform, and deliver data across a modern Data Lake architecture.
  • Develop and optimize distributed data processing workflows using Python and PySpark to handle large-scale datasets efficiently.
  • Implement and refine partitioning strategies for data lake storage frameworks to balance query performance with storage costs.

Requirements:

  • Solid experience with ETL processes and data pipeline development on AWS, with strong proficiency in Python and PySpark.
  • Thorough understanding of SQL, including complex queries (CTEs, window functions, aggregations) and translating workloads from legacy RDBMS.
  • Hands-on experience with AWS Glue, Athena, Redshift, and Data Lake architectures.

Benefits:

  • Premium Healthcare, Meal Voucher, and Maternity and Parental leaves.
  • Mobile services subsidy, Sick pay, and Life insurance.
  • CI&T University, Colombian Holidays, and Paid Vacations.

CI&T

CI&T helps large enterprises transform the potential of AI into real business impact. With 30 years of experience, we are 8,000 CI&Ters across more than 25 countries.

Apply for This Position