Data Engineer
ืคืืจืกื ืืชืืื ยท 0 ืืืขืืืื
ืฉืืืจื, ืืืฉื ืื ืืืืงืช ืืชืืื. ืคืชืืืช ืืฉืืื ืืื ื ืืืงืืช ืืื ืฉื ืืืช.
- At least 3 years of experience as a Data Engineer
- 3 years of experience with Object-Oriented Programming (OOP) languages
- 3 years of experience with Python
- Spark
- BSc in Computer Science, Engineering, Mathematics, or Statistics
- AWS services (Athena, Glue, Step Functions, EMR, Redshift, RDS, Lambda, SQS/SNS)
- Designing, developing, and optimizing complex solutions for transferring/processing large data volumes
- Optimization techniques and working with partitions across data formats such as Parquet, Avro, HDF5, Delta Lake
- Docker, Linux, CI/CD tools, and Kubernetes
- Data pipeline tools such as Airflow, Kubeflow
ืืืืฅ ืืชืืืืจ ืืืฉืจื ยท ืืชืขืืื ืืืืืืืืช
ืชืืืืจ ืืืฉืจื ืืืื
ืืืฉืจื ืืืงืืจืืช ยท ื ืฉืืจ ืืขืืืWe're Hiring: Data Engineer to Join Our AI Team!
๐ Lod | Hybrid - 1 day WFH
Join a team building the data infrastructure behind cutting-edge AI projects! You'll be responsible for ingesting large volumes of new data, gaining deep understanding and insight into it in close collaboration with Data Scientists, and designing and developing critical, large-scale data processes - both in cloud and on-prem environments.
Requirements:
โ At least 3 years of experience as a Data Engineer - required
โ 3 years of experience with Object-Oriented Programming (OOP) languages - required
โ 3 years of experience with Python - required
โ Hands-on experience with Spark for large-scale data - required
โ BSc in Computer Science, Engineering, Mathematics, or Statistics - required
Nice to have:
โญ At least 2 years of hands-on experience with AWS services such as Athena/Glue/Step Functions/EMR/Redshift/RDS - significant advantage
โญ Deep understanding of designing, developing, and optimizing complex solutions for transferring/processing large data volumes
โญ Understanding of optimization techniques and working with partitions across data formats such as Parquet, Avro, HDF5, Delta Lake
โญ Experience with Docker, Linux, CI/CD tools, and Kubernetes
โญ Experience with data pipeline tools such as Airflow, Kubeflow
โญ Understanding of machine learning concepts and processes
โญ Familiarity with GenAI solutions / prompt engineering - advantage
Team notes:
โญ AWS experience with core services - Glue, Step Functions, Lambda, SQS/SNS; familiarity with Redshift and Kafka, and working with APIs is a significant advantage. Candidates without any cloud environment experience won't be relevant.
โญ Extensive experience writing functional Python code and complex ETL processes, including building generic frameworks/templates for managing ETL workflows (controls, generic process management, parameterization, etc.)
โญ Experience deploying services - significant advantage; experience deploying models on SageMaker - especially significant advantage
โญ Experience writing microservices - significant advantage
ืฉืืืืช ืขื ืืืฉืจื
- ืืืฉืจื ืื ืฆืืื ื ืฉืืจ. ืื ืื ื ืืฆืืืื ืฉืืจ ืจืง ืืฉืืืขืกืืง ืืคืจืกื ืืืชื.
- At least 3 years of experience as a Data Engineer, 3 years of experience with Object-Oriented Programming (OOP) languages, 3 years of experience with Python, Spark, BSc in Computer Science, Engineering, Mathematics, or Statistics