
Innovation Starts Here
PySpark Data Engineer
Share this job
We Are
At Code1 Tech, we drive innovations that shape the future of enterprise technology. Our expertise spans Data Engineering, AI/ML, Cloud Solutions, and Full-Stack Development. We empower businesses with cutting-edge technology solutions, enabling digital transformation at scale. Join us to build impactful products with a passionate team of engineers and innovators.
About the Role
We are looking for a skilled PySpark Data Engineer with strong expertise in designing, developing, and optimizing large-scale data processing pipelines. The ideal candidate should have hands-on experience with Apache Spark, Python, Oracle SQL, and HDFS, along with a proven track record of building scalable data ingestion frameworks and processing high-volume datasets in production environments.
Experience: 6+ Years
Key Responsibilities:
- Design, develop, and maintain scalable Apache Spark (PySpark) data pipelines for batch and large-scale data processing.
- Build and enhance data ingestion frameworks integrating data from multiple source systems.
- Develop optimized ETL/ELT workflows for high-volume enterprise datasets.
- Process, transform, and optimize datasets exceeding 500GB while ensuring performance and reliability.
- Perform data cleansing, normalization, validation, and formatting to improve data quality.
- Write efficient Oracle SQL queries for data extraction, transformation, and analysis.
- Work with HDFS and distributed data processing environments.
- Optimize Spark jobs by tuning partitioning, shuffling, caching, and resource utilization.
- Automate manual data processing workflows using Python and Shell Scripting.
- Monitor, troubleshoot, and optimize production data pipelines for performance and scalability.
- Collaborate with data architects, analysts, and business stakeholders to deliver robust data solutions.
Mandatory Technical Skills:
- Hands-on expertise designing, building, and maintaining Apache Spark pipelines in production environments.
- Experience building scalable data ingestion frameworks integrating multiple source systems.
- Strong understanding of Spark Architecture including Driver/Executors, DAG, Partitioning, Shuffles, Caching, and Resource Management.
- Experience processing and transforming datasets larger than 500GB.
- Strong Oracle SQL and HDFS knowledge.
- Experience handling data cleansing, normalization, and data formatting processes.
- Strong Python, PySpark, and Shell Scripting skills.
- Experience automating manual data processing workflows using Python.
Ready to make an impact?
Apply now and join our team of innovators.