Tata Consultancy Services is hiring for a PySpark Developer. Desired experience: 4 to 8 years. Locations: Hyderabad, Chennai, Kolkata, Mumbai, Bhubaneswar, and Pune.
Job Description: As a Data Engineer, you will design, develop, and maintain data solutions for data generation, collection, and processing in a Big Data environment using predominantly PySpark/Python. Your role involves creating data pipelines, ensuring data quality, and implementing ETL processes to migrate and deploy data across systems using PySpark.
Roles & Responsibilities: - Design, develop, and maintain robust, scalable, high-performance Data Pipelines using PySpark. - Create data pipelines, ensure data quality, and implement ETL processes to migrate and deploy data across systems. - Migrate Ab Initio ETL applications into PySpark-based data pipelines. - Migrate on-premise workloads to cloud environments (AWS, Databricks, Snowflake) based on use cases. - Collaborate with cross-functional teams to identify and resolve data-related issues. - Stay updated with the latest advancements in data engineering and integrate innovative approaches for sustained competitive advantage.
Qualifications: - 4+ years of professional experience in Hadoop and PySpark/Python development. - Proven expertise in PySpark with experience handling large volumes of data. - 3+ years of working experience with AWS, Databricks/Snowflake, and Airflow. - Familiarity with CI/CD pipelines and version control systems (e.g., Git). - Strong debugging and problem-solving skills. - Excellent communication and collaboration skills.
Required Qualifications
4+ years of PySpark/Python development
3+ years of AWS, Databricks/Snowflake, and Airflow experience
Hadoop ecosystem expertise
Ab Initio ETL migration experience
CI/CD pipelines and Git version control
Large-scale data pipeline design
Preferred Qualifications
Cloud workload migration to AWS/Databricks/Snowflake