PySpark Data Engineer

TATA Consultancy Services Ltd

Pune, Maharashtra, India

Type: Full-Time Arrangement: Hybrid
3 - 6 Yrs800000 - 1500000(Annually)

Posted: 4 Aug 2026

Job Description

Design, develop, and maintain scalable ETL/ELT data pipelines using PySpark. Process and transform large datasets from various structured and unstructured sources. Build robust data workflows and optimize Spark jobs for performance and scalability. Develop and maintain data models, data lakes, and data warehouse solutions. Collaborate with Data Scientists, Analysts, and Business Teams to understand data requirements. Ensure data quality, integrity, security, and governance standards are maintained. Troubleshoot and resolve performance bottlenecks in Spark applications. Monitor data pipelines and automate operational processes. Implement CI/CD practices for data engineering workflows. Create and maintain technical documentation for data pipelines and processes.

Required Qualifications

  • PySpark expertise
  • ETL/ELT pipeline development
  • Python programming
  • SQL proficiency
  • Spark performance tuning
  • Data lake and data warehouse architecture
  • Data governance and security standards
  • CI/CD implementation for data workflows
  • Big data processing frameworks
  • Strong troubleshooting and debugging skills

Preferred Qualifications

  • Apache Airflow experience
  • Cloud platform proficiency (AWS/Azure/GCP)
  • Delta Lake or Apache Iceberg knowledge
  • Docker and Kubernetes familiarity
  • Advanced data modeling techniques
  • Agile/Scrum methodology experience

Skills Required

PySparkPythonSQLETL/ELT DevelopmentSpark OptimizationData Lakehouse DesignCI/CD PipelinesData GovernanceBig Data ProcessingCloud PlatformsGitLinux/UnixData ModelingAutomation ScriptingTechnical DocumentationProblem SolvingCross-functional Collaboration