Location: Pan India • 8+ years of experience in Azure Databricks technologies. • Hands-on experience in PySpark. • Strong understanding of ETL, BI, and Data Warehousing technologies. • Hands-on experience working with JSON and XML files as data sources. • Experience utilizing Auto Loader for incremental loading and explode functions for data flattening. • Proficiency in loading and managing Delta Live Tables. • Extensive experience with PySpark and Spark DataFrames. • Design, implement, and maintain robust data pipelines for ingestion, processing, and transformation within Azure. • Collaborate with data scientists and analysts to define data requirements and build effective data workflows. • Develop and maintain scalable data storage solutions including Azure SQL Database, Azure Data Lake, and Azure Blob Storage. • Leverage Azure Databricks to execute and optimize ETL operations. • Implement comprehensive data validation and cleansing procedures to ensure data quality, integrity, and reliability. • Optimize data pipelines for improved scalability, efficiency, and cost-effectiveness. • Monitor, troubleshoot, and resolve data pipeline issues to guarantee data consistency and availability. • Strong oral and written communication skills.
Required Qualifications
8+ years of Azure Databricks experience
PySpark expertise
ETL, BI, and Data Warehousing knowledge
JSON and XML data handling
Delta Live Tables implementation
Azure cloud storage services
Data pipeline architecture and development
Preferred Qualifications
Incremental loading with Auto Loader
Data validation and cleansing procedures
Pipeline optimization for scalability and cost-efficiency
Cross-functional collaboration with data science teams