Create and maintain optimal data pipeline architecture. Assemble large, complex data sets that meet functional / non-functional business requirements. Identify, design, and implement internal process improvements: automating manual processes, optimizing data delivery, re-designing infrastructure for greater scalability, etc. Build the infrastructure required for optimal ingestion, transformation, and Publishing of data from a wide variety of data sources using Python/Spark and AWS ‘big data’ technologies. Build analytics tools that utilize the data pipeline to provide actionable insights into customer acquisition, operational efficiency and other key business performance metrics. Work with stakeholders including the Executive, Product, Data and Design teams to assist with data-related technical issues and support their data infrastructure needs. Keep our data separated and secure across national boundaries through multiple data centers and AWS regions. Create data tools for analytics and data scientist team members that assist them in building and optimizing our product into an innovative industry leader. Work with data and analytics experts to strive for greater functionality in our data systems
Required Qualifications
Python expertise
Apache Spark expertise
AWS Big Data services proficiency
Data pipeline architecture design
ETL/ELT process optimization
Data security & compliance management
Preferred Qualifications
Azure Data Factory expertise
Databricks proficiency
SQL & NoSQL database knowledge
Experience with cloud data warehousing
Strong analytical & problem-solving skills
Skills Required
PythonApache SparkAWS Big DataData Pipeline ArchitectureETL/ELTCloud InfrastructureData SecurityAnalytics DevelopmentProcess AutomationData IntegrationSQLDatabricksAzure Data ServicesMachine Learning SupportAgile Methodologies