Project Role Description : Design, develop and maintain data solutions for data generation, collection, and processing. Create data pipelines, ensure data quality, and implement ETL (extract, transform and load) processes to migrate and deploy data across systems.
Must have skills : Data Engineering
Good to have skills : NA
Minimum 3 Year(s) Of Experience Is Required
Educational Qualification : 15 years full time education
Summary:
Build reliable data and knowledge pipelines that power enterprise AI, retrieval, and agentic applications.
Must have contributed to production data, search, or knowledge systems. Tutorial-based RAG projects and standalone prototypes are insufficient.
Hands-on experience using GitHub Copilot, Cursor or Claude Code with LlamaIndex or LangChain, vector databases and managed search platforms to generate connectors, parsing pipelines, retrieval logic and tests, while independently validating data quality, permissions and performance.
Roles & Responsibilities:
Develop pipelines for ingesting, cleaning, transforming, enriching, and indexing enterprise content. Implement document parsing, chunking, embeddings, metadata extraction, and retrieval. Build connectors to databases, files, APIs, and enterprise repositories. Apply data-quality, lineage, freshness, deduplication, and access-control rules. Monitor pipeline failures, retrieval performance, latency, and cost. Support evaluation and debugging of grounded AI responses. Write tested, maintainable, and production-ready code.
Professional & Technical Skills:
Python and SQL. Data pipelines, orchestration, APIs, search, and cloud storage. Embeddings, vector databases, hybrid retrieval, and RAG. Metadata, document processing, security, Git, CI/CD, and monitoring.