Position: AI / Data Engineer (Data Discovery Services)
Job Responsibilities
- Build and maintain Python data pipelines to pull enterprise metadata, enrich with taxonomy tags/ownership, and publish to the discovery platform.
- Tune/optimize search indexes (analyzers, boost fields, query testing) to improve result relevance.
- Build a semantic knowledge layer (document chunking, vector embeddings, semantic metadata enrichment) to support RAG/LLM retrieval.
- Maintain integrations syncing ontology/taxonomy changes into the platform.
- Troubleshoot data pipeline issues across Databricks and AWS Glue; trace root causes; add data quality checks.
- Build API endpoints and Model Context Protocol (MCP) servers exposing search/metadata to apps and AI agents.
- Design metadata pipelines for cross-domain dataset relationships and confidence-scored join catalogs.
- Analyze search patterns, collect feedback, and continuously improve discovery.
Required Qualifications
- Bachelorโs degree (CS/Data Science/Info Science/Engineering or related); Masterโs preferred.
- Proven production data pipeline delivery; strong data/software engineering background.
- Proficient Python; strong SQL.
- Experience with Databricks and AWS Glue; ETL/ELT fundamentals, data modeling, data quality, orchestration.
- AWS familiarity (S3, Lambda, API Gateway, Glue); OpenSearch/Elasticsearch experience.
- Metadata management/data cataloging knowledge; problem-solving, learning mindset; teamwork/communication.
Benefits (if eligible)
- Health coverage; wellbeing support programs; 401(k) and other protection benefits.
- Paid time off (flexible time off or annual vacation + holidays, depending on location/role).