Sr. Principal Data Engineer - Lakehouse Architecture
Eli Lilly and Company
August 06, 2026
Remote friendly (Indianapolis, IN)
United States
IT
What You Will Be Doing
- Design and implement comprehensive Lakehouse architecture solutions using Databricks and Snowflake.
- Define and carry out medallion architecture standards (Bronze/Silver/Gold) across data domains, ensuring data quality, lineage, and discoverability.
- Lead Unity Catalog governance design: schemas, access control policies, and data contracts.
- Build and maintain real-time and batch data processing systems using Apache Spark (PySpark/Scala), Kafka, and Databricks Structured Streaming.
- Architect scalable pipelines for structured, semi-structured, and unstructured data to deliver AI-ready data.
- Develop data transformation workflows using DBT, Airflow, or Databricks.
- Implement data governance frameworks (quality monitoring, lineage tracking, time travel, and security protocols).
- Build pipeline testing frameworks (unit tests, Great Expectations/dbt tests, schema validation).
- Define and publish data SLAs/SLOs; own incident response and root-cause analysis for pipeline failures.
- Drive adoption of modern data engineering standards (Infrastructure as Code, CI/CD, automated testing).
- Collaborate with data scientists/analysts/business partners; mentor 3β5 data engineers.
How You Will Succeed
- Mentor junior engineers; facilitate knowledge sharing.
- Lead multi-functional initiatives; drive technical consensus; growth mindset.
What You Should Bring
- Pharmaceutical/life sciences domain knowledge.
- Streaming tech: Kafka, Databricks Structured Streaming; data cataloging tools.
- Python and SQL; Apache Spark (PySpark/Scala).
- Cloud (AWS/Azure): S3, Glue, Redshift, IAM, CloudWatch.
- IaC (CloudFormation); containers/orchestration (Docker, Kubernetes).
- Databricks expertise (clusters, Delta Lake, Unity Catalog, Workflows, MLflow integration).
- Governance/security/compliance (e.g., GxP, HIPAA, GDPR).
- Airflow orchestration; strong dbt; data quality testing (Great Expectations/dbt tests); data observability.
- Delta Sharing or Apache Iceberg; DataOps/data mesh principles.
- ML platform integration exposure (MLflow, feature stores, or model serving).
Your Basic Qualifications
- Masterβs degree in CS/Engineering or related field.
- 3+ years Lakehouse experience (Databricks/Snowflake or similar).
- 7+ years data engineering experience; at least 3 years senior/lead capacity.
- Design and implement comprehensive Lakehouse architecture solutions using Databricks and Snowflake.
- Define and carry out medallion architecture standards (Bronze/Silver/Gold) across data domains, ensuring data quality, lineage, and discoverability.
- Lead Unity Catalog governance design: schemas, access control policies, and data contracts.
- Build and maintain real-time and batch data processing systems using Apache Spark (PySpark/Scala), Kafka, and Databricks Structured Streaming.
- Architect scalable pipelines for structured, semi-structured, and unstructured data to deliver AI-ready data.
- Develop data transformation workflows using DBT, Airflow, or Databricks.
- Implement data governance frameworks (quality monitoring, lineage tracking, time travel, and security protocols).
- Build pipeline testing frameworks (unit tests, Great Expectations/dbt tests, schema validation).
- Define and publish data SLAs/SLOs; own incident response and root-cause analysis for pipeline failures.
- Drive adoption of modern data engineering standards (Infrastructure as Code, CI/CD, automated testing).
- Collaborate with data scientists/analysts/business partners; mentor 3β5 data engineers.
How You Will Succeed
- Mentor junior engineers; facilitate knowledge sharing.
- Lead multi-functional initiatives; drive technical consensus; growth mindset.
What You Should Bring
- Pharmaceutical/life sciences domain knowledge.
- Streaming tech: Kafka, Databricks Structured Streaming; data cataloging tools.
- Python and SQL; Apache Spark (PySpark/Scala).
- Cloud (AWS/Azure): S3, Glue, Redshift, IAM, CloudWatch.
- IaC (CloudFormation); containers/orchestration (Docker, Kubernetes).
- Databricks expertise (clusters, Delta Lake, Unity Catalog, Workflows, MLflow integration).
- Governance/security/compliance (e.g., GxP, HIPAA, GDPR).
- Airflow orchestration; strong dbt; data quality testing (Great Expectations/dbt tests); data observability.
- Delta Sharing or Apache Iceberg; DataOps/data mesh principles.
- ML platform integration exposure (MLflow, feature stores, or model serving).
Your Basic Qualifications
- Masterβs degree in CS/Engineering or related field.
- 3+ years Lakehouse experience (Databricks/Snowflake or similar).
- 7+ years data engineering experience; at least 3 years senior/lead capacity.