Skip to main content
Eli Lilly and Company logo

Data Engineer

Eli Lilly and Company
September 16, 2026
Full-time
Remote friendly (San Francisco, CA)
Worldwide
IT
Role: Data Engineer supporting AI-driven drug discovery in Lilly and NVIDIA's partnership, focusing on building scalable data platforms for machine learning and scientific research. Responsibilities: develop and maintain pipelines for ingesting, transforming, and delivering biological, chemical, and experimental data; ensure data accuracy, traceability, and accessibility; enable rapid data flow from automated labs to AI models; implement data versioning and lineage for reproducibility; collaborate with laboratory and AI teams to optimize data processes. Requirements: Bachelor’s in Computer Science, Data Science, or related field; 5+ years in data engineering; proficiency in Python, SQL, distributed processing frameworks (Spark, Dask), cloud platforms (AWS, Azure), and data pipeline orchestration tools (Airflow, Dagster); experience with scientific data formats (SMILES, InChI, RDKit) is a plus. High-Value: supporting AI models in drug discovery; biological and chemical data management; cloud-based high-performance data systems; collaboration with scientists and lab teams. WorkSetup: hybrid model based in Silicon Valley, with three days onsite and two remote weekly.