Vir Biotechnology, Inc. logo

Associate Director, Principal Data Engineer

Vir Biotechnology, Inc.
August 20, 2026
Remote friendly (San Francisco, CA)
United States
IT
Associate Director, Principal Data Engineer

Responsibilities:
- Build AI-enabled research workflows and agents to help scientists search, analyze, summarize, and connect Vir-generated information.
- Develop AI agents for TCE research insights, Experimenta QC/QA review, and scientific/regulatory reporting.
- Design trusted LLM, RAG, knowledge-search, and agentic AI workflows with traceability, validation, governance, and human oversight.
- Lead development and integration of scientific platforms (e.g., SeqAssembler, HPD/dAIsY, SASTRY, OPAL Miner).
- Build scalable research and clinical data pipelines for bioinformatics, genomics, machine learning, and scientific decision-making.
- Architect and operate AWS cloud infrastructure for AI/ML, bioinformatics, clinical analysis, and large-scale scientific data.
- Strengthen data quality, metadata, lineage, observability, security, and governance.
- Partner cross-functionally to align data, cloud, application, and AI solutions with program needs.
- Provide senior technical leadership (architecture, engineering practices, mentoring, roadmap).

Qualifications:
- BS/MS in Computer Science, Data Engineering, Bioinformatics, Computational Biology, Engineering, or related technical field.
- 12+ years in data engineering, software engineering, scientific computing, cloud architecture, or scientific application.
- Strong hands-on Python and/or Java; scalable data pipelines, APIs/services, production-grade data.
- Experience with AI/ML-enabled scientific applications using LLMs (GPT, Claude, Gemini, etc.); RAG, knowledge search, agentic workflows, governance.
- Experience integrating LIMS/ELN/scientific data platforms or Experimenta for capture, analysis, and quality review.
- Strong AWS experience (compute/storage/networking/security/data processing/monitoring/automation/operations).
- Experience with Databricks, Snowflake, Redshift, Spark, Airflow, Nextflow, or similar.
- Experience supporting bioinformatics/genomics/ML/clinical or other scientific data.
- Knowledge of data modeling, metadata management, data quality, observability, governance, lineage.
- Demonstrated technical leadership and cross-functional communication.