Responsibilities:
- Develop, validate, and deploy predictive models for early-stage biologics process decisions (upstream performance, downstream purification, and CQA outcomes).
- Design hybrid models combining first-principles process understanding with data-driven methods.
- Build analytical frameworks supporting digital twin concepts and in-silico process optimization.
- Define data strategies for new programs (data to collect, structure, and connectivity across campaigns).
- Improve data quality, integration, and accessibility; help implement improved data infrastructure.
- Ensure solutions follow GxP and regulatory documentation/traceability requirements.
- Architect AI-driven analytical workflows (classical stats, supervised/unsupervised ML, retrieval-augmented generation, orchestrated pipelines, hybrid mechanisms).
- Automate/accelerate scientific workflows; scope, build, and validate solutions.
- Design statistically rigorous, resource-efficient experiments (DoE, Bayesian optimization, active learning).
- Support process characterization with statistical analysis, risk assessment, and CPPβCQA relationship identification.
- Translate analyses into actionable insights; document/present/defend work for internal reviews and regulatory submissions.
Qualifications:
Required:
- BS in CS/related + 7 years; MS + 6 years; or PhD + 2 years in IT/application development.
- Hands-on experience building/deploying ML/data science solutions in scientific/engineering environments.
- Expert Python (NumPy, pandas, scikit-learn, PyTorch/TensorFlow) and modern data engineering (cloud, big data, orchestration).
- Business analytics tools: R, Dataiku, AWS SageMaker, Spark, Tableau.
- Experience with DoE, Bayesian methods, or active learning.
- Familiarity with knowledge graphs and retrieval-augmented/orchestrated AI/LLM systems.
- GxP-regulated data science experience; FDA/EMA knowledge of process validation, CPV, and control strategy.
- Familiarity with MLOps and model lifecycle/deployment in regulated/enterprise environments.
- Ownership, solution-architect mindset, scientific integrity, credibility-based influence, and bias for high-impact problems.
Preferred:
- MS/PhD in data science/biostatistics/engineering/computational biology.
- 5+ years hands-on ML/data science in scientific/engineering environments.
- Experience with lab/manufacturing data (LIMS, MES, DeltaV/historian, eBR) and scalable process analytics pipelines.
- Biologics/bioprocess development experience.
- CMC/process characterization, scale-up, technology transfer, and/or BLA/IND support.
- Biopharmaceutical data familiarity; technology transfer/workflow and PPQ/PV exposure.
- Track record communicating analytical work (publications/regulatory/technical reports).