Senior Data Scientist, Biologics Discovery
Johnson & Johnson
September 01, 2026
On-site
Titusville, NJ
Clinical Research and Development
Senior Data Scientist (Biologics Discovery)
Position Summary
- Design featurization and dataset curation for antibody/protein sequence, constructs, assay, and biophysical data.
- Build and evaluate applied ML models; define evaluation frameworks to keep models trustworthy.
- Collaborate with data-infrastructure partners and In Silico Discovery (ISD) to strengthen ISDβs molecular property models.
Key Responsibilities
- Develop featurization and model-ready datasets.
- Specify features/labels/aggregation levels with data engineers; preserve raw representations.
- Curate, document, and version datasets for reproducibility and traceability.
- Partner with ISD on standardized, traceable training datasets; align on DS vs. ISD ownership.
- Frame ML problems around decision points in the DMTL cycle.
- Work with ontology and MLOps teams to ensure consistent semantics and reliable model deployment.
- Champion reproducibility, documentation, and responsible AI.
Qualifications
Required:
- M.S. or Ph.D. in CS/ML/Computational Biology/Bioinformatics/Statistics or related field.
- 2+ years applied ML (model development/evaluation/dataset curation) on complex scientific/biomedical data.
- Proficiency in Python (PyTorch, scikit-learn) and SQL.
- Experience converting heterogeneous experimental data into robust features/training sets; exposure to cloud training/data infrastructure.
- Understanding of evaluation/validation, leakage, and distribution shift.
- Ability to collaborate in a matrixed R&D environment.
Preferred:
- Biologics, antibody/protein sequence models, or protein language models.
- Active learning, Bayesian optimization, or sequence-based generative models.
- Familiarity with assay/biophysical data and developability endpoints.
- MLOps, experiment tracking, model monitoring.
- Ontologies/knowledge graphs for AI-ready datasets.
Benefits/Compensation (as stated)
- Base pay range: $109,000β$174,800; annual performance bonus eligibility; medical/dental/vision/life and disability/other listed insurance.
- Pension and 401(k) eligibility.
- Vacation up to 120 hours/year; sick time up to 40 hours/year; holiday/floating holidays up to 13 days/work-personal-family time up to 40 hours/year.
Locations
- Spring House, PA (strongly preferred), Titusville, NJ, or Raritan, NJ, USA; or Madrid, Spain. No remote option.
Position Summary
- Design featurization and dataset curation for antibody/protein sequence, constructs, assay, and biophysical data.
- Build and evaluate applied ML models; define evaluation frameworks to keep models trustworthy.
- Collaborate with data-infrastructure partners and In Silico Discovery (ISD) to strengthen ISDβs molecular property models.
Key Responsibilities
- Develop featurization and model-ready datasets.
- Specify features/labels/aggregation levels with data engineers; preserve raw representations.
- Curate, document, and version datasets for reproducibility and traceability.
- Partner with ISD on standardized, traceable training datasets; align on DS vs. ISD ownership.
- Frame ML problems around decision points in the DMTL cycle.
- Work with ontology and MLOps teams to ensure consistent semantics and reliable model deployment.
- Champion reproducibility, documentation, and responsible AI.
Qualifications
Required:
- M.S. or Ph.D. in CS/ML/Computational Biology/Bioinformatics/Statistics or related field.
- 2+ years applied ML (model development/evaluation/dataset curation) on complex scientific/biomedical data.
- Proficiency in Python (PyTorch, scikit-learn) and SQL.
- Experience converting heterogeneous experimental data into robust features/training sets; exposure to cloud training/data infrastructure.
- Understanding of evaluation/validation, leakage, and distribution shift.
- Ability to collaborate in a matrixed R&D environment.
Preferred:
- Biologics, antibody/protein sequence models, or protein language models.
- Active learning, Bayesian optimization, or sequence-based generative models.
- Familiarity with assay/biophysical data and developability endpoints.
- MLOps, experiment tracking, model monitoring.
- Ontologies/knowledge graphs for AI-ready datasets.
Benefits/Compensation (as stated)
- Base pay range: $109,000β$174,800; annual performance bonus eligibility; medical/dental/vision/life and disability/other listed insurance.
- Pension and 401(k) eligibility.
- Vacation up to 120 hours/year; sick time up to 40 hours/year; holiday/floating holidays up to 13 days/work-personal-family time up to 40 hours/year.
Locations
- Spring House, PA (strongly preferred), Titusville, NJ, or Raritan, NJ, USA; or Madrid, Spain. No remote option.