HPC Platform Product Lead
Zoetis
August 20, 2026
Remote friendly (Kalamazoo, MI)
United States
IT
POSITION RESPONSIBILITIES
- Product ownership & value management (25%): Own product vision/roadmap/service catalog/backlog with acceptance criteria; translate utilization/throughput/outcomes into investment asks, funding recommendations, and exec-ready value stories (showback/chargeback as applicable).
- HPC architecture & technical standards (20%): Define target-state architecture and standards across compute, storage, network, security, and tooling; drive multi-year evolution.
- SLURM scheduling leadership & L3 support (15%): Own scheduling strategy (partitions/queues, QoS, fairshare, accounting, reservations); lead configuration/troubleshooting and job efficiency analysis; provide technical backup.
- Operations, reliability, and observability (20%): Drive operational excellence with monitoring/alerting, dashboards, SLO/SLI, incident/problem follow-up, runbooks, and continuous improvement.
- Capacity planning, procurement, and lifecycle management (10%): Plan/execute compute & storage growth; manage allocations; coordinate procurement and lifecycle refresh.
- Governance, access, and user enablement (10%): Establish onboarding/entitlements/RBAC and auditability; improve support processes (ticketing/escalations/knowledge base); publish workload best practices (CPU/GPU optimization, execution standards, data movement).
EDUCATION & EXPERIENCE
Required:
- Bachelorโs in CS/Engineering/Information Systems (or equivalent).
- 7+ years in infrastructure/platform engineering or adjacent technical product roles.
- 5+ years HPC experience (on-prem/cloud/hybrid).
- Experience partnering with scientists/engineers to convert workload needs into prioritized platform capabilities and measurable value.
Preferred:
- Masterโs degree.
- HPC in regulated/highly governed environments.
- Product Owner/Agile delivery (backlog/acceptance criteria/roadmapping).
TECHNICAL SKILLS
- Strong SLURM administration: partitions/queues, QoS, fairshare, accounting, reservations; monitoring/troubleshooting; job efficiency analysis; user guidance.
- HPC fundamentals: Linux admin, CPU/GPU, high-speed networking, shared/parallel storage, capacity/performance planning.
- Automation/scripting (Bash/Python) and operational tooling; infra-as-code where applicable.
- Observability: utilization telemetry, job analytics, dashboards/alerting, operational reporting.
- Data lifecycle & protection: retention, tiered storage, archive/backup, restore testing, runbooks.
- Product/delivery: requirements discovery, value storytelling, roadmap/backlog prioritization, stakeholder management, communicating trade-offs and ROI.
- Product ownership & value management (25%): Own product vision/roadmap/service catalog/backlog with acceptance criteria; translate utilization/throughput/outcomes into investment asks, funding recommendations, and exec-ready value stories (showback/chargeback as applicable).
- HPC architecture & technical standards (20%): Define target-state architecture and standards across compute, storage, network, security, and tooling; drive multi-year evolution.
- SLURM scheduling leadership & L3 support (15%): Own scheduling strategy (partitions/queues, QoS, fairshare, accounting, reservations); lead configuration/troubleshooting and job efficiency analysis; provide technical backup.
- Operations, reliability, and observability (20%): Drive operational excellence with monitoring/alerting, dashboards, SLO/SLI, incident/problem follow-up, runbooks, and continuous improvement.
- Capacity planning, procurement, and lifecycle management (10%): Plan/execute compute & storage growth; manage allocations; coordinate procurement and lifecycle refresh.
- Governance, access, and user enablement (10%): Establish onboarding/entitlements/RBAC and auditability; improve support processes (ticketing/escalations/knowledge base); publish workload best practices (CPU/GPU optimization, execution standards, data movement).
EDUCATION & EXPERIENCE
Required:
- Bachelorโs in CS/Engineering/Information Systems (or equivalent).
- 7+ years in infrastructure/platform engineering or adjacent technical product roles.
- 5+ years HPC experience (on-prem/cloud/hybrid).
- Experience partnering with scientists/engineers to convert workload needs into prioritized platform capabilities and measurable value.
Preferred:
- Masterโs degree.
- HPC in regulated/highly governed environments.
- Product Owner/Agile delivery (backlog/acceptance criteria/roadmapping).
TECHNICAL SKILLS
- Strong SLURM administration: partitions/queues, QoS, fairshare, accounting, reservations; monitoring/troubleshooting; job efficiency analysis; user guidance.
- HPC fundamentals: Linux admin, CPU/GPU, high-speed networking, shared/parallel storage, capacity/performance planning.
- Automation/scripting (Bash/Python) and operational tooling; infra-as-code where applicable.
- Observability: utilization telemetry, job analytics, dashboards/alerting, operational reporting.
- Data lifecycle & protection: retention, tiered storage, archive/backup, restore testing, runbooks.
- Product/delivery: requirements discovery, value storytelling, roadmap/backlog prioritization, stakeholder management, communicating trade-offs and ROI.