Responsibilities:
- Serve as primary responder in scheduled on-call for production incidents across SOX- and FDA-regulated clinical applications; triage, escalate, and remediate per runbooks/incident playbooks.
- Execute approved production runbooks and maintenance scripts; document privileged actions for SOX ITGCs and FDA audit-trail requirements.
- Implement/maintain application-layer instrumentation (dashboards, alert thresholds, SLO monitors) and monitor SLO/error-budget consumption; escalate when budgets are at risk.
- Support releases/deployments: verify deployment health, monitor post-deploy behavior, and perform rollback when required.
- Participate in post-incident reviews: reconstruct timelines, identify contributing factors, and drive remediation action items to closure.
- Maintain audit-ready access logs, privileged-action records, and role-change documentation.
- Automate recurring operational work; track manual-intervention rate and reduce it by retiring runbook steps into automation.
- Author/update runbooks and known-issue documentation; improve operate discipline.
- Collaborate with application engineering to validate production fixes; contribute to on-call tooling and alert-quality improvements.
- Use AI coding assistants/agentic workflows for monitoring, triage, runbooks, and automation scripting (review as quality gate).
Required Qualifications:
- BS in CS/IS/SE or related, or equivalent experience.
- 5+ years in SRE/production ops/DevOps/platform engineering (production support).
- 2+ years on-call primary responder experience for Tier 1/business-critical systems.
- Experience executing production runbooks/maintenance/change procedures with privileged-action controls.
- Experience with production monitoring/alerting dashboards in an observability platform.
- Experience in post-incident review/blameless retrospectives.
- Hands-on use of AI coding assistants for operational automation.
Preferred Qualifications:
- Familiarity with SLO/error budgets and alerting design.
- Experience reducing toil via automation.
- Experience in SOX-controlled IT or CLIA/CAP-regulated environments.
- Cloud-native observability at the application instrumentation layer; telemetry/tracing/APM.
- Domain experience in clinical diagnostics/lab information systems/digital health.
- Familiarity with SAST/DAST and secure CI/CD; deployment pipelines/container orchestration/release automation.
- Relevant cloud/security certifications.
Physical/Other:
- Ability to sit/stand/work at a computer for extended periods.
- On-call includes required after-hours response (evenings/weekends); periodic travel may be required.