Staff Operational Support Engineer (L2) (37359129)
Job Description
OptiView is building a dedicated Operational Support (L2) team responsible for the stability, availability, and operational excellence of our 24/7 live video streaming, ads, player, and realtime delivery platforms.
\tAs an Operational Support Engineer (L2), you take endtoend ownership of customerimpacting production incidents once they are triaged by Level 1 support. You operate directly on production systems, lead live incident resolution, and act as the operational bridge between Support, Engineering, DevOps, and customers, particularly during highimpact live events.
\tThis is a handson, customerfacing role focused on incident ownership, production operations, automation, and operational scalability, not just reactive troubleshooting.
\tKey Responsibilities
\tIncident & Operational Support
\tTake ownership of escalated customer issues from Level 1 Support and drive them to resolution
\tTroubleshoot and resolve complex, high-impact production incidents affecting live streams, VOD playback, ad insertion, DRM, and real-time WebRTC services
\tOperate directly on production environments, including configuration changes, CDN adjustments, and corrective actions, following established operational procedures, including executing mitigations and emergency changes during live incidents when customer impact requires immediate action
\tLead or actively contribute to live incident bridges involving customers, internal teams, and partners
\tProvide clear, timely communication during incidents, including status updates and customer-facing explanations
\tInfrastructure as Code & Production Operations
\t· Work fluently with Infrastructure as Code (IaC) to understand, troubleshoot, and safely modify production environments
\t
\t· Leverage tools and frameworks such as:
\to Terraform
\to Helm
\to Kubernetes manifests
\to GitOps workflows
\to CI/CD and deployment pipelines
\tUse IaC as the primary mechanism for safe, auditable, and repeatable operational changes
\tCollaborate with Engineering and DevOps to improve deployment reliability and operational safety
\tValidate and execute infrastructure or configuration changes through codified workflows
\tAI-Driven Operations & Automation
\tLeverage AI tools and automation to enhance operational efficiency and incident response
\tContribute to and use:
\tAI-assisted incident triage and classification
\tAutomated runbook execution
\tAI-based pattern detection across incidents
\tIntelligent alert correlation and noise reduction
\tUse AI to:
\tGenerate or improve incident communications
\tAccelerate troubleshooting workflows
\tIdentify recurring patterns and systemic issues
\tDrive adoption of automation-first and AI-augmented operational practices
