The Associate Director will serve as the technical authority for Observability and Event Management. Lead the technology strategy, engineering, and operational excellence of the enterprise’s monitoring, telemetry, and event lifecycle capabilities. This role ensures our platforms, applications, and infrastructure are measurable, predictable, and resilient – enabling proactive detection, automated remediation, and data-driven engineering decisions.
Responsibilities:
- Own the observability strategy across logs, metrics, traces, events, and digital experience monitoring.
- Position observability as an enterprise service with clear service definitions, support models, standards, and adoption expectations.
- Influence the architecture and lead the engineering portion of modern telemetry pipelines: OpenTelemetry, distribution tracing, log aggregation, and metrics platforms.
- Lead the technology portion for Event Management, including event detection, correlation, enrichment, noise reduction, and automated response.
- Support and implement AIOps capabilities to improve signal-to-noise ratio, reduce toil, and accelerate root-cause identification.
- Drive automation-first approaches for alerting, remediation, and event routing.
- Partner with architects, platform, and cloud engineering teams, to ensure observability and monitoring coverage are embedded into new and existing services, applications, infrastructure, and platforms.
- Lead tool selection, roadmap, and vendor management for observability and event platforms.
- Lead and develop a team responsible for enterprise observability services, including coaching, performance management, prioritization of team capacity, skills development, and alignment of team execution to enterprise technology priorities.
- Mentor engineers and influence engineering culture towards proactive reliability and data-driven operations.