OBSERVABILITY / AZURE INFRASTRUCTURE
Observability that supports incident response
Connect telemetry to service ownership, useful alerts and recovery decisions.
Start with the service
Document what the application must do for its users. Use that context to choose signals such as response time, errors and connectivity, and agree thresholds with the teams responsible for service quality.
Give alerts an owner
An alert needs a recipient, a triage path and a next action. Identify which issues belong to SRE and which require application or networking expertise. Include links to the relevant dashboard and runbook.
Combine perspectives
External availability checks, Azure metrics and application telemetry describe different parts of the system. During triage, compare them to narrow the affected path instead of relying on one signal.
Review after incidents
Ask whether monitoring detected the failure and whether the notification helped the responding team. Adjust thresholds, routing and missing signals using incident evidence. Connectivity incidents in my work led to additional Azure VPN Gateway metric-based alerting.