>_ THE SRE EXPERIENCE

OBSERVABILITY / AZURE INFRASTRUCTURE

Observability that supports incident response

Connect telemetry to service ownership, useful alerts and recovery decisions.

By Gabriel Dinescu · Engineering notes

Start with the service

Document what the application must do for its users. Use that context to choose signals such as response time, errors and connectivity, and agree thresholds with the teams responsible for service quality.

Give alerts an owner

An alert needs a recipient, a triage path and a next action. Identify which issues belong to SRE and which require application or networking expertise. Include links to the relevant dashboard and runbook.

Combine perspectives

External availability checks, Azure metrics and application telemetry describe different parts of the system. During triage, compare them to narrow the affected path instead of relying on one signal.

Review after incidents

Ask whether monitoring detected the failure and whether the notification helped the responding team. Adjust thresholds, routing and missing signals using incident evidence. Connectivity incidents in my work led to additional Azure VPN Gateway metric-based alerting.

← Back to engineering notes