Skip to content

Observability 101

Six sessions to learn logs, metrics and traces by instrumenting a running marketplace and investigating incidents on your own stack.

You inherit a marketplace that is already in production, with customers placing orders every minute and almost no instrumentation. Your group adds one layer of observability per session, then uses all of it to find out what broke when incidents hit your stack.

  1. 1Setup
    blind stack
  2. 2Logs
    Logs
  3. 3Metrics
    LogsMetrics
  4. 4Traces
    LogsMetricsTraces
  5. 5Connect
    LogsMetricsTracesSLOs
  6. 6Incident
    LogsMetricsTracesSLOs
Your stack gains one layer of observability per session, then you use all of it.

Groups, labs and grading are on the course format page.

  1. 1. Why observability?Observability versus monitoring, the three pillars and environment setup.
  2. 2. LogsSoonStructured logging, log pipelines and searching logs in Datadog.
  3. 3. MetricsSoonMetric types, RED and USE, dashboards and a first monitor.
  4. 4. TracesSoonDistributed tracing, context propagation and trace-log correlation.
  5. 5. Connecting the pillarsSoonCorrelation, SLOs, error budgets and alerting that people trust.
  6. 6. Incident responseSoonThe incident lifecycle and the final graded exercise.

Each challenge and incident ends in a write-up: a Datadog Notebook in your team’s organization, with live queries as evidence and a link to the merge request that holds your fix. The final exercise is written the same way, as a post-mortem. The write-ups page shows the structure.