DevOps Observability Interview Questions — Practice Quiz

Reviewed by Mark Dickie · Last updated

8 questions on this page you can answer and see marked on the spot. Go to the first one

Observability in DevOps is the practice of instrumenting systems so you can answer questions about their internal state from the outside, using telemetry data the system emits while it runs. For a DevOps interview on observability, you should be able to explain the three pillars—metrics, logs, and traces—and when each one matters. Expect questions on SLOs and error budgets, how cardinality affects metric storage cost, and the difference between monitoring (watching known signals) and observability (asking new questions on the fly). You should also know common tools like Prometheus, Grafana, OpenTelemetry, Jaeger, and Loki well enough to describe a real pipeline.

ConceptWhat it meansLikely interview angle
MetricsNumeric time-series data sampled at intervalsHow do you avoid high-cardinality label explosion?
LogsStructured or unstructured event recordsWhy use structured logging and what fields are non-negotiable?
TracesRequest-scoped spans across service boundariesHow do you propagate trace context through async calls?
SLOs / error budgetsTarget reliability plus the budget for unreliabilityWalk through setting an SLO for a multi-service API.
AlertingRules that turn telemetry into notificationsWhat makes an alert actionable versus noise?

What does an observability interview typically cover?

  1. The three pillars of observability and how they complement each other.
  2. Cardinality: why adding a user_id label to a Prometheus counter can blow up storage.
  3. SLO design: picking a good SLI, setting an error budget, and deciding what happens when you burn through it.
  4. Distributed tracing basics: span context propagation, sampling strategies, and OpenTelemetry's role as a vendor-neutral standard.
  5. Alert fatigue: writing alerts that point to a specific service and a probable cause instead of a dashboard link.

How do monitoring and observability differ?

Monitoring tells you when something you already predicted has gone wrong; observability lets you investigate why an unexpected failure happened. In an interview, the clean way to frame it: monitoring needs preconfigured dashboards and thresholds, while observability gives you the raw telemetry to ask arbitrary questions at query time. Most teams need both, but an observability-first approach puts more energy into instrumentation and high-cardinality log and trace storage than into static alerts.

Key facts

  • Tarmac has 18 DevOps interview questions on this topic, 10 of them on this page, at difficulty 1–5 of 5.
  • Tarmac tracked 2,209 job postings asking for DevOps in September 2026.
  • Roles asking for DevOps advertise a median base salary of £80,000, across 671 job postings as of September 2026.
  • Tarmac last reviewed these DevOps interview questions on 5 October 2026.

At a glance

Questions10 shown · 18 in the bank
Difficulty1–5 of 5
FormatsMultiple choice, Code output, Multiple answer, Flashcard, Ordering, True / false, Short answer, Find the bug

What you'll review

  1. logs metrics traces
  2. structured logging
  3. error budgets

Practice questions

Try one before you open the answer. Pick an option and press Check; it's marked on the spot.

A user's request touches five microservices before it fails deep in the chain. What is the standard way to find every log line generated by that one request, across all five services?#

Options

Show answer

A correlation ID is a unique identifier generated at the point a request first enters the system and passed to every downstream service call, usually via a request header. Every service logs that same ID alongside its own entries, so a single search on the ID pulls every log line the request touched, across every service, into one causal chain. Distributed-tracing trace IDs serve the same correlating role. Without a shared ID, per-service request IDs differ at each hop and timestamps alone can't disambiguate concurrent traffic.

Why:

A correlation ID is a unique identifier minted at the edge of a request (an API gateway or the first service it hits) and threaded through every downstream call, usually as a header (e.g. X-Correlation-ID, or the trace ID a distributed-tracing system generates, which serves the same purpose). Every service that touches the request logs that ID alongside its own log lines, so a single search or filter on the ID reconstructs the request's full path through the system regardless of how many services it crossed. Searching each service's logs by timestamp alone is unreliable — under real concurrent traffic, many unrelated requests log in the same millisecond, and clock skew across hosts makes timestamp correlation worse, not better. Relying on each service's own locally-generated request ID gets the mechanism backwards: a request ID generated independently inside each service is, by definition, a different value per hop unless it is explicitly passed in and reused — that's exactly the bug the correlation-ID pattern exists to prevent. The claim that there is no standard mechanism for this is false; this is a solved, standard practice, and its absence is a common root cause of 'we can't tell what happened' during an incident.

A service is up but slow, and you cannot tell which downstream call is responsible. Which of the 'three pillars' of observability is designed to answer 'where in the request path is the time going?'#

Options

Show answer

Distributed tracing answers this. A trace stitches together the spans of a single request as it crosses service boundaries, with each span timed, so you can read off exactly which downstream hop dominates the latency. Metrics aggregate away the per-request path — they tell you that latency is up, not where. Logs are per-service discrete events that don't automatically reconstruct one request's journey without a shared trace or correlation id. The three pillars are logs, metrics, and traces; alerting is built on them, not a separate pillar.

Why:

The three pillars are logs, metrics, and traces, and each answers a different question. Metrics are cheap aggregate time-series — great for 'is latency up?' and dashboards/alerts, but they aggregate away the per-request path, so they tell you that something is slow, not where in a multi-service call. Logs are discrete events rich in detail, but on their own they are per-service and don't automatically reconstruct one request's journey across services. Distributed tracing is purpose-built for this: a trace ties together the spans of a single request as it crosses service boundaries, each span timed, so you can read off exactly which downstream hop dominates the latency. The claim that a single latency gauge tells you exactly which service is slow overstates what a single metric can do; the claim that grepping each service's log lines reconstructs the full causal request path automatically assumes logs auto-correlate across services (they don't without a shared trace/correlation id); the claim that alerting is the fourth pillar and pinpoints the slow hop directly is wrong — alerting is built on metrics/logs, it is not a fourth pillar and does not localize a slow span.

Two services log the same checkout event. Service A emits the first line, Service B the second. A log platform ingests both and you need to alert when amount > 100 for a given user_id. Which line lets you build that query reliably without brittle text parsing?#

A: User 4823 checked out for $142.50 successfully
B: {"event":"checkout","user_id":4823,"amount":142.50,"status":"ok"}

Options

Show answer
Line B — it is structured (JSON) with typed, named fields, so the platform indexes user_id and amount and you can query amount > 100 directly
Why:

Line B is structured logging: a machine-readable object with named, typed fields. A log platform parses it into indexed fields (user_id as a number, amount as a number), so a query like event:checkout AND amount > 100 is exact and survives wording changes. Line A is a human sentence — to extract the amount you'd need a fragile regex that breaks the moment someone changes 'checked out for $' to 'purchased', drops the dollar sign, or localizes the message, and the value arrives as text, not a comparable number. That brittleness is exactly why production systems standardize on structured logs with consistent field names (and a correlation/trace id) across services. The claim that free-text logs are easier for machines to query has it backwards; prose is easy for humans, hard for machines. The claim that structure makes no difference is wrong because structure is precisely what makes reliable field queries possible. The claim that numeric thresholds can never be evaluated from logs is false — you absolutely can threshold over a parsed numeric log field; metrics are often derived from such logs, but the log field itself is queryable.

Which statements about OpenTelemetry are correct? Select all that apply.#

Options

Pick every one that applies.

Show answer

OpenTelemetry is a vendor-neutral instrumentation layer covering traces, metrics, and logs through one API/SDK, so code is instrumented once regardless of destination. Context propagation — serializing an active span's trace ID and span ID into outgoing headers, such as the W3C Trace Context 'traceparent' header, and reading it back downstream — is what lets independently-generated spans from separate services assemble into a single trace. Because instrumentation is decoupled from the backend, switching observability vendors is normally a config change, not a rewrite. OpenTelemetry does not itself store or visualize data; that still requires a backend.

Why:

OpenTelemetry (a CNCF project) provides one vendor-neutral API/SDK surface across traces, metrics, and logs, so instrumentation is written once against the OpenTelemetry API rather than against a specific vendor's proprietary client. The mechanism that turns independently-created spans on separate services into a single coherent trace is context propagation: the active span's context is serialized into a standard header on outgoing calls and extracted on the receiving side to seed a child span with the correct parent — this is exactly what the W3C Trace Context standard formalizes, and it's the core concept enabling distributed tracing at all. Because the instrumentation API is decoupled from the destination, switching backends is normally an exporter/collector configuration change, not an application rewrite — this is the practical payoff of vendor neutrality and the main reason teams adopt it. The claim that OpenTelemetry itself replaces the need for an observability backend is false: OpenTelemetry is the instrumentation, collection, and export layer; it deliberately does not include a storage/visualization backend, so you still need something like Prometheus, Jaeger, Tempo, or a commercial APM to store and view the data. The claim that OpenTelemetry only covers distributed tracing is false: traces, metrics, and logs are all first-class OpenTelemetry signals, not an out-of-scope extra.

How do an SLI, an SLO, and an error budget relate to each other?#

Show answer

An SLI (service level indicator) is the actual quantitative measurement of some aspect of service behavior from the user's perspective — for example, 'percentage of requests that returned a non-5xx status over the last 28 days,' or a latency percentile. An SLO (service level objective) is a target set on top of that SLI — for example, '99.9% of requests succeed.' The error budget is what's left over: 100% minus the SLO, so a 99.9% SLO leaves a 0.1% error budget over the measurement window. Concretely, if the service handles 1,000,000 requests over 28 days, a 99.9% SLO permits roughly 1,000 failed requests before the budget is exhausted. The three form a stack, in order: you measure with the SLI, you set a target on that measurement with the SLO, and the error budget is the operational slack derived from that target — spent on deploy risk and experimentation while it remains, and its exhaustion is what triggers a reliability-first policy (freeze risky releases, redirect effort to fixes) until the service earns budget back.

Why:

This is foundational SRE vocabulary, and interviewers listen for the stack order rather than just the definitions in isolation: measure (SLI) → set a target on the measurement (SLO) → derive operational slack from the target (error budget). A strong answer also grounds it with a concrete numeric example (a 99.9% SLO on 1,000,000 requests over 28 days permits ~1,000 failures) and names the operational consequence of the budget — it isn't just an accounting artifact, it's the number that decides whether the team ships fast or slows down to firefight. Candidates who define all three correctly but can't say how spending/exhausting the budget changes team behavior are missing the practical half of the concept.

Order the steps of how a single distributed trace comes together, from a request first entering your system to an engineer viewing the assembled trace.#

Put these in order

Show answer

A distributed trace forms in five steps. First, a request with no existing trace context triggers the entry service to mint a new trace ID and open a root span. Before calling downstream, that span's context is injected into the outgoing request's headers (the W3C 'traceparent' header). The downstream service extracts that context and opens a child span recording its caller as parent — repeating at every hop. Each service exports its finished spans independently to a collector. Only then does the backend group every span sharing the trace ID and assemble them by parent/child order into one viewable trace.

Why:

This is the context-propagation mechanism that OpenTelemetry describes as the core concept enabling distributed tracing at all. It starts when a request arrives with no existing trace context, so the first service acts as the entry point: it mints a new trace ID and opens the root span. Before that service calls anything downstream, it serializes its current span's identifiers into the outgoing request — standardized today as the W3C Trace Context 'traceparent' header — so the propagation format is consistent across languages and vendors. The receiving service extracts that header, and instead of starting an unrelated new trace, it opens a child span that records the caller's span ID as its parent, which is what stitches the two services' work into one causal chain (this inject/extract pair repeats at every hop in a multi-service call). Each service exports its own finished spans, independently and typically batched, to a collector or backend — export doesn't wait for the whole request to finish. Only at the backend does assembly happen: every span sharing the trace ID gets grouped and ordered by its parent/child links into the waterfall view engineers actually read to see which hop dominated the latency. Getting inject/extract in the wrong order relative to the call, or forgetting that assembly is a backend-side step rather than something each service does locally, are the common misconceptions this ordering surfaces.

An SRE team runs a service with a 99.9% availability SLO. Which statements about the resulting error budget are correct? Select all that apply.#

Options

Pick every one that applies.

Show answer

The error budget is the allowed unreliability — 100% minus the SLO, so a 99.9% target permits 0.1% failure over the window. While budget remains, the team can spend it on shipping features and deploy risk; once it is exhausted, the policy freezes risky releases and prioritizes reliability work. Its value is turning the velocity-versus-stability debate into a shared, data-driven number. Spending zero budget is not the goal — it signals the SLO is too loose or reliability is over-invested.

Why:

An error budget is derived directly from the SLO: it is 100% minus the objective, so a 99.9% target permits 0.1% unreliability over the measurement window. That budget is a currency. While it is unspent the team has room to move fast — deploy often, take calculated risk — and when it runs out, the agreed error-budget policy kicks in: stop shipping risky changes and redirect effort to reliability until the service earns budget back. Its real organizational value is making the eternal velocity-vs-stability tension an objective, shared number instead of a turf war between developers who want to ship and ops who want stability. The claim that the goal is always to keep the error budget fully unspent is the classic trap: a budget that is never spent means your SLO is too loose (or you're over-investing in reliability users don't perceive) — you're leaving velocity on the table, so consistently spending zero is a signal to recalibrate, not a victory. The claim that the error budget is set by counting the number of incidents per quarter, independent of any SLO is wrong: the budget is computed from the SLO/SLI math, not by tallying incident counts.

If your application uses an OpenTelemetry tracing SDK, a bare console.log/print statement inside a traced function will automatically include that span's trace ID and span ID, with no logging-specific configuration, simply because the trace and the log statement run in the same process.#

Options

Show answer

False. A plain console.log/print call bypasses OpenTelemetry's logging pipeline entirely, so it will not include the active span's trace_id or span_id just because tracing and logging happen to run in the same process. Getting trace_id/span_id into log records requires OpenTelemetry's logging signal or bridge to read the active span context and inject those fields — sometimes a one-line configuration step, sometimes automatic once that integration is wired up, but never a byproduct of shared process memory alone.

Why:

False — correlating a log line with the active trace requires the OpenTelemetry logging signal or bridge to read the active span's context and inject trace_id/span_id into the emitted record; it doesn't happen just because tracing and logging code run in the same process. In practice this is often lightweight — the .NET SDK enables logs-to-activity correlation with no extra user code once OpenTelemetry logging is wired up, and Python needs an environment variable (OTEL_PYTHON_LOG_CORRELATION=true) plus adding %(otelTraceID)s/%(otelSpanID)s to the log formatter — but it is always a deliberate integration step. A bare console.log/print call bypasses that pipeline entirely and emits an ordinary line with no trace fields, regardless of which span happens to be active at that moment. This is exactly why 'log-trace correlation' is documented as its own setup step by every OpenTelemetry language SDK and observability vendor, rather than being an automatic side effect of co-located code.

What is 'cardinality' in the context of a metrics system like Prometheus, and why does adding a high-cardinality label — such as user_id, request_id, or a raw unnormalized URL path — to a metric cause production problems at scale?#

Show answer

Cardinality is the number of unique time series a metric produces, which equals the number of distinct combinations of its label values. A metric name plus one specific set of label key/value pairs identifies exactly one time series in the underlying store, and every new unique combination of values creates a brand-new series that has to be indexed and held in memory (and eventually on disk). Bounded labels like status_code or http_method take only a handful of values, so they multiply the series count by a small, known factor. Labels like user_id, session_id, request_id, or a raw URL path are effectively unbounded — a new value for every user, request, or unique path — so attaching one of them to a metric multiplies the series count by that unbounded factor, sometimes into the millions. In production this shows up as memory exhaustion in the time-series database (each series carries a real per-series memory floor), storage blowup, and slower queries because the engine has to scan far more series. Since teams typically respond to hitting that ceiling by dropping metrics or shortening retention, the fix is to keep unbounded-value fields out of metric labels entirely — aggregate them away, or route that data to logs or traces, which are built for high-cardinality point lookups, and reserve metrics for bounded, cheaply-aggregatable dimensions.

Why:

This question separates candidates who've only used a metrics dashboard from those who understand how a time-series database actually stores data. The interview point is the mechanism: a metric's identity is its name plus its full label set, so cardinality is combinatorial — each additional label multiplies the series count by the number of distinct values that label can take, not adds to it. A strong answer names the mechanism (unique label-value combination = new series = new memory/storage cost), gives the canonical bad examples (user_id, request_id, session_id, raw/unnormalized URLs), contrasts them with safe bounded labels (status_code, method, region), and states the production consequence (memory exhaustion, storage growth, slow queries, and the operational trap of teams dropping retention right when they need the history most). The fix — keep unbounded fields out of metric labels and use logs/traces for per-entity lookups — shows the candidate understands why the three observability signals are architecturally different tools, not interchangeable ones.

This Express middleware instruments every HTTP request with a Prometheus counter so the team can graph request volume and error rate per endpoint. After deploying, the Prometheus server's memory usage climbs without bound and eventually OOMs. What's the bug, and what's the fix?#

const httpRequestsTotal = new prom.Counter({
  name: "http_requests_total",
  help: "Total HTTP requests handled",
  labelNames: ["method", "path", "status_code"],
});

app.use((req, res, next) => {
  res.on("finish", () => {
    httpRequestsTotal.inc({
      method: req.method,
      path: req.originalUrl,
      status_code: res.statusCode,
    });
  });
  next();
});

Options

Show answer

path is set to req.originalUrl — the raw request path plus query string, including any embedded IDs like /users/48213/orders/91027?ref=email — so every distinct URL becomes a new label value and the counter grows one time series per unique URL ever hit, an unbounded set; the fix is to label with the matched route template (e.g. req.route.path, which yields /users/:id/orders/:id) and drop the query string, giving the label a small, bounded set of values regardless of traffic volume

Why:

The bug is cardinality explosion, and it's caused by the path label. req.originalUrl is the raw, unnormalized URL — it includes path parameters (/users/48213/...) and the query string, so nearly every request produces a URL that has never been seen before. Because a Prometheus time series is identified by metric name plus its full label set, every one of those distinct path values creates a brand-new, permanent time series that has to be indexed and held in memory — so memory grows roughly linearly with traffic instead of being bounded by the small number of actual endpoints the service exposes. The fix is to label with the matched route template the router resolved the request to (Express exposes this as req.route.path, giving a bounded value like /users/:id/orders/:id regardless of which user or order is requested) and to never include the query string in a label at all. The general rule this illustrates: any field whose value-space scales with users, requests, or time — IDs, session tokens, raw URLs, full timestamps — does not belong in a metric label; that kind of high-cardinality lookup is exactly what logs and traces are architected for, while a metrics store is architected for bounded aggregation. The claim that switching to res.on('close', ...) fixes the unbounded memory growth is a red herring — the finish-vs-close event timing has nothing to do with unbounded label growth. The claim that Prometheus counters leak memory by design and that calling .reset() on the counter on a timer fixes it misdiagnoses the failure as a client-library leak and proposes periodically discarding real data instead of fixing the label; it wouldn't even work, since new unique paths would keep arriving between resets and the underlying cause is untouched. Moving new prom.Counter(...) inside the app.use callback makes it worse: creating a new Counter object per request breaks the entire point of a counter (a single accumulating series) and does nothing about the unbounded path values driving the memory growth.

Related interview questions

Job market

See devops salaries and hiring demand from live job postings.

The other 8 questions

This page shows 10 and marks what you pick. That's as far as a page can go. A free account opens the other 8 and keeps every answer. What you miss comes back until it's right: after a day, then at longer gaps.

Start with this topic

Free · the whole bank · 100 marked answers per 30 days · written feedback on the paid plan

Interview booked? Practise the job ad instead

What moved, monthly

One email a month when the bulletin comes out: what moved in the markets we track, and the new question topics we published. Confirm your address to join. Unsubscribe any time.