DevOps & SRE Interview: 15 In-Depth Questions
Covers Kubernetes internals, SLI/SLO error budgets, automated rollouts, disaster recovery, and observability.
Questions reflect common real-world prompts. The three answer layers are illustrative examples, not real interview transcripts.
① Common plain answer
"I check error logs using kubectl logs, run kubectl describe to inspect the exit code, or exec into the container."
Listing basic commands misses underlying failure chains such as InitContainer deadlocks, OOMKilled events, configuration mount failures, and misconfigured health probes.
② Interviewer follow-up logic
③ Quantified high-score answer
Diagnosing Kubernetes pods stuck in CrashLoopBackOff requires a deterministic triage sequence isolating Linux container termination primitives from orchestrator lifecycle constraints. Running kubectl describe pod immediately inspects the container Last State structure to extract the termination reason and exit code. Exit code 137 signifies SIGKILL dispatched by the kernel OOM Killer when exceeding container cgroup memory limits, requiring heap analysis and limit recalibration. Exit codes 1 or 2 denote runtime initialization panics, extracted using kubectl logs --previous before container recreation. In our production payments cluster hosting 420 microservices, an aggressive Liveness probe configured with an initialDelaySeconds of 5 seconds repeatedly killed heavy JVM workloads taking 18 seconds to warm database connection pools, inducing a cascading CrashLoop outage. By deploying Kubernetes startupProbes with failureThreshold failure buffers to decouple initialization latency from ongoing health evaluations, we eliminated 99.2% of false-positive restarts while keeping steady-state liveness probe thresholds tight at 3-second failure intervals.
① Common plain answer
"We target four nines for server CPU utilization and API availability, monitor alert groups, and freeze deployments when metrics dip."
Confuses raw infrastructure capacity with real user satisfaction, lacking Critical User Journey SLI definitions and burn rate alerting policies.
② Interviewer follow-up logic
③ Quantified high-score answer
SRE rejects using raw CPU as an SLI. We define metrics around Critical User Journeys: for example, measuring payment checkout availability as the percentage of requests returning HTTP 2xx under 500ms. We set a 99.95% monthly SLO, yielding an error budget of 21.6 minutes. Our monitoring implements multi-window burn rate alerts, paging on-call engineers if 5% of the monthly budget burns in one hour. If monthly error budgets deplete past 80%, automated gates freeze non-critical product releases, redirecting engineering sprints toward stability hardening.
① Common plain answer
"I instrument all microservices with OpenTelemetry and export every trace to Elasticsearch or Jaeger for persistent querying."
Capturing 100% of telemetry traces bankrupts storage clusters, lacking two-tier sampling architectures, tail-based filtering, and columnar storage models.
② Interviewer follow-up logic
③ Quantified high-score answer
Collecting 100% of production traces is a recognized anti-pattern. We enforce a two-tier sampling architecture within our OpenTelemetry Collector cluster. Clients apply a 1% probabilistic head-based sample to establish throughput baselines. The collector cluster applies tail-based sampling, buffering spans in memory to force 100% persistence whenever requests contain HTTP 5xx responses, latencies exceeding 1.5 seconds, or high-priority tenant headers. Storing traces in ClickHouse columnar databases reduced infrastructure costs by 75% while preserving full forensic fidelity.
① Common plain answer
"During releases, I deploy to one server first, watch for runtime errors manually, and click the rollback button if issues appear."
Manual metric watching lacks statistical rigor, lacking automated statistical verification of canary cohorts against baseline production traffic.
② Interviewer follow-up logic
③ Quantified high-score answer
Canary releases require automated statistical verification. Using Argo Rollouts and Prometheus metrics, we route 5% of production traffic to canary pods alongside existing baseline instances. Over a ten-minute observation window, the controller evaluates key statistical indicators: 5xx error rate deltas, P99 latency regressions, and uncaught exception frequency. If metrics deviate beyond statistical variance thresholds, the controller halts traffic progression, executing automatic rollbacks within 30 seconds and broadcasting forensic incident telemetry directly to internal alerting channels.
① Common plain answer
"I write Terraform scripts to provision cloud servers and run Jenkins jobs to deploy application code directly to the Kubernetes cluster."
Lacks declarative reconciliation, Git-backed auditability, drift detection loops, and secrets management required for immutable infrastructure.
② Interviewer follow-up logic
③ Quantified high-score answer
IaC and GitOps enforce declarative desired-state management. We provision cloud networks, managed Kubernetes clusters, and storage buckets using modular Terraform with remote encrypted state locking. For application delivery, ArgoCD implements GitOps principles, establishing Git repositories as the immutable Single Source of Truth. ArgoCD continuously reconciles live cluster state against Git manifests, automatically correcting manual configuration drift. Dynamic secrets are injected securely at runtime using HashiCorp Vault and External Secrets operators.
① Common plain answer
"I restart the kubelet service on the host, inspect the system logs, or remove the node and re-join it to the cluster."
Blindly restarting services lacks systemic troubleshooting rigor, ignoring container runtime sockets, kernel hung tasks, disk pressure, and CNI routing failures.
② Interviewer follow-up logic
③ Quantified high-score answer
Investigating a NotReady node begins with running kubectl describe node to inspect the Conditions block, identifying DiskPressure, PIDPressure, or NetworkUnavailable flags. For disk pressure, we audit ephemeral storage and container logs under var lib containerd. If kubelet heartbeats fail, I review journalctl -u kubelet to identify containerd socket deadlocks or CNI routing failures across Calico pods. If hardware degradation is confirmed, we cordon and drain workloads immediately to maintain application availability while preserving kernel diagnostic traces.
① Common plain answer
"I add Redis caches when endpoints slow down, queue incoming requests on the frontend, or disable non-essential queries in the backend."
Static thresholds collapse under abrupt traffic spikes, lacking queue-delay load shedding algorithms (like CoDel) and circuit-breaking architectures.
② Interviewer follow-up logic
③ Quantified high-score answer
Load shedding must adapt dynamically to internal system saturation. Rather than relying on static rate limits that fail during sudden traffic surges, our edge gateways deploy adaptive queue-delay algorithms. We track inflight request queues and CPU saturation; when queue delays spike, low-priority background jobs and recommendation feeds are dropped first to safeguard core checkout transactions. Inter-service RPC calls configure circuit breakers via Envoy, failing fast to cached fallbacks when downstream error thresholds trip to prevent systemic collapse.
① Common plain answer
"We deploy data centers across two cities with bi-directional master-slave database replication, switching DNS records during failures."
Dual-master active-active replication over long distances guarantees write conflicts and data corruption, ignoring cell-based sharding principles.
② Interviewer follow-up logic
③ Quantified high-score answer
Active-active reliability requires cell-based routing and localized transactional boundaries. We shard workloads deterministically by customer ID hashes into independent geographic cells, ensuring 95% of write transactions finalize entirely within local region boundaries to eliminate cross-continent latency. Global configurations replicate via Raft consensus. During metropolitan disaster events, global traffic routing redistributes traffic away from impaired availability zones, pairing distributed locks with cell routing tables to eliminate split-brain write corruption.
① Common plain answer
"I deploy Prometheus to scrape all Kubernetes pod metrics and configure Alertmanager to broadcast alerts to email distribution lists."
Standalone Prometheus instances create single points of failure and storage limits, lacking long-term downsampling and alert suppression logic.
② Interviewer follow-up logic
③ Quantified high-score answer
We separate telemetry collection from long-term persistence. Lightweight Prometheus agents scrape metrics locally across clusters, streaming data via Remote Write to a centralized VictoriaMetrics or Thanos cluster for downsampling and multi-year retention. For alert governance, Alertmanager enforces multi-tiered deduplication: grouping alerts by service and environment tags, and applying inhibition rules so datacenter connectivity failures automatically suppress thousands of downstream microservice timeout alerts. This reduced daily on-call paging noise by over 80%.
① Common plain answer
"We connect data centers with dedicated private leased lines, configure host firewalls, and allow services to communicate via internal IPs."
Perimeter-based security permits unrestricted lateral movement once breached, lacking cryptographic workload identities, mutual TLS, and L7 policies.
② Interviewer follow-up logic
③ Quantified high-score answer
Zero Trust assumes the internal network is hostile. We deploy Istio and SPIRE across our hybrid Kubernetes clusters, issuing short-lived cryptographically verified X.509 certificates to each workload. The service mesh data plane enforces mutual TLS (mTLS) across all east-west traffic by default, encrypting cross-cluster network packets. Fine-grained AuthorizationPolicies restrict inter-service communication to strict allowlists, preventing lateral movement. Ingress access terminates static bastion passwords in favor of short-lived OIDC-backed JIT authorization tokens.
① Common plain answer
"We unplug network cables in test environments or kill random application pods to check if containers recover automatically."
Uncontrolled fault injection resembles destructive chaos, lacking steady-state hypotheses, blast radius containment, and automated kill switches.
② Interviewer follow-up logic
③ Quantified high-score answer
Chaos Engineering is controlled scientific experimentation. Every exercise adheres to disciplined lifecycle phases: First, we define quantifiable steady-state metrics, such as checkout success rates and P99 latency. Second, we formulate hypotheses: "Injecting 30% packet loss into risk analysis pods will trigger graceful circuit breaking without degrading primary checkout." Using Chaos Mesh, we inject targeted synthetic degradation strictly into canary traffic cohorts. Automated kill switches abort experiments within ten seconds if steady-state metrics breach safety thresholds.
① Common plain answer
"I tag all software engineers in our chat channels, tell everyone to inspect their components, and urge them to resolve issues quickly."
Lacks Incident Command System structure; unstructured chatter, absence of defined operational roles, and premature debugging prolong outage duration.
② Interviewer follow-up logic
③ Quantified high-score answer
As Incident Commander, my role is orchestrating structure, filtering operational noise, and accelerating service restoration. Upon opening the war room, I assign dedicated roles: an operational scribe to broadcast external status updates every 15 minutes, and lead technical investigators to evaluate telemetry without executive distraction. Incident triage strictly prioritizes mitigation over root-cause debugging: we restore traffic via rollback, canary isolation, or feature toggles rather than attempting live code patches in production. Forensic root-cause analysis begins only after customer impact reaches zero.
① Common plain answer
"We hold a retrospective meeting to identify who pushed the erroneous commit or misconfigured the server, and penalize responsible individuals."
Assigning individual blame encourages teams to conceal mistakes and fosters defensive silos, failing to fix underlying organizational vulnerabilities.
② Interviewer follow-up logic
③ Quantified high-score answer
Blameless culture operates on a core axiom: engineers make the best decisions possible given the context and tools available at the time. Post-mortems investigate why the system permitted dangerous actions and failed to contain the impact. We apply the 5 Whys technique: Why did automated CI gates fail to catch this regression? Why did alerts take ten minutes to page on-call engineers? We convert findings into prioritized Action Items with assigned owners and verifiable completion criteria, transforming outages into durable system hardening.
① Common plain answer
"I reject the release outright, cite company stability policies, and revoke deployment credentials until all verification tests pass."
Rigid gatekeeping creates adversarial friction, prompting business partners to view SRE as blockers rather than collaborative business enablers.
② Interviewer follow-up logic
③ Quantified high-score answer
SRE builds guarded superhighways rather than arbitrary roadblocks. When faced with high-stakes release deadlines, I collaborate with business leaders to understand commercial drivers and target launch windows. Instead of unilateral rejection, we craft an emergency mitigation framework: executing lightweight automated load tests on critical paths while disabling non-essential background modules. We pre-configure targeted traffic rate limiters at the gateway and agree on signed rollback thresholds, balancing velocity with baseline stability.
① Common plain answer
"I work overtime to build automated scripts when tasks pile up, delay manual support tickets, or hire interns to manage tickets."
Relying on manual heroics and ticket delays perpetuates burnout, lacking systematic toil measurement and adherence to the SRE 50% engineering rule.
② Interviewer follow-up logic
③ Quantified high-score answer
SRE tenets mandate capping repetitive operational toil at 50% of total engineering capacity. My elimination playbook follows three steps: First, we categorize manual tickets to quantify the top recurring operational requests, such as database credentials or log exports. Second, we apply our automation rule: any procedure performed manually three times must be automated into self-service workflows via our Internal Developer Platform. Third, we redirect reclaimed capacity toward high-leverage architectural projects, including autoscaling and chaos game days.
Keep practicing in another role
After DevOps & SRE Engineer, these are the adjacent roles to practice next
Backend Interview: 15 In-Depth Questions
Core Tech · Distributed Systems · BQ
View bank
Same tech stackFull-Stack Engineering Interview: 15 In-Depth Questions
End-to-End Delivery · Architecture · Full-Stack Perf · BQ
View bank
Skill extensionCybersecurity & Pen Testing Interview: 15 In-Depth Questions
Offensive & Defensive · Vulnerability Mining · Zero Trust · BQ
View bank
Don't see your role? Browse all 25 roles →
Finished the breakdown? Try a realistic mock interview
Start a round without signing up. Experience in-depth follow-up questions and surface your real project highlights.
No credit card required · Free 600 credits on signup