DevOps & SRE Interview: 15 In-Depth Questions

Covers Kubernetes internals, SLI/SLO error budgets, automated rollouts, disaster recovery, and observability.

How AI interview works
15 real questions·3 categories·Interviewer follow-up logic per question

Questions reflect common real-world prompts. The three answer layers are illustrative examples, not real interview transcripts.

15 questionsClick a question to expand the 3 layers

① Common plain answer

"I check error logs using kubectl logs, run kubectl describe to inspect the exit code, or exec into the container."

Listing basic commands misses underlying failure chains such as InitContainer deadlocks, OOMKilled events, configuration mount failures, and misconfigured health probes.

② Interviewer follow-up logic

When zombie processes exhaust host PID limits and degrade node capacity, how do you configure pause containers or shared PID namespaces?If Pods experience OOMKilled events while JVM heap metrics remain well below thresholds, what diagnostic techniques isolate off-heap memory leaks?In automated deployment pipelines, how do PreStop hooks and terminationGracePeriodSeconds parameters ensure zero-downtime connection draining?

③ Quantified high-score answer

Diagnosing Kubernetes pods stuck in CrashLoopBackOff requires a deterministic triage sequence isolating Linux container termination primitives from orchestrator lifecycle constraints. Running kubectl describe pod immediately inspects the container Last State structure to extract the termination reason and exit code. Exit code 137 signifies SIGKILL dispatched by the kernel OOM Killer when exceeding container cgroup memory limits, requiring heap analysis and limit recalibration. Exit codes 1 or 2 denote runtime initialization panics, extracted using kubectl logs --previous before container recreation. In our production payments cluster hosting 420 microservices, an aggressive Liveness probe configured with an initialDelaySeconds of 5 seconds repeatedly killed heavy JVM workloads taking 18 seconds to warm database connection pools, inducing a cascading CrashLoop outage. By deploying Kubernetes startupProbes with failureThreshold failure buffers to decouple initialization latency from ongoing health evaluations, we eliminated 99.2% of false-positive restarts while keeping steady-state liveness probe thresholds tight at 3-second failure intervals.

Finished the breakdown? Try a realistic mock interview

Start a round without signing up. Experience in-depth follow-up questions and surface your real project highlights.

Create free account

No credit card required · Free 600 credits on signup