LLM Inference & Infra Interview: 15 In-Depth Questions
Covers KV cache fragmentation, PagedAttention, tensor parallelism, speculative decoding, and low-bit quantization.
Questions reflect common real-world prompts. The three answer layers are illustrative examples, not real interview transcripts.
① Common plain answer
"The model must cache intermediate key-value tensors for every token, and vLLM manages memory better to reduce wastage."
Surface tool familiarity misses underlying physical mechanics: internal fragmentation from pre-allocations, external fragmentation, and virtual paging.
② Interviewer follow-up logic
③ Quantified high-score answer
In autoregressive large language model inference, sequential KV cache expansion makes contiguous tensor allocation unsustainable, wasting up to 75% of GPU High Bandwidth Memory (HBM) through severe internal and external memory fragmentation. PagedAttention addresses this physical bottleneck by translating operating system virtual memory paging into accelerator architectures. Instead of pre-reserving contiguous physical memory buffers for maximum context windows, it decomposes the KV cache into fixed-size physical blocks (typically 16 or 32 tokens) linked through an indexed logical block table, allocating physical pages non-contiguously on-demand. In our production cluster serving 70B parameter models under 1,600 concurrent client streams, deploying PagedAttention collapsed KV cache memory waste from 72% to under 4%, facilitating a 3.2x increase in effective batch throughput on identical 8xH100 GPU server topologies. Furthermore, incorporating reference-counted Copy-on-Write semantics enables multiple parallel decoding branches and multi-turn system prompt prefixes to share identical physical blocks without memory duplication until output tokens diverge.
① Common plain answer
"If first token latency is slow, I add GPUs; if decoding is slow, I decrease batch size or compress model weights."
Fails to differentiate physical compute bounds: Prefill compute-bound matrix multiplications versus Decode memory-bound bandwidth constraints.
② Interviewer follow-up logic
③ Quantified high-score answer
TTFT and TPOT represent opposing hardware constraints. The Prefill phase computes all prompt tokens concurrently, making it compute-bound: optimization requires FlashAttention kernels maximizing Tensor Core occupancy and Chunked Prefill to interleave long prompts into iteration batches. Conversely, the Decode phase generates single tokens autoregressively, making it memory-bandwidth bound (constrained by HBM throughput). Optimization focuses on enlarging batch sizes to amortize weight transfers, running speculative decoding, or deploying FP8 quantization to reduce memory traffic.
① Common plain answer
"FlashAttention is a fused CUDA kernel for attention calculations that makes execution several times faster on GPUs."
Ignores physical memory bandwidth limits, lacking understanding of matrix tiling, online softmax normalization, and avoiding roundtrips to high-latency HBM.
② Interviewer follow-up logic
③ Quantified high-score answer
Standard Attention materializes the complete N-by-N attention matrix in high-latency HBM, causing severe memory bandwidth bottlenecks. FlashAttention introduces tiling and online softmax: Q, K, and V matrices are loaded in small tiles into high-speed on-chip SRAM (delivering 10TB/s+ bandwidth). By maintaining running maximums and normalization factors mathematically, softmax computes incrementally without writing intermediate attention matrices back to HBM. FlashAttention-2 and 3 further optimize warp scheduling and leverage Hopper TMA asynchronous memory transfers.
① Common plain answer
"A smaller model drafts several words quickly, then the larger model checks them, accepting valid tokens or regenerating mistakes."
Lacks rigorous understanding of tree attention masking, modified rejection sampling mathematics, and probabilistic alignment with target models.
② Interviewer follow-up logic
③ Quantified high-score answer
Speculative decoding trades inexpensive draft compute for parallelized target verification. The serving engine deploys an efficient draft model alongside a target model: the draft model generates K candidate tokens rapidly. The target model evaluates the context and all K candidates in a single forward pass utilizing tree attention masks. Applying rejection sampling based on probability ratios confirms M valid tokens per step. Because verification matches the latency of single-token decoding, maintaining acceptance rates above 70% accelerates throughput by two-to-threefold.
① Common plain answer
"I convert 16-bit floating point weights into 8-bit or 4-bit integers to halve VRAM and accelerate speeds without losing accuracy."
Overlooks activation outliers that destroy quantization precision, missing the distinction between weight-only (W4A16) and full INT8 Tensor Core execution.
② Interviewer follow-up logic
③ Quantified high-score answer
Quantization degradation stems from high-magnitude activation outlier channels. GPTQ executes second-order error compensation via inverse Hessian matrices, updating unquantized weights row-by-row for W4A16 weight-only storage. AWQ observes that weights are not equally salient: protecting the top 1% of salient weights based on activation magnitude retains FP16 precision, quantizing only residual weights. SmoothQuant enables true W8A8 matrix multiplications by mathematically migrating outlier scaling factors from activations into weights, unlocking hardware INT8 Tensor Core acceleration.
① Common plain answer
"If a model exceeds single-GPU VRAM, we partition layers across cards or split matrix multiplications in half to compute concurrently."
Lacks understanding of Megatron-LM row/column partitioning, All-Reduce synchronization points, and pipeline bubble minimization.
② Interviewer follow-up logic
③ Quantified high-score answer
TP and PP address distinct communication and memory boundaries. Tensor Parallelism partitions matrix multiplications within individual layers: in MLP blocks, first-layer weights are split column-wise and second-layer weights row-wise, requiring a single All-Reduce synchronization per layer. Because communication occurs per token, TP demands ultra-high-speed NVLink interconnects within single server chassis. Pipeline Parallelism splits layers across GPUs, introducing idle bubbles; engines mitigate bubbles during inference using micro-batching. Standard topology applies TP=8 intra-node, scaling PP or Data Parallelism inter-node.
① Common plain answer
"We assign high-compute servers to process incoming prompts, and lower-tier servers to generate output text, sending data over networks."
Ignores how mixed workloads cause compute-bound prefill requests to induce catastrophic jitter on memory-bound decode latency, lacking RDMA KV transfer designs.
② Interviewer follow-up logic
③ Quantified high-score answer
Coupling prefill and decode phases inside identical nodes creates severe resource contention: compute-heavy prompt processing stalls ongoing autoregressive streams. PD Disaggregation decouples clusters into specialized node pools: Prefill pools run large-batch prompt processing at near-100% compute saturation. Upon completion, intermediate KV Caches transfer zero-copy over high-speed RDMA/RoCEv2 networks directly into Decode pools. Decode nodes run isolated autoregressive steps with stable TPOT. This disaggregation doubles aggregate cluster throughput while reducing P99 latencies by 60%.
① Common plain answer
"We cache embeddings of common system prompts, so when user queries share identical prefixes, the engine skips redundant computation."
Treats prefix caching as static hash lookups, failing to explain Radix-Tree block management, reference counting, and cache-aware routing.
② Interviewer follow-up logic
③ Quantified high-score answer
Prefix caching eliminates redundant prefill compute across agent workflows and few-shot prompts. In engines like SGLang and vLLM, a Radix Tree indexes historical token sequences against physical KV Cache blocks. When requests arrive, the engine executes longest-prefix matching: matching prefix blocks reuse existing VRAM allocations directly, computing prefill solely for novel suffix tokens. Pair this with Cache-Aware Routing at the gateway layer, dispatching requests sharing system prompts to identical physical GPU instances, slashing cluster prefill compute costs by over 70%.
① Common plain answer
"We deploy Kubernetes clusters, install the NVIDIA GPU Operator, and configure deployment manifests requesting eight GPU devices per pod."
Standard Kubernetes schedulers lack NVLink topology awareness, frequently allocating GPUs across NUMA nodes and halving tensor-parallel bandwidth.
② Interviewer follow-up logic
③ Quantified high-score answer
Distributed LLM serving demands hardware-aware scheduling. We deploy Kubernetes schedulers (Volcano or Kueue) with Node Feature Discovery: scheduling pods requiring Tensor Parallelism exclusively across GPUs co-located on identical NUMA nodes interconnected via full-mesh NVLink, eliminating PCIe cross-switch degradation. For autoscaling, KEDA monitors real-time request queue depths and KV Cache saturation. To accelerate pod cold starts, we implement peer-to-peer container image distribution and parallel safetensor direct memory mapping, compressing spin-up durations from ten minutes to under 40 seconds.
① Common plain answer
"We increase max_tokens configurations in model settings, install additional server memory, and run larger H100 GPUs with expanded VRAM."
Ignores quadratic attention complexity, lacking understanding of RingAttention sequence parallelism, sparse attention sinks, and RoPE extrapolation.
② Interviewer follow-up logic
③ Quantified high-score answer
Conquering 128K+ contexts requires combining architectural sparsification with sequence parallelism. Algorithmically, we deploy sparse attention mechanisms (such as StreamingLLM retaining initial attention sinks and sliding local windows) to break quadratic complexity barriers. Infrastructure-wise, we implement Sequence Parallelism via RingAttention or DeepSpeed-Ulysses: partitioning input token sequences along the sequence dimension across ring-connected GPUs. Each card computes localized attention blocks while asynchronously passing KV tensors in a ring, fitting 1M token contexts across multi-GPU clusters alongside NTK-aware RoPE position scaling.
① Common plain answer
"We inspect GPU utilization and VRAM occupancy on Grafana, and collect API request throughput and average latency using Prometheus."
Traditional server metrics miss LLM-specific indicators: Time to First Token (TTFT), Time Per Output Token (TPOT), and KV Cache occupancy.
② Interviewer follow-up logic
③ Quantified high-score answer
LLM APM requires capturing specialized generation telemetry across three layers: At the hardware layer, we record GPU power draw, SM utilization, memory bandwidth, and NVLink saturation. At the inference engine tier, we expose real-time runtime metrics: KV Cache allocation percentages, waiting queue dwell times, TTFT distributions, TPOT percentiles (P50/P90/P99), and prefill preemption counts. At the API gateway tier, we map token counts to financial cost metrics, enabling immediate triage to isolate whether performance degradation stems from compute saturation, memory bandwidth limits, or network transport latency.
① Common plain answer
"I restart crashed Docker container pods immediately, request business teams throttle API traffic in group chats, and bring nodes back online."
Blindly restarting pods under load causes immediate secondary crashes, lacking gateway backpressure, queue draining, and graduated traffic restoration.
② Interviewer follow-up logic
③ Quantified high-score answer
Mitigating systemic GPU OOM crashes requires immediate load shedding followed by phased recovery. Step one is perimeter containment: the ingress gateway trips circuit breakers, returning HTTP 429 backpressure to incoming requests and shedding traffic to protect surviving nodes. Step two stabilizes the serving engine: schedulers reject new prefill jobs while completing in-flight decode generations, preventing work-in-progress state loss. Step three identifies root causes: tracing payloads reveals whether malicious prompt bombs bypassed input validation limits. After tightening gateway input guards, traffic is gradually restored.
① Common plain answer
"I explain how expensive commercial APIs are, emphasize data sovereignty requirements, and tell business teams they are obligated to use internal models."
Relying on security mandates without addressing user experience creates resentment, lacking empirical benchmarking and dedicated performance tuning.
② Interviewer follow-up logic
③ Quantified high-score answer
Resolving stakeholder dissatisfaction demands empirical transparency and targeted performance engineering. Rather than arguing policy, I benchmark real production traffic: running automated scripts evaluating proprietary serving against commercial APIs across TTFT, TPOT, and uptime to identify concrete bottlenecks. I launched a one-week optimization sprint: upgrading to modern vLLM engines, enabling Chunked Prefill, and deploying FP8 quantization, cutting P90 latency by 50% to surpass cloud baselines. Delivering an analysis highlighting data privacy, custom fine-tuning, and 80% lower cost per token restored stakeholder trust.
① Common plain answer
"I instruct researchers to rewrite custom operators using standard PyTorch primitives, informing them that non-standard layers cannot be deployed."
Dismissing novel research stalls algorithmic innovation, lacking the engineering ownership to analyze bottlenecks, author custom Triton kernels, and benchmark gains.
② Interviewer follow-up logic
③ Quantified high-score answer
AI infrastructure bridges research innovation with production serving velocity. When evaluating non-standard operators, I partner with research scientists to understand the architectural intent: quantifying the algorithmic accuracy gains. If the gains are substantial, I author custom Triton or CUDA kernels to implement the operator natively, compiling them into TensorRT-LLM or vLLM custom plugin registries. We run precision regression suites to verify numerical accuracy within 1e-3 tolerance while benchmarking throughput, preserving algorithmic breakthroughs without compromising production serving SLAs.
① Common plain answer
"We tell leadership that we saved many GPUs, writing technical progress updates into weekly reports to show our team works hard."
Technical jargon fails to convince finance and executive leadership, lacking commercial financial translation like cost per million tokens and CAPEX savings.
② Interviewer follow-up logic
③ Quantified high-score answer
Communicating with executive leadership requires translating technical acceleration into commercial financial metrics. I instituted our corporate LLM Cost and Efficiency Compass, tracking cost per million tokens alongside throughput capacity per GPU card. During quarterly reviews, I demonstrated that by deploying speculative decoding, PagedAttention, and dynamic prefix caching, our infrastructure absorbed a 3x surge in corporate LLM usage without procuring additional GPU hardware. This avoided millions in cloud leasing and capital expenditures, proving infrastructure engineering as a strategic profit-enabler.
Keep practicing in another role
After LLM Inference & Infra Engineer, these are the adjacent roles to practice next
AI Agent & LLM Interview: 15 In-Depth Questions
Workflows · Tool Calling · Retrieval · Evaluation · BQ
View bank
Same tech stackAlgorithm & ML Interview: 15 In-Depth Questions
Machine Learning · Engineering · BQ
View bank
Common pivotBig Data Engineer Interview: 15 In-Depth Questions
Real-Time Data Warehouse · Stream Processing · ETL · BQ
View bank
Don't see your role? Browse all 25 roles →
Finished the breakdown? Try a realistic mock interview
Start a round without signing up. Experience in-depth follow-up questions and surface your real project highlights.
No credit card required · Free 600 credits on signup