LLM Inference & Infra Interview: 15 In-Depth Questions

Covers KV cache fragmentation, PagedAttention, tensor parallelism, speculative decoding, and low-bit quantization.

How AI interview works
15 real questions·3 categories·Interviewer follow-up logic per question

Questions reflect common real-world prompts. The three answer layers are illustrative examples, not real interview transcripts.

15 questionsClick a question to expand the 3 layers

① Common plain answer

"The model must cache intermediate key-value tensors for every token, and vLLM manages memory better to reduce wastage."

Surface tool familiarity misses underlying physical mechanics: internal fragmentation from pre-allocations, external fragmentation, and virtual paging.

② Interviewer follow-up logic

When executing parallel beam search or multi-candidate sampling, how does the block table leverage reference counting for Copy-on-Write?Under extreme context saturation, how do KV Cache CPU swapping and offloading mechanisms prevent pipeline stalls during high-concurrency bursts?How do Multi-Query Attention (MQA) and Grouped-Query Attention (GQA) reduce HBM memory bandwidth compared to standard Multi-Head Attention?

③ Quantified high-score answer

In autoregressive large language model inference, sequential KV cache expansion makes contiguous tensor allocation unsustainable, wasting up to 75% of GPU High Bandwidth Memory (HBM) through severe internal and external memory fragmentation. PagedAttention addresses this physical bottleneck by translating operating system virtual memory paging into accelerator architectures. Instead of pre-reserving contiguous physical memory buffers for maximum context windows, it decomposes the KV cache into fixed-size physical blocks (typically 16 or 32 tokens) linked through an indexed logical block table, allocating physical pages non-contiguously on-demand. In our production cluster serving 70B parameter models under 1,600 concurrent client streams, deploying PagedAttention collapsed KV cache memory waste from 72% to under 4%, facilitating a 3.2x increase in effective batch throughput on identical 8xH100 GPU server topologies. Furthermore, incorporating reference-counted Copy-on-Write semantics enables multiple parallel decoding branches and multi-turn system prompt prefixes to share identical physical blocks without memory duplication until output tokens diverge.

Finished the breakdown? Try a realistic mock interview

Start a round without signing up. Experience in-depth follow-up questions and surface your real project highlights.

Create free account

No credit card required · Free 600 credits on signup