Transformers are one of the most important architectures in modern Natural Language Processing (NLP), Generative AI, and Large Language Models (LLMs). Introduced in the 2017 paper Attention Is All You Need, the Transformer architecture replaced the recurrent processing used by earlier sequence models with attention-based computation.
This guide explains the Transformer architecture from the basics to advanced concepts such as Query, Key, Value (QKV), positional encoding, scaled dot-product attention, multi-head attention, residual connections, Layer Normalization, Feed-Forward Networks, causal masking, cross-attention, KV cache, RoPE, GQA and FlashAttention.
Why Were Transformers Needed? #
Before Transformers, sequence-to-sequence systems commonly used RNNs, LSTMs and GRUs. These models process tokens sequentially, which creates a major limitation during training because later positions depend on earlier recurrent states.
Major Problems with RNNs and LSTMs #
- Sequential processing: Tokens are processed step by step, limiting parallelization during training.
- Long-range dependencies: Information and gradients may become difficult to preserve over long sequences.
- Context bottleneck: Early encoder-decoder systems compressed a source sequence into a fixed-size representation.
- Limited scalability: Sequential computation makes very large-scale pre-training more difficult.
How Does the Transformer Solve These Problems? #
The Transformer removes recurrence and uses attention mechanisms to allow tokens to interact with one another through parallel matrix operations. This makes training highly parallelizable and provides direct interactions between token positions.
| Property | RNN/LSTM | Transformer |
|---|---|---|
| Processing | Sequential | Highly parallel during training |
| Long-range interaction | Through recurrent steps | Direct attention between tokens |
| GPU parallelization | Limited by recurrence | Well suited to matrix operations |
| Scalability | More difficult | Highly scalable |
Transformer Architecture: Overview #
The original Transformer consists of an Encoder Stack and a Decoder Stack. The encoder creates representations of the source sequence, while the decoder uses the target context and encoder representations to generate output.
Token Embeddings #
A neural network cannot directly process words as text. Tokens are converted into numerical representations called embeddings.
An embedding maps a token ID into a dense vector representation that can be processed by neural network layers.
Static vs Contextual Representations #
| Type | Meaning |
|---|---|
| Static embedding | A token has a fixed representation independent of the sentence. |
| Contextual representation | The representation changes according to surrounding tokens. |
For example, the word bank can refer to a financial institution or the side of a river. Attention allows the representation of a token to incorporate information from surrounding context.
Positional Encoding #
Self-attention by itself does not inherently provide left-to-right sequence order. Therefore, Transformer models need a positional mechanism so that the model can distinguish different token positions.
For example, the sentences “Dog bites man” and “Man bites dog” contain the same words but have different meanings because the order is different.
Sinusoidal Positional Encoding #
The original Transformer uses sinusoidal positional encoding. The formulas are:
PE(pos, 2i) = sin(pos / 10000(2i / dmodel))
PE(pos, 2i+1) = cos(pos / 10000(2i / dmodel))
The positional vector is added element-wise to the token embedding:
Input Representation = Token Embedding + Positional Encoding
Adding positional information preserves the model dimension. Concatenating the two vectors would increase the representation dimension.
Memory Trick: Positional Encoding tells the Transformer where each token is located.
Query, Key and Value (Q, K, V) #
The central mechanism of Transformer attention uses three learned representations: Query (Q), Key (K), and Value (V).
| Vector | Simple Meaning |
|---|---|
| Query | What the current token is looking for. |
| Key | What a token contains or advertises for matching. |
| Value | The actual information that is passed forward. |
The input matrix X is projected into Q, K and V using learned matrices:
Q = XWQ
K = XWK
V = XWV
Real-Life Example #
Imagine searching YouTube. Your search phrase behaves like a Query. Video titles and tags behave like Keys. The actual video content behaves like the Value.
Scaled Dot-Product Attention #
Scaled dot-product attention is the mathematical core of Transformer attention.
Attention(Q, K, V) = Softmax((QKT) / √dk)V
Step-by-Step Attention Process #
- Generate Q, K and V.
- Calculate QKT to obtain compatibility scores.
- Divide the scores by √dk.
- Apply Softmax to obtain attention weights.
- Multiply the attention weights by V.
Why Do We Divide by √dk? #
As the dimensionality of Q and K increases, dot products can become large. Large logits can push Softmax toward very sharp probabilities and reduce useful gradient signal.
Dividing by √dk controls the scale of the dot products and helps keep Softmax in a more useful range during training.
Memory Trick: √dk acts as a Softmax scale or gradient stabilizer.
Worked Attention Example #
Consider two tokens: Money and Bank. Suppose the source guide gives the following values:
Vmoney = [1, 4]
Vbank = [3, 7]
After calculating QKT, scaling by √2 and applying Softmax, the example produces approximately:
| Token | Attention Weights |
|---|---|
| Money | [0.33, 0.67] |
| Bank | [0.11, 0.89] |
The resulting contextual representations are approximately:
Ymoney = [2.34, 6.01]
Ybank = [2.78, 6.67]
This demonstrates the central idea of attention: the representation of a token is updated by combining information from relevant tokens.
Exam Sequence: Q/K/V → Dot Product → Scale → Softmax → Weighted Sum.
Multi-Head Attention #
Instead of performing only one attention operation, Transformers use multiple attention heads. Each head can learn relationships in a different representation subspace.
The formula is:
MultiHead(Q,K,V) = Concat(head1, …, headh)WO
For example, with dmodel = 512 and 8 heads, each head can operate on a 64-dimensional representation.
Memory Trick: Multi-head attention is like using multiple cameras to view the same scene from different angles.
Residual Connections #
Transformer blocks use residual or skip connections to provide a direct path around a sub-layer.
Residual Output = x + SubLayer(x)
Residual connections help information and gradient signals flow through deep Transformer stacks.
Layer Normalization #
Layer Normalization normalizes the feature dimensions of an individual token representation. It is different from Batch Normalization, which uses statistics across batch samples.
LayerNorm also contains learned parameters such as gamma (γ) for scale and beta (β) for shift.
Position-wise Feed-Forward Network #
Each Transformer block contains a Feed-Forward Network (FFN), which applies a non-linear transformation independently to each token position.
FFN(x) = ReLU(xW1 + b1)W2 + b2
In the original-style configuration, the representation can expand from 512 → 2048 → 512.
Causal Masking #
Causal masking is used to prevent a decoder token from looking at future target tokens.
For example, when predicting token 2, the model must not use token 3 because token 3 is the future answer.
A conceptual three-token causal mask is:
[ 0 -∞ -∞ ] [ 0 0 -∞ ] [ 0 0 0 ]
The mask is added to the attention scores before Softmax. Values represented by −∞ receive approximately zero probability after Softmax.
Memory Trick: Causal Mask = Do not look into the future.
Cross-Attention #
Cross-attention connects the Decoder with the Encoder output in an encoder-decoder Transformer.
| Component | Source |
|---|---|
| Query | Decoder |
| Key | Encoder output |
| Value | Encoder output |
Therefore:
Q = XdecoderWQ
K = YencoderWK
V = YencoderWV
For translation, the decoder can use cross-attention to identify which source tokens are relevant to the target token currently being generated.
Memory Trick: Cross-Attention = Q from Decoder, K/V from Encoder.
Encoder Architecture #
The encoder repeatedly applies self-attention and feed-forward processing with residual connections and normalization.
Decoder Architecture #
The Decoder has an additional masked self-attention stage and, in an encoder-decoder Transformer, a cross-attention stage that receives Key and Value representations from the Encoder.
Training vs Inference #
| Aspect | Training | Inference |
|---|---|---|
| Target sequence | Ground-truth target available | Generated dynamically |
| Processing | Can be parallelized with causal masking | Autoregressive token-by-token generation |
| Teacher forcing | Used | Not used in the same way |
| Causal structure | Prevents future-token leakage | Maintained during generation |
Teacher Forcing #
Teacher forcing means feeding the ground-truth previous target tokens to the decoder during training rather than relying on previously generated predictions.
Autoregressive Generation #
During inference, the model predicts a token, appends that token to the generated sequence, and then predicts the next token. This continues until an end condition is reached.
Linear Layer and Softmax #
At the decoder output, a Linear layer maps the hidden representation to vocabulary-sized logits.
Hidden representation → Linear Layer → Vocabulary logits → Softmax → Token probabilities
Softmax converts the logits into probabilities over possible next tokens.
Transformer Computational Complexity #
Full self-attention has approximately:
O(n² · d)
where n is the sequence length and d is the representation dimension.
The quadratic term appears because QKT produces an n × n attention matrix.
| Sequence Length | Pairwise Attention Entries |
|---|---|
| 100 | 10,000 |
| 1,000 | 1,000,000 |
| 10,000 | 100,000,000 |
This quadratic behavior becomes an important challenge for very long context windows.
BERT vs GPT vs T5 #
| Model | Architecture | Main Idea |
|---|---|---|
| BERT | Encoder-only | Bidirectional representation learning; commonly used for understanding tasks. |
| GPT-style models | Decoder-only | Causal autoregressive language generation. |
| T5 | Encoder-decoder | Sequence-to-sequence transformation. |
Memory Trick: BERT = Understand, GPT = Generate, T5 = Transform.
Modern Transformer Concepts #
KV Cache #
KV Cache stores previously calculated Key and Value representations during autoregressive generation so that historical K/V states do not need to be recomputed repeatedly.
Rotary Position Embedding (RoPE) #
RoPE is a positional technique that applies position-dependent rotations to Query and Key representations.
Grouped-Query Attention (GQA) #
GQA allows multiple Query heads to share Key/Value heads, reducing the Key/Value cache memory requirement during inference.
FlashAttention #
FlashAttention is an exact, IO-aware attention implementation designed to improve GPU memory access and efficiency without approximating the mathematical attention operation.
Common Transformer Misconceptions #
| Misconception | Correct Understanding |
|---|---|
| Self-attention automatically knows word order. | Positional information is required to represent sequence order. |
| Multi-head attention simply multiplies parameters by the number of heads. | The model dimension is divided among heads and the outputs are combined. |
| LayerNorm is the same as BatchNorm. | LayerNorm normalizes feature dimensions for each token representation. |
| Causal masking allows future tokens. | Causal masking blocks future positions. |
| Self-attention and cross-attention are the same. | Self-attention uses one sequence; cross-attention connects different streams such as Decoder Q with Encoder K/V. |
Important Transformer Formulas #
| Concept | Formula |
|---|---|
| Query | Q = XWQ |
| Key | K = XWK |
| Value | V = XWV |
| Attention | Softmax((QKT)/√dk)V |
| Residual | x + SubLayer(x) |
| FFN | ReLU(xW1 + b1)W2 + b2 |
| Self-Attention Complexity | O(n² · d) |
Important Points for Exams and Interviews #
- Transformers were introduced in 2017 in Attention Is All You Need.
- Transformers replace recurrent processing with attention-based computation.
- Positional information is necessary because self-attention alone does not provide sequence order.
- Q, K and V are learned projections of token representations.
- Attention follows the sequence: Q/K/V → QKT → scaling → Softmax → weighted V.
- Multi-head attention captures information from multiple representation subspaces.
- Residual connections provide a direct path around sub-layers.
- LayerNorm normalizes feature dimensions per token.
- FFN provides position-wise non-linear transformation.
- Causal masking prevents a decoder from seeing future tokens.
- Cross-attention uses Decoder Q and Encoder K/V.
- BERT is encoder-only, GPT-style models are decoder-only, and T5 is encoder-decoder.
- Full self-attention has quadratic sequence-length complexity.
- KV Cache improves autoregressive inference by reusing previous Key/Value states.
Frequently Asked Questions #
What is a Transformer? #
A Transformer is an attention-based neural network architecture introduced in 2017 that processes sequence relationships using attention rather than recurrent processing.
What is self-attention? #
Self-attention allows tokens within the same sequence to calculate how strongly they should attend to one another using Query, Key and Value representations.
What are Q, K and V? #
Query represents what a token is looking for, Key represents what a token can be matched against, and Value represents the information passed into the attention output.
Why is positional encoding needed? #
Because self-attention does not inherently encode the order of tokens. Positional information provides sequence-order information.
Why is attention divided by √dk? #
Scaling controls the magnitude of the dot-product scores and helps prevent Softmax from becoming excessively saturated.
What is causal masking? #
Causal masking prevents a decoder position from attending to future target positions.
What is cross-attention? #
Cross-attention connects the decoder with encoder representations. The Decoder produces Query vectors while the Encoder provides Key and Value vectors.
What is the difference between BERT and GPT? #
BERT is an encoder-only Transformer architecture, while GPT-style models are decoder-only architectures designed for causal autoregressive generation.
Quick Revision #
Transformer → Embedding → Positional Information → Q/K/V → Scaled Dot-Product Attention → Multi-Head Attention → Residual + LayerNorm → FFN → Decoder Mask → Cross-Attention → Linear → Softmax
Attention = Softmax((QKT)/√dk)V
Self-Attention = Same sequence provides Q, K and V
Cross-Attention = Decoder Q + Encoder K/V
Good memory chain: Ask → Match → Retrieve.
Conclusion #
The Transformer architecture changed sequence modeling by replacing recurrent processing with attention-based computation. Its core ideas include positional information, Query-Key-Value projections, scaled dot-product attention, multi-head attention, residual connections, LayerNorm and Feed-Forward Networks.
For encoder-decoder Transformers, the Decoder additionally uses causal masking and cross-attention. These concepts form the foundation for important Transformer families such as BERT, GPT-style models and T5 and provide the architectural basis for many modern LLM systems.
Final memory line: Transformer = Token Representation + Position + Attention + Multi-Head Processing + Stable Deep Blocks + Autoregressive/Sequence-to-Sequence Generation.
Top 50 Transformer Interview Questions and Answers #
If you are preparing for a Machine Learning Engineer, AI Engineer,
Deep Learning Engineer, NLP Engineer, Generative AI Engineer, or LLM Engineer
interview, knowing Transformer definitions alone is not enough.
A strong interviewer may start with a simple question such as
“What is a Transformer?” and then progressively ask:
Why?, How?, What happens if?,
What is the computational cost?, and finally
How would you solve this problem in production?
Therefore, these 50 questions are arranged from
Easy → Intermediate → Advanced → Senior/Production.
The later questions are intentionally more situation-based and require
reasoning rather than memorization.
Level 1: Easy Transformer Interview Questions #
These questions test whether you understand the fundamental building blocks
of Transformer architecture.
1. What is a Transformer? #
A Transformer is a neural network architecture based primarily on
attention mechanisms, especially self-attention.
It was introduced in the paper Attention Is All You Need.
Unlike RNNs and LSTMs, Transformers do not require recurrent sequential
processing. This allows many token positions to be processed in parallel
during training.
The key idea is that a token can directly interact with other tokens instead
of passing information step-by-step through recurrent states.
For example, in:
“The animal didn’t cross the road because it was tired.”
the representation of “it” can attend to relevant words
such as “animal” to build contextual information.
Interview Answer:
“A Transformer is an attention-based neural network architecture that models
relationships between tokens using self-attention. Unlike RNNs, it removes
recurrence and allows highly parallelized sequence processing during training.”
2. Why were Transformers introduced when RNNs already existed? #
RNNs process sequences sequentially. The hidden state at position
t depends on the state at position t-1.
This creates several limitations:
- Sequential computation limits training parallelism.
- Long-range dependencies can be difficult to learn.
- Information has to pass through many recurrent steps.
- Vanishing and exploding gradients can make optimization difficult.
Transformers replace recurrence with attention, allowing tokens to directly
interact with other tokens.
Interview Answer:
“Transformers were introduced to overcome the sequential bottleneck of RNNs
and improve the modeling of long-range relationships while enabling much
greater parallelism during training.”
3. Explain the basic Transformer architecture. #
The original Transformer contains an Encoder and a
Decoder.
The Encoder processes the input sequence.
The Decoder generates the target sequence.
An Encoder block contains self-attention and a feed-forward network along with
residual connections and normalization.
A Decoder block contains masked self-attention, cross-attention, and a
feed-forward network.
4. What is Self-Attention? #
Self-attention allows each token to calculate how strongly it should interact
with other tokens in the same sequence.
The input is projected into three representations:
- Query (Q) — what the token is looking for.
- Key (K) — what information a token offers for matching.
- Value (V) — the information that will actually be retrieved.
The standard scaled dot-product attention formula is:
Attention(Q,K,V) =
Softmax((QKT) / √dk)V
5. What are Query, Key, and Value? #
A useful interview analogy is a search system.
-
Query:
What am I looking for? -
Key:
What information does this token advertise for matching? -
Value:
What information should I retrieve if this token is relevant?
They are learned projections of the input:
Q = XWQ
K = XWK
V = XWV
6. What is Multi-Head Attention and why is it useful? #
Multi-Head Attention performs several attention operations in parallel.
Instead of forcing one attention mechanism to learn every relationship,
different heads can learn different patterns.
For example, different heads may learn relationships involving:
- syntactic dependencies
- semantic relationships
- token-to-token relationships
- different contextual patterns
The outputs of the heads are concatenated and passed through a projection layer.
7. What is positional encoding and why is it necessary? #
Self-attention does not inherently encode the order of tokens.
Consider:
Dog bites man
and:
Man bites dog
The same words appear, but their order changes the meaning.
Positional information gives the model information about token positions.
In the original Transformer, sinusoidal positional encoding was used:
PE(pos, 2i) =
sin(pos / 100002i/dmodel)
PE(pos, 2i+1) =
cos(pos / 100002i/dmodel)
8. What is a residual connection? #
A residual connection adds the input of a sublayer to its output:
Output = x + Sublayer(x)
This creates a direct path for information and gradients through deep
Transformer networks.
Without residual paths, optimization of a deep stack of Transformer layers
can become more difficult.
9. What is Layer Normalization? #
LayerNorm normalizes the feature dimensions of a token representation.
Unlike BatchNorm, LayerNorm does not depend on statistics calculated across
the batch.
This makes it suitable for sequence models where sequence lengths and batch
composition can vary.
10. What is the purpose of the Feed-Forward Network? #
The Feed-Forward Network, or FFN, applies a nonlinear transformation to each
token representation independently after attention.
A simplified form is:
FFN(x) =
ReLU(xW1 + b1)W2 + b2
Attention primarily mixes information between tokens, while the FFN transforms
the representation of each token.
Level 2: Intermediate Transformer Interview Questions #
These questions move beyond definitions and test whether you understand the
internal mechanics of attention and Transformer computation.
11. Why do we divide attention scores by √dk? #
The Query-Key dot product can become large as the dimensionality
dk increases.
If the values entering Softmax become very large, Softmax can become highly
peaked. One or a few positions may receive almost all the probability mass,
while the remaining probabilities become extremely small.
This can result in very small gradients and make optimization less stable.
Scaling by:
√dk
controls the magnitude of the scores before Softmax.
Strong Interview Answer:
“We divide by √dk to control the variance and magnitude of the
dot-product attention scores. Without scaling, large scores can saturate
Softmax and produce very small gradients.”
12. What does each element of QKT represent? #
Suppose:
Q ∈ Rn × dk
and:
K ∈ Rn × dk
Then:
QKT ∈ Rn × n
The element at row i and column j represents
the similarity between the Query of token i and the Key of
token j.
This is why the attention matrix contains approximately
n² pairwise interactions.
13. If dmodel = 512 and there are 8 attention heads, what is the dimension of each head? #
The dimension per head is:
dhead = dmodel / h
Therefore:
512 / 8 = 64
Each head operates on 64 dimensions.
If there are 16 heads:
512 / 16 = 32
Each head would then operate on 32 dimensions, assuming the total model
dimension remains 512.
14. Does increasing the number of attention heads automatically multiply the number of parameters? #
No. The answer depends on how the architecture is parameterized.
In standard Multi-Head Attention, the total model dimension is divided among
the heads. The projection matrices operate on the overall model dimension,
rather than independently using the full model dimension for every head.
Therefore, simply saying:
“8 heads means 8 times more parameters”
is incorrect.
What does change is the way the representation is partitioned across heads.
15. Why does the Transformer need both Attention and an FFN? #
These components perform different operations.
| Component | Main Function |
|---|---|
| Attention | Mixes information between token positions |
| FFN | Applies nonlinear transformation independently to each token |
A useful way to remember this is:
Attention = communication between tokens
FFN = computation/transformation within each token representation
16. What happens if positional information is removed? #
The model loses explicit information about sequence order.
Self-attention can still calculate relationships between token representations,
but the model no longer has a direct signal telling it which token occurred
first, second, third, and so on.
Therefore, sequences containing the same tokens in different orders can become
difficult to distinguish.
This is especially problematic for language because word order often changes
meaning.
17. What is causal masking? #
Causal masking prevents a token from attending to future tokens.
For example, when predicting token 5, the model may use tokens 1–5 but must
not use tokens 6, 7, 8, and so on.
This prevents future information from leaking into the prediction.
18. What would happen if causal masking were accidentally disabled? #
During training, the model could see future target tokens while predicting the
current token.
For example, when predicting:
“AI”
the model might be allowed to see future target words that would not exist yet
during real generation.
This creates future-information leakage.
The model could therefore achieve misleadingly strong training behavior without
having the same information available during autoregressive inference.
Interviewer Follow-up:
“How would you detect this bug?”
A good answer would include checking the attention mask, comparing training
and inference behavior, testing with controlled sequences, and verifying that
future positions receive blocked attention.
19. What is Cross-Attention? #
Cross-attention allows one sequence to attend to another sequence.
In the original Encoder-Decoder Transformer:
- Query comes from the Decoder.
- Key comes from the Encoder output.
- Value comes from the Encoder output.
The Decoder therefore uses its current representation to decide which parts
of the Encoder output are relevant.
20. Why does the Decoder use masked self-attention before cross-attention? #
The Decoder first needs to understand its own already-generated or target-side
context without looking at future tokens.
Masked self-attention provides this causal context.
Cross-attention then connects that Decoder representation to the Encoder output.
So the conceptual flow is:
Level 3: Advanced Transformer Interview Questions #
These questions are designed to test deeper understanding of Transformer
mathematics, complexity, training behavior, and inference.
21. Why is self-attention O(n²d)? #
Let:
- n = sequence length
- d = representation dimension
The main quadratic operation is:
QKT
If Q contains n queries and K contains n keys, the model calculates
approximately:
n × n = n²
pairwise interactions.
Each interaction involves a d-dimensional representation, leading to the
commonly stated complexity:
O(n²d)
22. If the sequence length doubles, what happens to the quadratic attention matrix? #
Suppose the original sequence length is:
n
The attention matrix contains:
n²
entries.
If the sequence length becomes:
2n
the matrix contains:
(2n)² = 4n²
Therefore, the number of pairwise attention positions becomes approximately
4 times larger.
This is the core reason long-context Transformer computation can become
expensive.
23. Why are Transformers good at long-range dependencies? #
In an RNN, information from token 1 may need to pass through many recurrent
steps before influencing token 500.
In self-attention, token 500 can directly attend to token 1 in the same
attention operation.
Therefore, the interaction path between distant tokens can be much shorter.
However, this benefit comes with the quadratic attention cost discussed above.
24. Why is training highly parallelizable but autoregressive inference is not? #
During training, the target sequence is known.
A causal mask prevents each position from seeing future positions, but many
positions can still be processed simultaneously using matrix operations.
During autoregressive inference, the model does not know the next token until
it predicts it.
Therefore:
Token 1 → Token 2 → Token 3 → Token 4 → …
must be generated sequentially.
This distinction is one of the most important concepts in Transformer
inference engineering.
25. What is KV Cache and why is it important? #
During autoregressive generation, previous tokens have already produced their
Key and Value representations.
Recomputing those K/V representations at every decoding step is unnecessary.
KV Cache stores them so that later decoding steps can reuse them.
The main benefit is avoiding repeated computation for historical K/V states.
26. What is the KV Cache trade-off? #
KV Cache improves decoding efficiency, but it consumes memory.
As the sequence grows, more Key and Value tensors must be stored.
Therefore:
Longer context + larger model + more concurrent requests
can result in substantial KV-cache memory usage.
This creates an important inference trade-off:
Compute saved ↔ Memory consumed
27. What happens to KV Cache memory when context grows from 4K to 32K? #
The number of cached tokens increases by:
32K / 4K = 8
So, assuming all other factors remain constant, the amount of K/V state that
must be stored grows approximately by a factor of 8.
This is why long-context inference can become memory-intensive.
In a real serving system, concurrency, batch size, precision, number of layers,
and number of K/V heads also affect total memory usage.
28. What is GQA and why is it useful? #
Grouped-Query Attention allows multiple Query heads to share Key and Value
heads.
Compared with standard Multi-Head Attention, fewer K/V heads can reduce the
amount of K/V state that must be stored.
This is especially useful for autoregressive inference because KV Cache memory
can become a significant bottleneck.
29. What is RoPE and what problem does it solve? #
RoPE stands for Rotary Position Embedding.
It introduces positional information by applying position-dependent rotations
to Query and Key representations.
Because attention is calculated using Q and K, modifying them in a
position-dependent way allows positional relationships to influence attention.
This differs conceptually from simply adding a separate positional vector to
the token embedding.
30. What is FlashAttention? #
FlashAttention is an exact attention algorithm designed to improve the
efficiency of attention computation, particularly on GPUs.
The key idea is to reduce inefficient memory movement by using
tiling and memory-aware computation.
It does not obtain its speedup by simply approximating attention.
The attention calculation remains exact while the implementation is optimized
for the GPU memory hierarchy.
Level 4: Senior and Production-Oriented Transformer Questions #
These questions are intentionally closer to the type of reasoning expected
when interviewing for experienced AI/ML and inference engineering roles.
Current role descriptions emphasize profiling, optimization, KV-cache behavior,
autoregressive decoding, GPU performance, memory constraints and production
serving. :contentReference[oaicite:1]{index=1}
31. Your model has good accuracy but very high inference latency. How would you investigate? #
I would not immediately change the model.
First, I would establish where the latency is coming from.
Step 1 — Define the latency metric
- Time to First Token (TTFT)
- Inter-Token Latency (ITL)
- Total request latency
- Tokens per second
Step 2 — Profile the workload
- GPU utilization
- GPU memory usage
- Memory bandwidth
- Attention kernel time
- CPU-GPU synchronization
- Data-transfer overhead
Step 3 — Inspect Transformer-specific bottlenecks
- Context length
- KV Cache
- Attention implementation
- Batching efficiency
- Number of K/V heads
Only after identifying the bottleneck would I choose an optimization.
Senior Interview Answer:
“I would first profile the inference path and separate TTFT from decode
latency. Then I would determine whether the bottleneck is compute, memory
bandwidth, KV-cache pressure, batching, synchronization, or the attention
kernel. I would optimize the measured bottleneck and validate the result using
latency, throughput, memory, and quality metrics.”
32. Your Transformer fits during training but causes GPU OOM during inference. Why? #
This can happen even though inference does not perform backpropagation.
Possible causes include:
- Much longer inference context.
- KV Cache consuming substantial memory.
- Large inference batch size.
- Multiple concurrent requests.
- Temporary inference buffers.
- Different precision or memory configuration.
- Memory fragmentation.
A common mistake is to look only at model parameter memory.
For large autoregressive models, runtime state such as KV Cache can be a major
part of the memory footprint.
33. Production latency increases sharply when context length increases. What would you check? #
I would first reproduce the problem using controlled context lengths such as:
2K → 4K → 8K → 16K → 32K
Then I would measure:
- TTFT
- Decode latency
- Attention execution time
- KV-cache memory
- GPU utilization
- Memory bandwidth
If the problem scales strongly with context length, I would investigate the
quadratic attention workload and the increasing KV-cache state.
I would then evaluate optimized attention, cache-management strategies and
architecture choices such as GQA where appropriate.
34. GPU utilization is only 30% and inference is slow. Is the GPU the problem? #
Not necessarily.
Low GPU utilization can occur when the workload is limited by:
- Memory bandwidth
- CPU preprocessing
- CPU-GPU synchronization
- Small workloads
- Kernel launch overhead
- Data movement
- Network communication
Therefore, GPU utilization alone is not sufficient to identify the bottleneck.
A profiler should be used to determine whether the workload is
compute-bound, memory-bound, communication-bound, or overhead-bound.
35. How would you determine whether attention is compute-bound or memory-bound? #
I would profile both computation and memory behavior.
Important measurements include:
- GPU compute utilization
- Memory bandwidth utilization
- Kernel execution time
- Arithmetic intensity
- Data movement
- GPU occupancy
If the hardware is spending most of its time performing arithmetic operations,
the workload may be compute-bound.
If the limiting factor is moving data between memory levels, it may be
memory-bound.
This distinction determines what optimization should be attempted next.
36. You have two Transformer models with similar accuracy. One is twice as fast. Which would you deploy? #
I would not decide using accuracy alone.
I would compare the models using the actual production requirements:
| Metric | Why It Matters |
|---|---|
| Accuracy / Quality | Ensures acceptable model behavior |
| Latency | Determines user experience |
| Throughput | Determines serving capacity |
| GPU Memory | Affects deployment constraints |
| Cost per request | Affects production economics |
| Reliability | Important under production load |
The correct engineering decision depends on the service-level requirements,
not simply on which model has higher benchmark accuracy.
37. You increase the context window from 8K to 64K. What problems would you expect? #
Several issues can appear.
- Higher attention computation.
- Larger attention matrices.
- Higher KV-cache memory usage.
- Higher latency.
- Lower maximum concurrency.
- Potential GPU memory pressure.
The key observation is that increasing context length is not simply a matter
of changing one configuration value. It can significantly affect both
computation and memory.
38. How would you reduce Transformer inference memory without immediately changing the model architecture? #
I would investigate several layers of the inference stack:
- Reduce unnecessary context.
- Use appropriate numerical precision.
- Optimize KV-cache storage.
- Improve batching and scheduling.
- Use memory-efficient attention.
- Reduce redundant intermediate buffers.
The correct choice depends on whether the memory is primarily consumed by
parameters, activations, KV Cache, temporary buffers, or concurrent requests.
39. A Transformer model suddenly produces NaN loss during training. How would you debug it? #
I would debug systematically instead of immediately changing the learning rate.
- Check the input data for NaN or infinite values.
- Check the attention logits before Softmax.
- Verify that the attention mask is valid.
- Check for numerical overflow in mixed-precision computation.
- Inspect gradient norms.
- Check the learning rate and optimizer configuration.
- Verify that normalization layers are behaving correctly.
- Reproduce the issue with a smaller batch or controlled input.
The goal is to determine whether the NaN originates from the data,
forward pass, attention computation, loss, or backward pass.
40. Attention weights become almost one-hot very early during training. What might be happening? #
One possibility is that attention logits are becoming too large.
If the logits have very large differences, Softmax can become highly peaked.
This may cause one token to receive nearly all the attention weight.
I would investigate:
- Magnitude of Q and K
- Scaling by √dk
- Normalization behavior
- Learning rate
- Numerical precision
This question tests whether the candidate understands the relationship between
dot-product magnitude, Softmax saturation and gradient behavior.
41. You are asked to design a low-latency Transformer inference service. What would you consider? #
I would begin by defining the service requirements:
- Target latency
- Expected requests per second
- Maximum context length
- Maximum output length
- Quality requirements
- GPU memory constraints
- Concurrency requirements
Then I would design around the actual bottlenecks.
At the Transformer level, I would consider:
- KV Cache management
- GQA
- Efficient attention kernels
- Batching
- Scheduling
- Appropriate precision
Finally, I would benchmark the complete serving system rather than evaluating
only model-level inference time.
42. Why can KV Cache become the bottleneck in a large-scale serving system? #
KV Cache grows with the amount of context stored for active sequences.
In a serving system, the problem is not only one request.
There may be hundreds or thousands of concurrent sequences.
Therefore:
Per-request KV memory × Number of active requests
can become a major GPU memory constraint.
This can reduce concurrency even when the model parameters themselves fit
comfortably in memory.
43. Why can GQA improve inference scalability? #
GQA reduces the number of distinct Key and Value heads.
Since the KV Cache stores K/V representations, fewer K/V heads can reduce the
memory required per cached token.
Lower KV-cache memory can allow:
- More concurrent sequences
- Longer contexts
- Less memory pressure
- Potentially better inference efficiency
However, the architecture must be evaluated for quality and workload-specific
performance.
44. FlashAttention is exact. How can it be faster without approximating attention? #
The important distinction is between the mathematical operation and the way it
is implemented on hardware.
Standard implementations may materialize or repeatedly move large intermediate
attention data through GPU memory.
FlashAttention uses tiling and a memory-aware computation strategy so that
data movement is reduced.
Therefore, the speedup comes primarily from improving the
memory access pattern and hardware utilization, not from
changing the mathematical attention result into an approximation.
45. Your model has excellent benchmark results but poor production performance. Why? #
A benchmark may not represent the actual production workload.
Production may introduce:
- Different context lengths
- Higher concurrency
- Different batch sizes
- Variable request lengths
- KV-cache pressure
- Network overhead
- CPU preprocessing
- Scheduling overhead
Therefore, a model can perform well on a benchmark while behaving differently
under real production traffic.
A senior engineer should benchmark using realistic workload distributions,
not only a single idealized example.
Level 5: Very Hard / Senior Interview Questions #
These final questions are designed to force the candidate to reason about
trade-offs rather than simply recall Transformer terminology.
46. Your inference server is memory-bound. Would adding a larger GPU automatically solve the problem? #
Not necessarily.
A larger GPU may provide more memory capacity or bandwidth, but the actual
bottleneck must first be identified.
For example, if the system is limited by:
- CPU preprocessing
- Network communication
- Synchronization
- Inefficient kernels
- KV-cache layout
- Memory access patterns
then simply adding more GPU compute capacity may not solve the real problem.
A strong answer is:
“Profile first, identify the bottleneck, then select hardware or software
optimization.”
47. A team wants to increase context length from 16K to 128K. What questions would you ask before approving the change? #
I would ask:
- What percentage of real requests actually require 128K?
- What is the expected latency increase?
- What is the KV-cache memory requirement?
- How will concurrency change?
- What GPU memory is available?
- Does model quality actually improve at 128K?
- What happens to cost per request?
- What is the maximum acceptable latency?
- Can long-context requests be handled differently from normal requests?
This is a system-level question: increasing context is an architectural and
operational decision, not simply a model configuration change.
48. Design a debugging strategy for a Transformer whose inference latency suddenly doubled after a deployment. #
I would first determine whether the model itself changed or whether the serving
environment changed.
Step 1: Compare versions
- Model weights
- Model configuration
- Inference framework
- CUDA/runtime version
- Attention kernel
- Hardware
Step 2: Compare workload
- Average input length
- Average output length
- Concurrency
- Batch size
Step 3: Profile
Compare kernel execution times, GPU utilization, memory bandwidth and
CPU/network overhead against the previous deployment.
Step 4: Isolate
Run the same model with the same input distribution on both versions.
This separates a workload change from an implementation regression.
Only after isolating the regression would I change the implementation.
49. You need to choose between reducing model size and optimizing inference kernels. How would you decide? #
I would first define the actual bottleneck and business constraint.
If the model is already accurate enough but compute cost is too high,
reducing model size may help.
If the model has acceptable computational requirements but spends too much time
in inefficient kernels or memory movement, kernel optimization may provide
better results without changing model quality.
I would compare:
| Option | Potential Benefit | Potential Risk |
|---|---|---|
| Smaller Model | Lower compute and memory | Potential quality loss |
| Kernel Optimization | Lower latency without changing model | Engineering complexity |
| GQA | Lower KV memory | Architecture/quality trade-offs |
| Efficient Attention | Better memory behavior | Hardware/framework dependency |
The correct answer is not automatically one option. The decision should be
based on measured bottlenecks and production requirements.
50. You are the senior engineer responsible for improving a Transformer serving system. What would your optimization strategy be? #
I would use a measurement-driven optimization process.
1. Define the targets
- TTFT
- Inter-token latency
- Throughput
- GPU memory
- Cost per request
- Quality
2. Establish a realistic benchmark
Use realistic distributions of prompt length, output length, concurrency and
request arrival patterns.
3. Profile
Identify whether the bottleneck is:
- Compute
- Memory bandwidth
- KV Cache
- Attention kernels
- CPU overhead
- Network communication
- Scheduling
4. Optimize the measured bottleneck
- Efficient attention
- KV-cache optimization
- GQA where appropriate
- Batching
- Precision optimization
- Kernel optimization
- Scheduling improvements
5. Measure again
Every optimization must be evaluated against the same benchmark.
A senior-level answer should explicitly mention that an optimization is not
successful merely because one metric improves. The final decision must consider
latency, throughput, memory, cost, reliability and model quality.
Strong Senior Interview Answer:
“I would use a measurement-driven optimization loop: define production SLOs,
build a representative benchmark, profile the complete inference path, identify
the dominant bottleneck, apply the smallest appropriate optimization, and
re-measure latency, throughput, memory, cost and quality. I would avoid
optimizing components that profiling shows are not on the critical path.”
Transformer Interview Difficulty Map #
| Level | Questions | What Interviewer Tests |
|---|---|---|
| Easy | 1–10 | Transformer fundamentals and terminology |
| Intermediate | 11–20 | Attention mechanics and architectural reasoning |
| Advanced | 21–30 | Mathematics, complexity and inference concepts |
| Senior | 31–45 | Debugging, profiling and production optimization |
| Very Hard | 46–50 | Architecture decisions and senior-level engineering reasoning |
What Makes These Questions Different? #
A weak Transformer interview question asks:
“What is KV Cache?”
A stronger interview question asks:
“Your 32K-context production model is running out of GPU memory. How would you
determine whether KV Cache is the cause, and what would you change?”
The second question tests multiple skills at once:
- Understanding of KV Cache
- Memory reasoning
- Profiling
- Debugging
- Architecture knowledge
- Trade-off analysis
That is the style this question set is designed to develop.
How to Answer Transformer Questions in a Real Interview #
For easy questions, give a direct definition followed by the reason.
For advanced questions, use:
- Concept — What is happening?
- Reason — Why does it happen?
- Mechanism — How does the Transformer implement it?
- Cost — What is the computational or memory impact?
- Trade-off — What do we gain and lose?
- Production — How would you measure or debug it?
For example, if asked about KV Cache, do not stop at:
“KV Cache stores previous Keys and Values.”
A stronger answer continues:
“It avoids recomputing historical K/V states during autoregressive decoding.
The trade-off is increased memory usage, which becomes important with long
contexts and high concurrency. I would therefore monitor cache memory,
sequence length, concurrency and decode latency.”
Final Transformer Interview Checklist #
- Transformer architecture
- Self-Attention
- Query, Key and Value
- Scaled Dot-Product Attention
- Multi-Head Attention
- Positional Information
- Residual Connections
- LayerNorm
- Feed-Forward Network
- Causal Masking
- Cross-Attention
- Training vs Inference
- Autoregressive Generation
- Attention Complexity
- KV Cache
- Long-Context Memory
- RoPE
- GQA
- FlashAttention
- GPU Profiling
- Compute vs Memory Bottlenecks
- Inference Latency
- GPU OOM Debugging
- NaN Debugging
- Production Architecture
Final Takeaway #
The goal of Transformer interview preparation should not be to memorize
50 definitions. A strong candidate should be able to move from:
“What is this?”
to:
“Why does it exist?”
then:
“How does it work mathematically?”
and finally:
“What would you do if this failed in production?”
That progression is what makes these questions useful for both
entry-level Transformer interviews and more advanced
AI/ML and inference-engineering interviews.