Second interview track
AI Interview Roadmap
36 basic, high-frequency AI interview questions with interview-ready answers. Focused on the fundamentals of LLMs, prompting, RAG, agents and production AI — enough to build a strong foundation before moving into advanced topics.
36 questions6 categoriesQuestions + answersSenior/Staff foundation
How to prepare: first answer each question yourself in 30–60 seconds. Then compare with the answer shown. For Senior/Staff interviews, add one production example and one trade-off to your answer.
1. LLM Fundamentals
6 questions
The minimum LLM knowledge expected from a senior backend/software engineer.
Q01
What is an LLM?
Answer: A Large Language Model is a neural network trained on large amounts of text to predict the next token. In applications, it can generate, classify, summarize and transform text based on the supplied context.
Q02
What is a token?
Answer: A token is a unit of text processed by the model. It can be a word, part of a word, punctuation, or another text fragment. Token count affects context limits, latency and cost.
Q03
What is an embedding?
Answer: An embedding is a numeric vector representing the semantic meaning of text. Similar concepts tend to have nearby vectors, which makes embeddings useful for semantic search and RAG.
Q04
What is a context window?
Answer: The context window is the amount of input and output context a model can process for a request. Larger context can improve access to information but increases cost and does not automatically guarantee better answers.
Q05
What is temperature?
Answer: Temperature controls how much randomness is used during generation. Lower values generally make output more deterministic; higher values produce more varied output. For factual structured tasks, lower temperature is usually preferred.
Q06
Why do LLMs hallucinate?
Answer: An LLM generates likely text rather than directly querying a truth database. Missing context, ambiguous prompts and model limitations can therefore produce plausible but incorrect answers. Grounding, retrieval, validation and constrained outputs reduce the risk.
2. Prompting & Tool Calling
6 questions
Questions that test whether you can reliably control an LLM from software.
Q07
What makes a good production prompt?
Answer: Clearly define the task, relevant context, constraints, expected output format and failure behavior. Keep instructions unambiguous and test the prompt against representative examples.
Q08
Zero-shot vs few-shot prompting?
Answer: Zero-shot gives instructions without examples. Few-shot includes examples of the desired behavior. Few-shot can improve consistency when the task has a specific format or subtle classification boundary.
Q09
How do you get structured JSON from an LLM?
Answer: Prefer a provider's structured-output or schema/function-calling capability when available. Validate the response in application code and retry or recover when validation fails.
Q10
What is function/tool calling?
Answer: The model decides which predefined tool to call and supplies structured arguments. Your application executes the tool, returns the result to the model, and the model can then produce the final response.
Q11
Prompt injection vs normal user input?
Answer: Prompt injection occurs when untrusted content attempts to override the application's instructions or manipulate tool use. Treat retrieved documents and user input as untrusted data, enforce permissions outside the prompt, and validate tool actions.
Q12
When should you use an LLM versus normal code?
Answer: Use normal deterministic code when rules are explicit and correctness must be guaranteed. Use an LLM where language understanding, generation or ambiguity provides value. Strong systems usually combine both.
3. Embeddings & RAG Basics
6 questions
The core GenAI topic for enterprise applications.
Q13
What is RAG?
Answer: Retrieval-Augmented Generation retrieves relevant external information and places it into the model's context before generation. This lets an application answer using current or private data without retraining the model for every document change.
Q14
Why do we need a vector database?
Answer: It stores embeddings and efficiently finds vectors similar to a query embedding. This provides semantic retrieval over large collections of documents.
Q15
What is chunking?
Answer: Chunking splits documents into smaller pieces before embedding them. Good chunks preserve enough meaning to answer questions while staying small enough for efficient retrieval and context usage.
Q16
How do you choose chunk size?
Answer: Start from document structure and the expected question size rather than a universal number. Evaluate retrieval quality, context size and answer quality, then tune using real queries.
Q17
What is hybrid search?
Answer: Hybrid search combines lexical retrieval such as BM25 with semantic vector retrieval. Lexical search is strong for exact terms, identifiers and names, while vector search handles semantic similarity.
Q18
Why use reranking?
Answer: Initial retrieval is optimized for recall and speed. A reranker can examine the top candidates more deeply and reorder them so the most relevant chunks are passed to the LLM.
4. RAG Interview Questions
6 questions
Basic RAG troubleshooting and design questions worth knowing cold.
Q19
Explain a basic RAG architecture.
Answer: Documents are ingested, cleaned, chunked and embedded into an index. At query time, the question is embedded or searched, relevant chunks are retrieved and optionally reranked, then the LLM generates an answer using those chunks as context.
Q20
What happens if RAG retrieves the wrong documents?
Answer: The problem is primarily retrieval quality. Inspect the query, chunking, embedding model, metadata filters, top-k and ranking. Improve retrieval before trying to fix generation with a larger prompt.
Q21
What if RAG retrieves the correct document but the answer is wrong?
Answer: Separate retrieval from generation. Check whether the relevant chunk was actually included, whether the prompt clearly requires grounding, whether context was truncated, and whether the model incorrectly interpreted the evidence.
Q22
How do you reduce hallucinations in RAG?
Answer: Provide high-quality retrieved evidence, instruct the model to answer only from supported context, require citations where appropriate, and allow the model to say it does not know. Evaluate groundedness with representative test cases.
Q23
How do you handle document updates?
Answer: Track document versions and IDs, re-embed changed content, remove or invalidate deleted chunks, and keep metadata that supports freshness. For critical systems, design an explicit indexing pipeline rather than relying on manual updates.
Q24
How do you secure RAG?
Answer: Authorization must be enforced before retrieved content reaches the model. Store tenant/permission metadata, filter retrieval by the caller's permissions, and avoid relying on the prompt to enforce access control.
5. Agentic AI Basics
6 questions
Keep the first agent questions practical rather than research-heavy.
Q25
What is an AI agent?
Answer: An agent is an LLM-driven system that can decide what actions to take, call tools, observe results and continue toward a goal. A normal LLM call simply generates an answer from its current context.
Q26
Agent vs workflow?
Answer: A workflow follows a predefined sequence of steps. An agent can choose the next action dynamically. Use workflows when the process is predictable; use agents when decisions or tool selection genuinely need model-driven flexibility.
Q27
What is ReAct?
Answer: ReAct is a pattern where the model reasons about the current task, selects an action or tool, observes the result, and continues until it reaches a stopping condition.
Q28
How do you stop an agent from looping?
Answer: Set maximum steps or token budgets, detect repeated states/actions, define explicit termination conditions, and make tools idempotent where possible. Never depend only on the model deciding when to stop.
Q29
How do you handle failed tools?
Answer: Validate tool inputs and outputs, use bounded retries for transient failures, apply timeouts, return structured error information to the agent, and provide a fallback or human escalation for important operations.
Q30
When should you avoid agents?
Answer: Avoid them when a deterministic workflow can solve the problem more cheaply, predictably and safely. Agentic behavior adds latency, cost and more failure modes.
6. Production AI Basics
6 questions
The senior-level basics that turn a demo into an engineering system.
Q31
How do you reduce LLM latency?
Answer: Use an appropriately sized model, stream responses, reduce unnecessary context, parallelize independent operations, optimize retrieval, and cache deterministic or reusable work.
Q32
How do you reduce LLM cost?
Answer: Control input/output tokens, use smaller models where sufficient, cache repeated work, route requests to different models based on complexity, and measure cost per successful task rather than only cost per request.
Q33
What should you monitor in an LLM application?
Answer: Track latency, token usage, cost, errors, model versions, retrieval quality, tool failures and user/task outcomes. For RAG and agents, trace the intermediate retrieval and tool steps as well.
Q34
How do you evaluate a RAG system?
Answer: Create a representative evaluation dataset and measure retrieval relevance/recall plus answer correctness and groundedness. Combine automated metrics with human review for important cases.
Q35
RAG vs fine-tuning?
Answer: Use RAG when the model needs access to changing or private knowledge. Fine-tuning is more appropriate when you need to change model behavior, style or task performance based on examples. They can also be combined.
Q36
What is a good AI system-design answer?
Answer: Start with requirements and scale, then design ingestion/data flow, retrieval or model calls, APIs, storage, caching, async processing, security, evaluation, observability, failure handling, latency and cost. Explain trade-offs rather than drawing boxes only.
OPTIONAL📚 Explore More — Advanced AI Topics
Not required for the core roadmap. Open these only when you want deeper GenAI, RAG, Agentic AI and production knowledge.
7. LLM Deep Dive
6 topics
Useful when interviews go beyond application-level LLM knowledge.
37 Transformer architecture
Know: Self-attention, positional information, feed-forward layers, residual connections and why transformers scale well.
38 Self-attention vs cross-attention
Know: Self-attention relates tokens within the same sequence; cross-attention lets one representation attend to another.
39 Decoder-only vs encoder-decoder models
Know: Decoder-only models are common for generation; encoder-decoder models separate input understanding from output generation.
40 Model parameters and model size
Know: Parameters are learned weights. More parameters can increase capability but also memory, compute and serving cost.
41 Quantization
Know: Represent weights or activations with lower precision to reduce memory and inference cost, usually with some quality trade-off.
42 Inference vs training
Know: Training learns model parameters; inference uses the trained model to produce outputs for new inputs.
8. Advanced RAG
7 topics
Explore these after the basic RAG architecture is comfortable.
43 Metadata filtering
Know: Restrict retrieval using tenant, document type, date or access-level metadata.
44 Query rewriting
Know: Transform an ambiguous or conversational query into a retrieval-friendly query.
45 HyDE
Know: Generate a hypothetical document and use its embedding to improve retrieval for some queries.
46 Parent-child retrieval
Know: Retrieve small chunks for precision but provide a larger parent section for surrounding context.
47 Multi-vector retrieval
Know: Represent different aspects of a document with multiple vectors when one embedding is insufficient.
48 GraphRAG
Know: Combine retrieval with a knowledge graph or graph-derived structure to reason over relationships.
49 Retrieval evaluation
Know: Measure whether relevant documents are retrieved using metrics such as recall@k and precision@k.
9. Agentic AI — Deeper Concepts
7 topics
For interviews that go beyond a single tool-calling agent.
50 Agent memory
Know: Short-term conversation state and longer-term persisted information, with retention and privacy rules.
51 Planning in agents
Know: Decompose a goal into actions using explicit plans, model-generated plans or deterministic workflows.
52 Reflection
Know: Review an intermediate result, identify problems and attempt an improvement with bounded iterations.
53 Multi-agent systems
Know: Multiple specialized agents collaborate through defined messages or tasks, adding coordination trade-offs.
54 Agent orchestration
Know: Control state, tools, retries, routing, budgets and termination outside the model where possible.
55 Human-in-the-loop agents
Know: Pause for human approval before high-impact actions such as financial or destructive operations.
56 Agent state machines
Know: Explicit states and transitions make agent workflows easier to reason about, retry, observe and recover.
10. Model Adaptation
6 topics
Useful when interviewers ask when prompting or RAG is not enough.
57 Fine-tuning
Know: Further train a pretrained model on task-specific examples to change behavior or improve a target task.
58 RAG vs fine-tuning
Know: RAG supplies changing knowledge at inference time; fine-tuning changes model behavior from training examples.
59 LoRA
Know: Train small low-rank matrices while keeping most base-model weights frozen, reducing fine-tuning cost.
60 QLoRA
Know: Combine quantization of the base model with LoRA adapters for memory-efficient fine-tuning.
61 Instruction tuning
Know: Train on instruction-and-response examples so a model follows user instructions more effectively.
62 When not to fine-tune
Know: Avoid it when changing knowledge is the main problem or prompting/RAG is sufficient.
11. Production AI Architecture
8 topics
Senior/Staff-level follow-ups for building reliable AI systems.
63 AI gateway
Know: Centralize authentication, model routing, rate limits, logging, policies and provider abstraction.
64 Model routing
Know: Route requests based on complexity, latency, cost, modality or quality requirements.
65 Semantic caching
Know: Reuse responses for semantically similar requests when correctness and freshness requirements allow.
66 Guardrails
Know: Use input/output validation, policy checks, tool permissions and deterministic controls.
67 AI observability
Know: Trace prompts, model calls, retrieval, tools, latency, tokens, errors and outcomes end-to-end.
68 Model fallback strategy
Know: Use bounded fallbacks for provider failures, rate limits or model errors without uncontrolled retries.
69 AI rate limiting
Know: Apply per-user, tenant and application limits, with token-aware controls where appropriate.
70 AI system-design trade-offs
Know: Discuss quality, latency, cost, availability, security, freshness and operational complexity.
These are foundation questions. Advanced topics such as GraphRAG, HyDE, LoRA/QLoRA, multi-agent orchestration and advanced evaluation can be added later as a separate advanced track.