How to Decide Chunk Size in Any Project: Complete Interview Guide A new interview guide outlines a systematic framework for determining chunk size in retrieval-augmented generation projects, emphasizing that chunk size is a design decision based on document type, retrieval precision, token budget, LLM context window, and latency. The guide provides specific ranges—legal documents 1000-2000 words, news articles 400-800 words—and includes a calculation example using Claude 3 Sonnet's 200K token window, allocating 60% for context, 15% for prompt, 15% for response, and 10% safety buffer, resulting in 10-15 chunks of 2000 characters each. It also lists embedding model limits, such as OpenAI's text-embedding-3-large at 8,191 tokens and Cohere at 1,024 tokens, which constrain chunk size before embedding. This guide covers the decision-making framework that interviewers expect you to explain, with real-world trade-offs and problem-solving approach. What Interviewer Wants to Hear: "Chunk size is NOT a fixed number - it's a design decision based on multiple factors: document type, retrieval precision, token budget, LLM context window, and latency requirements. I use a systematic approach: analyze constraints → model requirements → test with metrics → iterate." Interview Question: “How would chunk size differ for legal documents vs news articles?” Answer Template: Legal documents 1000-2000 words :- Why: A clause or section is an indivisible unit of meaning- Risk: Breaking mid-clause creates ambiguity in retrieval- Example: Contract terms must be complete to be accurate News articles 400-800 words :- Why: Readers expect paragraph-level information- Benefit: Smaller chunks allow precise topic retrieval- Trade-off: May need overlap to connect ideas Ask yourself these questions: 1. PRECISION REQUIREMENT □ High precision needed? → SMALLER chunks 400-600 words Example: Legal discovery, medical diagnosis □ Moderate precision? → MEDIUM chunks 700-1200 words Example: Technical support, FAQ systems □ Broad context OK? → LARGER chunks 1500-2500 words Example: General knowledge Q&A 2. RETRIEVAL TYPE □ Exact match needed? → SMALLER chunks Example: "Find this specific clause" □ Semantic match? → MEDIUM chunks Example: "Find information about refunds" □ Topic match? → LARGER chunks Example: "Explain machine learning basics" 3. LATENCY REQUIREMENTS □ Real-time < 500ms ? → SMALLER chunks Faster embedding, faster search □ Sub-second < 2s ? → MEDIUM chunks □ Can tolerate delay? → LARGER chunks More context means better answers 4. COST SENSITIVITY □ Budget tight? → SMALLER chunks Fewer embeddings to compute □ Can spend? → LARGER chunks More embeddings = better precision Interview Answer Example: "For a chatbot answering customer support tickets, I'd analyze:- Precision: Medium-high accuracy matters - Retrieval: Semantic users phrase questions differently - Latency: < 2 seconds acceptable for chat - Cost: Moderate budget This suggests 600-1000 word chunks with 15% overlap." Key calculation interviewers expect: CONTEXT BUDGET = LLM CONTEXT WINDOW 0.6 Reserve 60% for safety Example with Claude 3 Sonnet 200K tokens total tokens = 200 000safe budget = total tokens 0.6 120,000 tokensreserved for prompt = 20 000 tokensreserved for response = 10 000 tokensavailable for context = 90 000 tokens Now work backwards from chunks:tokens per 1000 chars = 250 Rough estimateavailable chars = 90 000 / 250 1000 ~360,000 characters If chunk size = 2000 chars, how many chunks can we fit?max chunks = 360 000 / 2000 ~180 chunks But typically use only top-5 to top-10 chunksreasonable budget = 10 2000 20,000 chars = 5,000 tokens Interview Answer: "Given Claude's 200K token window, I'd allocate:- 60% for context 120K tokens - 15% for prompt 30K tokens - 15% for response 30K tokens - 10% safety buffer 20K tokens If each 2000-char chunk ≈ 500 tokens, I can safely fit 10-15 chunks. Working backward: 15 chunks 2000 chars = 30,000 chars maximum context. For better retrieval quality, I'd use parent-child chunking where top-5 most relevant child chunks expand to their parent chunks for full context." Most embedding models have input limitations: EMBEDDING LIMITS = { "OpenAI text-embedding-3-large ": "8,191 tokens ~6,000 words ", "Anthropic own embeddings ": "varies", "Cohere": "1,024 tokens safe limit ", "Local models ONNX ": "512-2048 tokens typically",} This limits chunk size BEFORE embedding: If embedding model accepts 1024 tokens max: That's roughly 750-1000 words or 3000-4000 characters Safe chunk size < embedding limitCHUNK SIZE = min 2000 chars, embedding model limit Interview Question: How does embedding model capacity affect your chunk size decision? Answer: "Embedding models have token limits. For example, Cohere's model safely handles 1024 tokens. Since 1 token ≈ 4 characters:- 1024 tokens ≈ 4000 characters max Even if I want 2000-char chunks for my use case, I must verify the embedding model can handle it. If it can't, I either:1. Use a more capable embedding model OpenAI's 8K model 2. Reduce chunk size 1500-2000 chars 3. Split chunks: embed smaller, but retrieve with parent for context The embedding model is often the limiting factor." class ChunksizeMetrics: """ What interviewers expect you to measure: """ def init self : self.metrics = { Retrieval Quality "Precision@K": "Of top-K results, how many relevant?", "Recall@K": "Of all relevant docs, how many in top-K?", "NDCG@K": "Normalized Discounted Cumulative Gain ranking quality ", "MRR": "Mean Reciprocal Rank position of first relevant result ", Cost & Performance "Embedding Time": "Seconds to embed all chunks", "Search Latency": "Seconds to find top-K results", "Storage Size": "GB needed to store all embeddings", "Cost Per Query": "$ cost API calls for embeddings ", Answer Quality "Hallucination Rate": "% of made-up information in answers", "Citation Accuracy": "% of citations point to correct source", "Context Coverage": "% of relevant information in context", "Token Efficiency": "Useful info per token in context" } Interviewer Question: How would you evaluate if your chunk size is optimal? Answer: "I'd run experiments with 3 chunk sizes: small 400 words , medium 800 words , large 1500 words . For each, I measure: 1. RETRIEVAL QUALITY does RAG find relevant info? - Precision@5: "Of top 5 results, are they relevant?" - Recall@10: "Of all relevant documents, how many appear?" - Target: Precision 0.8, Recall 0.7 bash 2. COST & LATENCY can we afford it? - Embedding cost: $ per 1M chunks - Search latency: milliseconds - Storage: GB for all embeddings - Target: <100ms latency, <1KB per chunk metadata 3. ANSWER QUALITY does LLM generate good responses? - Hallucination rate: Manual review of 50 samples - Citation accuracy: Does answer reference correct chunks? - Token efficiency: Useful info per token - Target: <5% hallucinations, 90% citations accurate Then I pick the size that balances these metrics best." class ChunkSizeExperiment: """ Structure your testing like a real data scientist """ CHUNK SIZES = 400, 600, 800, 1000, 1500 Words EVALUATION SET = 100 Test queries def run experiment self : results = {} for size in self.CHUNK SIZES: Step 1: Create chunks of this size chunks = create chunks text, size words=size Step 2: Embed all chunks embeddings = embed all chunks storage.index embeddings Step 3: Run test queries metrics = { "precision 5": 0.0, "recall 10": 0.0, "latency ms": 0.0, "cost usd": 0.0, "hallucination rate": 0.0, } for query in EVALUATION SET: results top5 = retrieve query, k=5 Evaluate relevance, measure latency, etc. results size = metrics return self.analyze results results def analyze results self, results : """ Present findings professionally """ print """ CHUNK SIZE ANALYSIS RESULTS ============================ Size | Precision | Recall | Latency | Cost | Halluc. ------|-----------|--------|---------|-------|-------- 400w | 0.85 | 0.72 | 45ms | $0.8 | 3.2% 600w | 0.88 | 0.78 | 50ms | $1.2 | 2.1% 800w | 0.86 | 0.82 | 55ms | $1.6 | 1.8% ⭐ BEST 1000w | 0.82 | 0.80 | 65ms | $2.0 | 1.5% 1500w | 0.79 | 0.75 | 80ms | $3.0 | 2.0% RECOMMENDATION: 800 words - Highest Recall 0.82 - Good Precision 0.86 - Reasonable latency 55ms - Manageable cost - Lowest hallucination rate 1.8% """ Interview Presentation: "I'd create a simple experiment with 5 chunk sizes and run 100 test queries. For each size, I measure precision, recall, latency, and cost. Based on the results, 800-word chunks provide the best balance: high recall fewer missed docs , good precision fewer irrelevant results , and reasonable cost. The key insight: larger chunks give better recall more context , but smaller chunks give better precision less noise . 800 words is the sweet spot." class DomainSpecificChunkSizes: """ Interview tip: Show you understand domain context """ DOMAINS = { "LEGAL": { "size": "1000-2000 words", "reason": "Sections/clauses are legal units", "example": "Contract clause must be complete", "key metric": "Precision Recall accuracy critical ", "overlap": "20% preserve clause boundaries ", }, "MEDICAL": { "size": "500-1000 words", "reason": "Patient outcomes depend on complete context", "example": "Symptoms + test results + diagnosis", "key metric": "Recall Precision miss nothing ", "overlap": "20% connect symptoms to outcomes ", }, "E-COMMERCE": { "size": "400-800 words", "reason": "Product info is naturally separated", "example": "Product specs, reviews, shipping info", "key metric": "Speed real-time product search ", "overlap": "10% less critical ", }, "CUSTOMER SUPPORT": { "size": "600-1000 words", "reason": "Q&A pairs with explanation", "example": "Question + answer + examples", "key metric": "User satisfaction answers must be complete ", "overlap": "15%", }, "TECHNICAL DOCS": { "size": "500-1000 words", "reason": "API docs, parameters need to stay together", "example": "Function signature + params + examples", "key metric": "Accuracy wrong example breaks code ", "overlap": "15%", }, "NEWS/MEDIA": { "size": "400-600 words", "reason": "Articles are already well-written units", "example": "One news story = one natural unit", "key metric": "Latency real-time relevance ", "overlap": "10%", }, } Interview Answer Example: "For a medical chatbot vs an e-commerce bot, chunk sizing would be very different: MEDICAL diagnosing symptoms :- Chunk size: 700-1000 words- Why: Symptoms, tests, diagnosis, treatment must be together- Metric: Optimize for RECALL don't miss anything - Overlap: 20% very important - connect related symptoms E-COMMERCE product recommendation :- Chunk size: 500-800 words - Why: Products are naturally separate; specs are self-contained- Metric: Optimize for SPEED <100ms and cost- Overlap: 10% less critical The key difference: Medical prioritizes completeness; e-commerce prioritizes speed." ❌ "We just use 512-token chunks like everyone else" → Shows no independent thinking ❌ "Bigger chunks are always better" → Ignores retrieval precision vs recall trade-off ❌ "We never tested different chunk sizes" → Suggests no systematic approach ❌ "Chunk size doesn't matter much" → Shows ignorance of its impact ✅ Instead say: "We systematically tested 5 different chunk sizes on our evaluation set, measuring precision, recall, latency, and cost. Based on the results, 800-word chunks provided the best balance for our use case." Answer: "Good question. This is where recursive chunking + parent-child hierarchy becomes critical. Approach:1. DETECT boundaries: Identify natural sections, paragraphs, sentences2. RECURSIVE splitting: Try to split at paragraph level first3. FALL BACK gracefully: If paragraph chunk size, split by sentences4. PRESERVE context: Use parent chunks to keep full section Example: If I want 800-word chunks but a legal clause is 1200 words:- Don't: Break the clause loses meaning - Do: Keep clause as one parent chunk Split into 2-3 child chunks for embedding Retrieve child chunks + expand to full parent for LLM This preserves semantic integrity while optimizing for retrieval." Answer: "Variable-length documents need adaptive chunking: 1. ANALYZE document length - If < 2000 words: Use as single chunk - If 2000-10000 words: Split into 3-5 chunks - If 10000 words: Use hierarchical chunking 2. USE overlap intelligently - Short docs: 10% overlap minimal redundancy - Long docs: 20% overlap ensure continuity - Large docs: 25% overlap bridge multiple chunks 3. IMPLEMENT dynamic sizing chunks = if doc length < 2000: chunks = entire doc else: base size = doc length / num sections 0.8 chunks = recursive split doc, size=base size This ensures consistency while respecting document structure." Answer: "Mixed content requires special handling: STRATEGY:1. DETECT content type if is table: chunk size = 1000 chars tables are dense elif is code: chunk size = 500 chars code needs precision else: chunk size = 2000 chars normal text 2. KEEP units intact - Table: Never split a row across chunks - Code: Never split a function/class - Text: Split at paragraph boundary 3. ADD metadata Each chunk stores: - content type: 'text', 'table', 'code' - importance: 'high', 'normal', 'low' - structure: original section/subsection 4. USE in retrieval Query for code → prioritize code chunks Query for stats → prioritize table chunks This ensures retrieval quality across mixed content." Answer: "Production data often differs from dev data. I'd implement: 1. MONITORING in production - Track precision/recall on real queries - Monitor hallucination rate - Watch for queries that fail - Collect user feedback 2. A/B TESTING for changes Before changing chunk size, run experiment: - Control: Current chunk size 95% traffic - Test: New chunk size 5% traffic - Measure: Impact on user satisfaction, metrics 3. GRADUAL ROLLOUT - Week 1: Test new size with 5% queries - Week 2: Expand to 20% if metrics good - Week 3: Full rollout or rollback 4. FALLBACK PLAN If new size performs worse: - Immediately revert to previous size - Investigate issue - Re-test with adjusted size This ensures production stability while allowing optimization." Print this before your interview Interviewers love when you reference a framework. CHUNK SIZE DECISION MATRIX========================== SMALL 400-600w | MEDIUM 800-1200w | LARGE 1500-2500w Precision/Recall Trade | High Precision | Balanced | High RecallLatency | Fast <50ms | Moderate 50-80ms | Slow 80ms Cost | Low | Medium | HighSemantic Preservation | At Risk | Good | ExcellentEmbedding Model Load | Low | Medium | HighLLM Context Budget | Large margin | Comfortable | TightBest For | E-commerce | Support/FAQ | Legal/Medical | News | Tech docs | Research YOUR USE CASE:- Domain: - Priority: Speed? Cost? Accuracy?- Document length: Short / Medium / Long- Precision vs Recall: Which matters more? RECOMMENDED SIZE: words When asked “How do you decide chunk size?”, end with this: "In summary, it's a three-step process: 1. ANALYZE constraints - Document type legal vs news vs code - Precision vs recall needs - Latency and cost budgets - Model capabilities 2. EMPIRICAL testing - Test 3-5 different sizes - Measure precision, recall, latency, cost - Evaluate hallucination rate - Select size that balances trade-offs 3. PRODUCTION monitoring - Track real-world metrics - Run A/B tests before changes - Implement fallback plan - Iterate based on user feedback The key insight: There's no universal answer. It's a design decision based on YOUR specific constraints and priorities. For this project, I would recommend size words because specific reasons related to their domain ." “I’ve researched this topic through:” How to Decide Chunk Size in Any Project: Complete Interview Guide https://pub.towardsai.net/how-to-decide-chunk-size-in-any-project-complete-interview-guide-81558cb49052 was originally published in Towards AI https://pub.towardsai.net on Medium, where people are continuing the conversation by highlighting and responding to this story.