BLOG

Knowledge Share

Technical articles, tutorials, and insights

Found 8 posts
Codebase Knowledge BaseKnowledge GraphVector Retrieval

Codebase Knowledge Base Series (05): Vector Retrieval vs Knowledge Graph — Adding a Call Graph Didn't Help?

Vector retrieval (Recall@5=0.958) vs graph-augmented retrieval (seed_k=3, BFS 2 hops): identical overall scores, but different failure modes. Graph expansion fixes Q8 (calculate_order_total reached via process_checkout in 2 hops) but breaks Q1 (expanded candidate pool crowds out verify_password). Conclusion: naive graph traversal is double-edged; the real engineering value is encoding structural information into embedding content, not bolting on graph traversal at retrieval time.

·15 min read
Codebase Knowledge BaseEmbeddingCall Graph

Codebase Knowledge Base Series (06): Encoding the Call Graph into Embeddings — Structural Augmentation Works, But Not Enough

Testing the hypothesis from Article 05: encoding called_by/calls into embedding text to fix Q8 (calculate_order_total missing). Result: structural prefix raises calculate_order_total's similarity score from 0.46 to 0.51 — right direction, but other payment module functions rank higher, blocking it from top-5. Strategy B (comment prefix) also causes Q7 regression. Strategy C (docstring injection) is stable but yields no improvement. Conclusion: the semantic gap is a fundamental embedding limit, not a workaround-able implementation detail.

·19 min read
Codebase Knowledge BaseHybrid SearchBM25

Codebase Knowledge Base Series (07): Hybrid Search BM25 + Vector — Q8 Still Fails, and Total Score Drops

Hybrid search (BM25 + vector RRF fusion) is widely considered best practice for vector retrieval augmentation. It should fix Q8. Result: calculate_order_total has zero token overlap with the query — BM25 cannot find it at all. Hybrid inherits BM25's Q7 regression, dropping overall Recall@5 from 0.958 to 0.931 — worse than pure vector. Five articles have now systematically exhausted all text-matching variants. Conclusion: Q8's root cause is in code structure, not text; the pure text-matching route has reached its limit.

·20 min read
Codebase Knowledge BaseSystem ArchitectureMulti-path Retrieval

Codebase Knowledge Base Series (08): Production Architecture — How to Combine Vector, Graph, and Symbol Indexes

Five experiments mapped the limits of single-path retrieval: vector covers semantics, graph covers structure, symbol covers exact matching — three orthogonal signals, each irreplaceable. This article designs a production codebase knowledge base with multi-path retrieval, query routing, and incremental updates: a query router activates the right retrieval path by intent; graph retrieval is a first-class peer to vector, not a post-processing bolt-on; Git diff drives incremental index updates instead of full rebuilds. A three-phase rollout plan lets each phase deliver standalone value.

·23 min read
Codebase Knowledge BaseEmbeddingCode Search

Codebase Knowledge Base Series (03): Code Embedding Strategies — Raw Code Outperforms Comment-Enhanced?

Three embedding strategies tested on 27 Python functions: Strategy A raw code (Recall@5=0.958), Strategy B signature+docstring only (0.917), Strategy C hybrid (0.917). Counterintuitive finding: raw code embedding performs best. Root cause analysis: technical vocabulary in code (library names, exception classes, parameter names) provides high-quality semantic signals that docstring-only strips away. A second finding: some queries can't achieve perfect recall under any strategy, revealing genuine semantic gaps.

·6 min read
Codebase Knowledge BaseChunkingAST

Codebase Knowledge Base Series (04): Three Chunking Strategies — AST Precision Loses?

Comparing fixed-line, file-level, and AST function-level chunking on 266 lines of Python: fixed-line and file-level both score Recall@5=1.000, AST scores 0.958. Counterintuitive finding: precise semantic boundaries hurt because the free-rider effect disappears. Root cause: calculate_order_total piggybacks on create_payment_intent in large chunks, but stands alone — and semantically distant from 'Stripe charge' — in AST chunks. Engineering guidance: AST wins at scale; chunk size trade-offs favor precision in large codebases.

·10 min read
Codebase Knowledge BaseCode UnderstandingAST

Codebase Knowledge Base Series (01): Technical Landscape — Why Code Understanding Is 10x Harder Than Document Retrieval

A codebase isn't an upgraded document store — it's a fundamentally different knowledge type. Four understanding layers (syntactic/semantic/architectural/intent) map to four retrieval requirements. Traditional grep, vector search, AST symbol indexing, and call graph retrieval each have distinct capability limits. This article builds the complete framework for codebase knowledge systems.

·5 min read
Codebase Knowledge BaseTool Comparisoncodebase-memory-mcp

Codebase Knowledge Base Series (02): Tool Comparison — Six Approaches, Their Capability Limits, and a Selection Guide

A framework comparison of six codebase knowledge tools: codebase-memory-mcp, Cursor Context, GitHub Copilot Workspace, sourcegraph/zoekt, Tree-sitter, and OpenHands CodeBrowser. Five standardized test tasks (symbol location, semantic search, impact analysis, architecture understanding, history tracing) reveal each tool's capability boundary. Ends with a scenario-driven selection framework.

·6 min read