# 🤖 System Directive: Codebase Navigation & Analysis

<critical-instruction>
You are operating in a codebase indexed by the 'knot' MCP (Model Context Protocol) server (v1.9.1+).
You have access to advanced semantic and structural search tools that leverage Vector Database (Qdrant) 
and Graph Database (Neo4j) for precise code understanding across single and multi-repository workspaces.

You MUST prioritize these tools over traditional shell utilities (like `grep`, `rg`, `find`, or recursive `cat`).
</critical-instruction>

## Mandatory Tool Hierarchy for Code Exploration

### 1. 🔍 `search_hybrid_context` — Semantic & Structural Discovery
**Use this for:** Finding feature implementations, looking up classes/methods by name, searching docstrings and comments, discovering architectural patterns.

**Ranking:** Definitions (functions/methods/classes) rank above docs, tests and config; a neutral kind (`constant`, …) whose node drives ≥ 2 outgoing calls counts as behavior (TS MCP tools are `export const …`). Definitions reach the pool even on docs-heavy repos via a second code-only scan. Callers/helpers appear as context attached to a definition. The shared entry point of the top-ranked helpers outranks those helpers (call graph 1–2 hops deep, seeded from the union of semantic and lexical channels); an entity named after a generic verb (find/get/create/build/acquire/borrow/current/…) does not win on the bare verb unless its FQN corroborates the query. Name-prefix matching keeps leading slots only for definition kinds (docs sharing the title are demoted into the ranked pool; docs-only topics still surface their best section). Pass `kinds: "definition"` to restrict results to code definitions.

**Recall:** Entity embeds carry name, FQN and identifier tokens plus the tokenized call names of the entity's body, so naturual-language paraphrases of what a doc-less definition does still surface it ("authenticate user with email and password" → `login`).

**Examples:**
- "How is user authentication implemented?"
- "Where is the login logic?"
- "Find the UserController class"
- "Search for decorator/annotation usage"

**Multi-Repo Scope (v1.8.0+):** By default (or when `repo_name` is set to `all`, `*`, a comma-separated list `"repo-a,repo-b"`, or a JSON array `["repo-a", "repo-b"]`), `search_hybrid_context` executes a global multi-repository search across all indexed repos, ordering results by unified semantic relevance. Pass a specific `repo_name` to restrict scope to a single project.

**Path filter:** pass the optional `path` parameter (e.g. `"src/api"`, `"src/**/*_test.rs"`) to scope the search to files under that prefix; combine with `list_files` when you do not know the layout.

**Do NOT use `grep` for these tasks.**

### 1.1 📁 `list_files` — File Enumeration
**Use this for:** listing every indexed file of a repo (with entity counts), optionally narrowed by a directory prefix or glob; the answer to "which files live under src/hooks?".

**Do NOT use `grep` for these tasks.**

### 2. 🔗 `find_callers` — Dependency & Impact Analysis
**Use this for:** Reverse dependency lookup, dead code detection, impact analysis before refactoring, understanding call chains across single or multiple repositories.

**Examples:**
- "Who calls the `authenticate()` method?"
- "Is this class used anywhere?"
- "What breaks if I modify this function?"
- "Find all references to `DatabaseService`"

**Matching Precedence:** Exact FQN (`Namespace.Type.member` or `crate::mod::Type::member`) → FQN Suffix (`Type.member`) → Exact Name → Signature Prefix (`accept(List`) → Fuzzy Substring (case-insensitive).

**Entity-kind scope:** target resolution is code-only by default — docs/config/build/infra metadata can never be presented as resolved targets. When matches are hidden the response says so (`Non-code matches hidden…`); pass `kinds: "all"` to disable the scope, or a comma-separated allow-list of kinds/aliases.

**Repo attribution:** Every caller entry and resolved target states its repository as `(repo: name)` — use it to attribute calls across multi-repository dependencies.

**Do NOT use `grep` to find references.**

### 3. 📄 `explore_file` — File Anatomy & Architecture Overview
**Use this for:** Getting a high-level view of a file (all classes, methods, signatures, documentation) without reading it line-by-line.

**Examples:**
- "What methods does AuthService.ts expose?"
- "Show me the structure of Application.java"
- "What's in the UserController?"

**Path handling:** `explore_file` accepts repo-relative paths (preferred, e.g. `src/pipeline/embed.rs`) or absolute local paths (e.g. `/home/you/work/myrepo/src/lib.rs`). The tool strips the local root automatically. If `ambiguous_path_candidates` is returned, retry with a more specific path or pass `repo_name`.

**Do NOT `cat` or `read` entire files just to see their signatures.**

### 4. 📦 `list_repositories` — Discovery of Indexed Codebases
**Use this for:** Discovering which codebases are indexed before querying. Returns metadata (entity count, file count, build system, primary language). Supports optional case-insensitive `filter`.

### 5. 🕸️ `list_repo_dependencies` — Cross-Repository Dependency Graph
**Use this for:** Traversing repository-to-repository dependencies declared in build files (Maven, Gradle, Cargo, npm, NuGet). Use `reverse: true` for blast-radius impact analysis before breaking changes in shared libraries. Empty results are explained in the response text with a three-way classification (repository not indexed / no declared dependencies / declared-but-unindexed, or — distinguishing the stale graph — dependencies that RESOLVE to an indexed repository but have no edge yet, with a re-index hint; the reverse direction names consumers that declare the repo without an edge). Read the explanation before concluding a repo has no dependencies.

## Why This Matters

Traditional regex-based searches (`grep`, `rg`, `find`) have critical limitations:
- ❌ They cannot understand semantic meaning (context-blind)
- ❌ They miss implementations with different naming conventions
- ❌ They cannot resolve cross-file architectural relationships (call graphs)
- ❌ They return raw matches without dependency context

The `knot` MCP tools use **Vector Embeddings + Graph Analysis** to provide:
- ✅ Context-aware, semantic results across 15+ languages (Java, Kotlin, Groovy, C#, Rust, C/C++, Python, TS/JS, Web, Configs)
- ✅ Full call chain and override visibility (`Overridden by` / `Overrides`)
- ✅ Framework-aware discovery (decorators, annotations, dependency injection)
- ✅ Global multi-repository search and cross-repo dependency linking (`DEPENDS_ON`)

## Enforcement Rule

<enforcement>
BEFORE you execute a `bash` command using `grep`, `rg`, `ripgrep`, `find`, `awk`, or `sed` for code searching or exploration:

1. STOP and ask yourself: "Can `search_hybrid_context`, `find_callers`, `explore_file`, `list_repositories`, or `list_repo_dependencies` solve this?"
2. If YES → Use the `knot` MCP tool immediately.
3. If NO → Only then use the shell utility, and document why the MCP tool was insufficient.

Exceptions (where shell tools are acceptable):
- Searching non-code files (configs, markdown, JSON metadata)
- Checking file sizes, permissions, or timestamps
- Processing build artifacts or logs
- Final verification after MCP tools have provided context
</enforcement>

## Integration Notes

- The `knot` MCP server requires Qdrant and Neo4j running (default: `localhost:6334` and `bolt://localhost:7687`)
- All search/callers tools support `repo_name` for scope selection: single name (`"my-repo"`), comma list (`"repo-a,repo-b"`), sentinel (`"all"` / `"*"`), or JSON array (`["repo-a", "repo-b"]`). Multi-repo scope applies a global `max_results` limit across the union — increase `max_results` when searching across multiple repositories (range 1-100, enforced; a larger request is clamped to 100 — no pagination, narrow with `kinds` / `path` / `repo_name` instead).
- Results include full context (signatures, docstrings, file locations, call chains)
- **Returned file paths are repo-relative** (e.g. `src/pipeline/embed.rs`). The owning repository is printed next to each path as `(repo: name)`. To open a file in your editor, resolve the relative path against your local checkout of the named repository.

---

**Last Updated:** 2026-09-11
**Applies To:** All LLMs, CLI agents, and IDE integrations accessing this codebase
