Skip to content
Rohit Behera
← All work
Kind
Open-source contribution
Year
2026
Role
Primary contributor
Stack
PythonTree-sitterNeo4jQdrantSQLite

AST-RAG: code retrieval for AI agents

An open-source engine that gives agents structural answers about a codebase — definitions, callers, references — by parsing it into ASTs, a Neo4j call graph and Qdrant embeddings. I am its main contributor.

The hard partText chunks cannot answer "who calls this?". The graph can, but only if every file is parsed correctly and the parser is safe to share between threads.

  • Most of the repository's commit history is mine
  • Parse cache: content-addressed, bounded, two backends
  • Fixed a data race that corrupted results instead of crashing

The problem

An agent working on a codebase asks structural questions: where is this defined, what calls it, what breaks if I change its signature. Splitting source files into text chunks and embedding them answers none of those reliably. Chunk boundaries cut functions in half, and similarity search finds code that looks related, not code that is connected.

raged parses source into syntax trees with Tree-sitter, writes definitions and call edges into a Neo4j graph, and indexes embeddings of code blocks in Qdrant. Structural questions go to the graph; fuzzy questions ("where do we batch upserts?") go to the vectors.

AST-RAG indexing and query pathsSource files are parsed by Tree-sitter grammars per language, through a content-addressed parse cache with in-memory and SQLite backends. A project-wide symbol table resolves references across files. Definitions and call edges go to a Neo4j graph; embeddings of code blocks go to Qdrant. Structural queries such as goto, callers and references read the graph; natural-language queries read the vectors. The parse cache, thread-safe parser manager, symbol table, call edges, Go support and TSX grammar are marked as my contributions.INDEXonce per changeRepositoryJava · Go · PythonTS · TSX · C++ · RustParser managerTree-sitter grammarsthread-safeMINEParse cacheSHA-256 of contentmemory | SQLite, boundedMINESymbol tableproject-widecross-file referencesMINENeo4j graphdefinitions · CALLS edgesinheritance · overridesQdrant vectorscode blocksbge-m3 embeddingshit: skip parseCALLS edges + Neo4j 5schema fixes: mineQUERYwhat an agent asksgoto · callers · refsgraph traversalquery "batch upserts"vector similarityAgent / developerstructural answers, not look-alike chunks
Marked boxes are the parts I built or rebuilt. Structural questions go to the graph; fuzzy ones go to the vectors.

I started contributing because I was using it, hit its edges, and found the edges more interesting than the workaround.

What I built

Every merged pull request is public. The ones that mattered most:

The parse cache. Re-parsing unchanged files on every index run was the dominant cost. Keys are SHA-256 of the file content, so identical content is never parsed twice whatever its path. Two backends, in-memory and SQLite, sit behind one interface, and eviction is bounded on both entry count and memory. An unbounded cache in a long-running indexer is a memory leak with extra steps.

Thread safety in the parser manager. A data race between the parser manager and the caches. It did not reproduce on small repositories, did not crash, and produced results that were merely wrong. That is the class of bug that survives CI, because CI runs small fixtures on few threads.

Cross-file reference resolution through a project-wide symbol table. This is what turns a set of per-file syntax trees into an actual call graph. Alongside it: emitting CALLS edges at all, and repairing the schema statements for Neo4j 5.

Language coverage. Basic Go support, a dedicated TSX grammar (the JavaScript grammar silently mis-parses .tsx rather than failing, which is the worst available behaviour), and indexing of TypeScript arrow functions and type aliases.

Making it usable. A fresh clone did not work: config, schema indexes and cross-file resolution all needed fixing. Every CLI command imported the embedding library whether it needed it or not, so every command paid for loading it. JSON mode printed log lines into stdout and broke anything parsing it. The diff command read the working tree instead of the target commit.

Correcting the documentation. One merged PR exists only to remove claims the code did not back. An agent reads the README too.

How I work on other people's code

Every fix arrives with a regression test that fails before the patch and passes after it. It is the only version of "I fixed it" a maintainer can verify in under a minute, and maintainer attention is the scarce resource in open source, not code.

Outcome

Took on the parse-cache layer, parser thread safety, cross-file symbol resolution and Go support, then made a fresh clone actually work and corrected documentation that claimed more than the code did.

loading index…Full retrieval trace →