How a 570-Line JSON File Saves 85% of AI Token Waste (and the CO2 That Goes With It)

Every time you ask an AI coding assistant to "find where useAppStore is defined," it launches a multi-step expedition: glob for files, grep for patterns, read candidates one by one. For a modest 30-file React project, that's easily 5–10 tool calls and 15,000+ tokens of file content ingested — just to answer a question that has a three-word answer: src/stores/appStore.ts:23.

We built a dead-simple fix: a static JSON index of your codebase that the AI reads once instead of exploring repeatedly. Here's what happened when we iterated it down to the essentials.

The Problem: AI Assistants Explore Like Tourists

AI coding tools (Claude Code, Cursor, Copilot, etc.) don't have a mental model of your project. Every session starts from zero. When they need context, they:

  1. Glob for file patterns (**/*.ts, **/*.tsx)

  2. Grep for symbol names across matches

  3. Read each candidate file to confirm

  4. Repeat for every follow-up question

For our palm reading app (28 source files, React + TypeScript + Zustand), a typical "understand the codebase" exploration consumed ~50,000 input tokens across 15–20 tool calls. That's before the assistant writes a single line of code.

The Fix: One JSON File, Generated Once

We built Codebase Indexer, a lightweight Go tool that uses tree-sitter ASTs to extract three things:

The AI reads this once at conversation start, then knows exactly where everything is.

codebase-indexer -root . -out .context/relationships.json

The Iteration: From 2,970 Lines to 570

The first version indexed everything: every local variable, every function call, every destructured binding. It worked, but it was bloated. We went through five rounds of pruning:

Version Lines What Changed
v1: Everything 2,970 Calls, locals, destructurings, duplicates — all indexed
v2: Remove calls + dedup 2,255 Dropped calls arrays (noisiest, least useful field)
v3: Filter locals 314 Only exported symbols, stripped all local defines
v4: Fix type exports 470 Re-added defines for files with actual exports
v5: Full export detection 570 export const/function now detected alongside types

The key insight at each stage: what does the AI actually need to navigate? Not local variables. Not call expressions. Not destructured const [foo, setFoo] bindings. It needs:

  1. Where is this symbol defined? → symbols

  2. What depends on what? → dependencyGraph

  3. What does this file export? → files[x].exports

Everything else is noise that costs tokens without aiding navigation.

The Numbers

Token savings per session

Scenario Without Index With Index Savings
"Find where X is defined" ~15,000 tokens (5-10 tool calls) ~2,000 tokens (1 read) 87%
"Understand the codebase" ~50,000 tokens (15-20 tool calls) ~4,000 tokens (1 read) 92%
Typical feature implementation ~80,000 tokens exploration overhead ~12,000 tokens 85%
Impact analysis ("what breaks if I change X?") ~25,000 tokens ~2,000 tokens 92%

The index file itself costs ~4,000 tokens to ingest. It pays for itself on the first question.

Across a workday

A developer working with an AI assistant typically runs 20–40 conversations per day. Conservatively:

Monthly (solo developer)

For a team of 10

The CO2 Impact

AI inference has a real carbon footprint. The exact numbers vary by provider, hardware, and energy mix, but published research gives us reasonable estimates.

Per-token emissions

Based on Luccioni et al. (2023) and infrastructure reports from major cloud providers:

Solo developer savings

Team of 10

At scale (1,000 developers)

These are conservative estimates. They don't account for output tokens (which cost more energy), retries from failed explorations, or the compound effect of the AI making better decisions faster with good context.

Why This Works Better Than RAG or Embeddings

You might wonder: why not use vector embeddings or retrieval-augmented generation?

  1. Zero infrastructure. A Go binary and a JSON file. No vector database, no embedding model, no chunking strategy to tune.

  2. Deterministic. The same codebase always produces the same index. No relevance scoring to debug.

  3. Complete. Every export is indexed. Embedding-based retrieval can miss symbols that don't appear in semantically similar contexts.

  4. Fast. The indexer runs in <1 second for most projects. It can run on every commit via a git hook.

  5. Readable. The AI (and you) can read the JSON directly. Try that with a vector database.

How to Use It

1. Install

git clone https://tangled.org/metaend.eth.xyz/codebase-indexer/
cd codebase-indexer
go build -o ~/go/bin/codebase-indexer .

Make sure ~/go/bin is in your PATH.

2. Generate

codebase-indexer -root . -out .context/relationships.json

3. Tell your AI assistant

Add to CLAUDE.md or AGENTS.md:

Before exploring unfamiliar parts of the codebase, consult `.context/relationships.json`

This file contains:
- **symbols**: Global symbol-to-location index (find any function/var instantly)
- **dependencyGraph**: Which files depend on which
- **files**: Per-file exports and defines

4. Auto-update (optional)

# Git hook
cp post-commit .git/hooks/ && chmod +x .git/hooks/post-commit

# Or watch mode
./watch-index.sh

The Takeaway

The most effective optimization isn't a smarter model or a bigger context window. It's not sending tokens you don't need to send.

A 570-line JSON file, generated in under a second, eliminates 85% of the exploration overhead that AI coding assistants burn through on every session. That's real money saved, real latency reduced, and real carbon not emitted.

The best part: it took five iterations to get from "index everything" to "index only what matters." The same principle applies to the index itself — less is more, as long as it's the right less.


Codebase Indexer is open source. Supports JavaScript, TypeScript, Go, and Python via tree-sitter. Adding a new language is ~50 lines of Go.