Repository navigation
Proposal: Context File Growth: Consolidation + Smart Retrieval #19
Description
Activity
Thanks @v0lkan 🙏 .
This directly affects agent effectiveness.
And looks like you've started to feel the pain.
Pain is good; it drives structure.
"As projects mature, the context files will bloat and the agent gets worse at picking what's relevant": %100 agree with that.
imho, Phase 1 alone ( relevance-aware retrieval) would be a significant improvement with no file format changes.
It's a pragmatic take.
Let me dig into the codebase and see what comes out.
- added 2 commits that reference this issue
on Feb 20, 2026 thanks @v0lkan for coming up an idea with this and addressing the issue. i do suffer from the same problem when it comes to use ctx long sessions and/or large codebases. i was thinking what could be done and my reasoning is as follows.
if we reason from first principles about how human knowledge scales versus how LLMs process data, the root cause isn't actually retrieval, and the solution isn't batch consolidation.
The problem is that we are treating our context files as append-only ledgers.
Ledgers scale linearly to infinity; human technical knowledge does not. When you learn three things about React hooks, your brain doesn't store three isolated log entries—it updates your relational mental model of React.
Furthermore, trying to solve context limits by moving knowledge into Vector DBs (Embeddings) or fine-tuned weights (LoRAs) is a trap. Embeddings are opaque and still eat context tokens upon injection. LoRAs are probabilistic—they are great for implicit coding style, but they will confidently hallucinate strict rules.
Here is a first-principles alternative that honors your core constraint (human-readable, flat, git-tracked Markdown) while permanently solving the context bloat, leveraging bleeding-edge architectures like Test-Time Training (TTT-LM) and Semantic Ontologies.
The Root Problem
Files like LEARNINGS.md grow without bound because the agent is "appending." This forces expensive token-truncation and requires brittle batch-consolidation. We need to shift the data structure from a Chronological Ledger to a State-Based Wiki, and upgrade our inference engine to handle infinite sequences natively.
Proposal: The "Living Ontology" & TTT Architecture
1. Write-Time Compression (The Semantic Markdown AST)
Instead of the agent appending "Learning 48" to the bottom of the file, we rigidly structure context files as an Abstract Syntax Tree (AST) using Markdown headers (## Domains -> ### Topics).
When an agent learns a new edge case about WebSockets, it does not append. It uses a new /ctx-update skill to patch the existing ## WebSockets node. Knowledge naturally compresses at write-time. The git diff shows the evolution of a concept, not just a growing list.
To make this robust without databases, we introduce machine-readable (but human-friendly) tags:
Time-To-Live (TTL): > Expires: 2026-12-01. Knowledge decays. ctx drift no longer counts entries; it flags expired TTLs for revalidation.
Validity Tags: ``. If the agent's environment doesn't match, the AST parser drops the node entirely before the prompt is built.
Semantic Edges: *(Derived from: ../LEARNINGS.md#auth-bug)*. We use standard markdown anchors to link a rule in DECISIONS.md directly to the "why" in LEARNINGS.md.
2. Context-as-a-Dependency (CaaD)
Because our context is now a structured tree rather than a flat log, it can be shared across repositories just like an npm package.
Repositories define dependencies in context.yaml (e.g., @acme-corp/react-conventions).
Running ctx install performs an AST Tree Merge, pulling collective knowledge into local files under read-only namespaced headers (## @global/React).
Lexical Shadowing: If a rule under ## 🏠 Local/React contradicts the global registry, local always overrides global.
3. The Engine Layer: Test-Time Training (TTT) & Native Caching
We stop trying to carefully budget 4,000 tokens.
Standard Transformers use a KV Cache, which grows linearly and crashes VRAM. But by adopting models based on Test-Time Training (TTT-LM) (or leveraging advanced native KV caching), the model replaces the KV cache with a hidden state trained via gradient descent during inference.The Result: The memory footprint remains constant regardless of sequence length.
Instead of guessing which 4,000 tokens to push, ctx agent streams the entire parsed Markdown Ontology into the model's test-time weights. The agent processes your entire repository's rules with effectively zero ongoing VRAM bloat per turn.
Alternatives Considered
Approach | Why it's a Trap in this Paradigm -- | -- Embeddings / Vector DBs | Violates the flat-file constraint. RAG is a "guess and push" model that still injects tokens into the context window, blowing up budgets anyway. LoRA / Fine-Tuning | Weights are for implicit style, context is for explicit facts. LoRAs are probabilistic and will hallucinate strict rules. Periodic Consolidation | Requires expensive LLM batch processing. Highly prone to losing technical nuance. Treats the symptom (ledger bloat) instead of the disease.Implementation Phases
Phase 1: Data Structure Shift (Highest ROI, No File Changes)
Restructure existing .md files into strict Header/Topic trees.
Deprecate "append to bottom". Build the /ctx-update skill to patch existing headers.
Modify ctx drift to enforce AST structure and scan for expired TTL tags (> Expires: YYYY-MM-DD).
Phase 2: Context Registries (The CaaD Fleet Update)
Introduce context.yaml for defining organizational dependencies.
Build ctx install to fetch remote Markdown files and perform an AST merge into local files.
Implement Contextual Validity parsing to drop AST nodes based on the local package.json dynamically.
Phase 3: Infinite Engine Migration (TTT / Advanced Caching)
Transition the ctx agent backend to utilize native Prompt Caching or a TTT-LM runner.
Remove the --budget 4000 constraint. Feed the entire parsed AST into the model prefix, relying on the model's fixed memory state to retain the rules indefinitely.
Open Questions
TTT Exact Recall vs. Lossy Compression: While TTT elegantly compresses infinite context into a fixed state, it acts as lossy compression. Can it reliably recall exact API strings (e.g., sk_test_...) buried deep in the ontology, or do we need a hybrid attention/TTT approach for hard constants?
TTL Enforcement: When ctx drift flags an expired learning, should we allow the agent to automatically write a test script to attempt to re-validate it, or must TTL renewals strictly be human-approved?
Governance: Should agents be allowed to automatically propose /ctx-publish PRs to the global company registry when they patch a novel bug into the local ontology?
- added 2 commits that reference this issue
on Feb 27, 2026 - Reacted by ersanReacted by ersan and Pires
I would like to propose one additional improvement to the memory retention flow.
During a session, an LLM naturally tries to satisfy the user’s immediate needs and keep building on the current direction of the conversation. However, a session can sometimes drift into an unhelpful direction, and the model may continue reinforcing that drift simply because it is trying to remain coherent with the ongoing conversation.
In those cases, the issue is not only whether the content should be preserved, but whether it should be treated as useful context at all.
Instead of relying only on explicit user tagging such as “keep” or “discard,” ctx could use implicit feedback from later turns to decide whether a portion of the conversation should become durable context, be summarized lightly, or be excluded from persistent memory.
For example, if the user later corrects the direction, abandons the topic, says that a path was not useful, or returns to a previous framing, that could be treated as a signal that the drifted portion should not be promoted into long term context.
This does not necessarily require a large LLM for every decision. A lightweight retention classifier could run at the end of each session and score conversation segments before anything is written into long term context. The score could combine explicit user instructions, correction signals, topic abandonment, project relevance, stability, and reuse potential.
Segments with high retention scores could be promoted into durable context. Uncertain segments could be summarized lightly or surfaced for user review. Low scoring drift could simply be excluded from persistent memory.
This would complement smart retrieval and consolidation by addressing an earlier stage of the pipeline: not only “what should we retrieve from existing context?” but “what should be allowed to become long term context in the first place?”
This may reduce context bloat without forcing users to manually curate every session, while still preserving the project’s core principle that the final context remains human readable and user controlled.
Excellent suggestions @Ayshine .
I think it's about time to put this into the backlog after cleaning up the current
.context/TASKS.md.It's about time for
ctxto start "dreaming".I may create a more detailed write up and share it in the internal (inverted funnel) GitHub discussion with the usual suspects.
Reacted by Ayşin SancıI have some doubts though.
Most good harnesses limit the
Read()tool. For Claude Code for example, after 3000 lines read tool will reject and the agent will use things likeripgrep.Still an interesting idea; need to think about this a bit more.
- added 2 commits that reference this issue
on Jul 7, 2026
Problem
Files that grow every session —
LEARNINGS.md,DECISIONS.md,CONVENTIONS.md— accumulate entries without bound. After weeks of activeuse, these files become expensive to load into agent context and dilute
signal with entries that aren't relevant to the current task.
The problem isn't storage (flat files are fine on disk). The problem is
retrieval:
ctx agent --budget 4000must decide what to include, andtoday it truncates by token count rather than by relevance.
Constraints
the core design principle that context is human-readable, git-tracked
Markdown.
them harder to find. A learning from 3 months ago might still be the most
critical thing an agent needs today.
LEARNINGS-01.md,LEARNINGS-02.mdloses the single-file simplicity that makes ctx work.Agents would need to know which file to read, and humans lose the ability
to grep one file.
learnings and decisions are often permanent. "Don't use
go installinhooks" is as true on day 100 as day 1.
Proposal: Three-Layer Approach
1. Periodic Consolidation (reduces without burying)
A
/ctx-consolidateskill that:.context/archive/learnings-YYYY-QN.mdfor referenceExample: 5 separate learnings about hook edge cases become 1 consolidated
entry covering all the gotchas, with the 5 originals archived.
Key distinction: consolidation ≠ archival. Archival moves entries out.
Consolidation replaces verbose entries with tighter ones — the file stays
useful, just denser.
2. Relevance-Aware
ctx agentRetrievalInstead of truncating by token count,
ctx agentscores entries:TASKS.mdentriesthan old learnings (conventions are always relevant; learnings are
situational)
one-line summary rather than full inclusion
The files stay flat and append-only. The presentation layer gets smart.
This is the highest-impact change: it solves the problem without touching
the files at all.
3. Soft Caps with Nudges
ctx driftwarns when a file exceeds a threshold:Not enforcement — just a nudge in the existing maintenance workflow. The
threshold is configurable via
.contextrc.Alternatives Considered
Implementation Phases
Phase 1: Smart Retrieval (highest impact, no file changes)
Enhance
ctx agentto score entries against current tasks before includingthem. This alone solves the "too much context" problem without any file
format changes.
internal/cli/agent/to score entriesTASKS.mdfor relevance matching(high weight) → decisions (medium) → learnings (scored)
Phase 2: Drift Nudges (low effort)
Add entry count checks to
ctx drift:.contextrcPhase 3: Consolidation Skill (full solution)
Build
/ctx-consolidateskill:.context/archive/Open Questions
consolidated-from:metadata fieldlinking back to the originals?
files)?
ctx agentbudget allocation weights be configurable, or aresensible defaults sufficient?
Labels
enhancement,context-management,agent-ux