Skip to content

The cognitive loop: how agent-ways works

This document is a walk-through of the agent-ways cognitive architecture. It assumes you know what Claude Code is and nothing beyond that. It is the document to read when you want to understand how the pieces fit together — ways, progressive disclosure, what persists across sessions, and the awareness layer — without diving into the individual ADRs.

If you want to decide a specific tradeoff, read an ADR. If you want the theoretical framing, read the cognitive loop and awareness layer design note. If you want to build a way, read the hooks-and-ways guide. This document sits one level above all of those: it tells the story of how the system composes.

The problem this system addresses

Long collaborations with Claude degrade silently. Not because Claude forgets — it doesn't have memory to forget — but because context fills, attention drifts, and the model that worked at the start of a session stops working later without anyone noticing. The specific failure modes:

  • Guidance dilution. You told Claude important things early in the session, but forty turns later they're buried under task-specific details.
  • Context pressure. Compaction happens at the worst possible moment, in the middle of a thought.
  • Cross-session amnesia. Lessons learned yesterday are invisible today.
  • Interaction friction. You have to re-brief Claude on context Claude could have observed for itself.
  • Silent drift. Claude works on something tangentially related to the goal without anyone noticing until the drift is large.

agent-ways is the composition of mechanisms that address these failures together. No single mechanism solves any of them; the system works because the pieces cover each other's blind spots.

The governing principle: substrate separation

The most important thing to understand is that agent-ways treats Claude's reasoning capacity as the expensive substrate, and everything else as cheaper substrates. Deterministic computation runs in shell scripts, compiled binaries, file I/O, and tool invocations. Reasoning runs in inference, and one small, bounded inference runs outside Claude's context window. The design principle is:

Do the cheap work in a cheap substrate so the expensive substrate can think about things that matter.

This pattern runs through every layer of the system:

  • Way matching is a compiled Rust binary doing embedding math. Claude never decides which ways are relevant. The matcher proposes candidates before Claude sees anything, and on the prompt lanes a small hosted model (the relevance judge, ADR-196) answers yes or no for each one. The judge runs in the ways agent, a per-user daemon, never in Claude's context, and its cost is capped in candidates, input and time. On any failure the matcher's decision stands, and gate.mode: off removes it. The relevance judge gives the limits.
  • Event logging is the ways binary appending one JSON line per fire or near-miss to a file, with no inference involved. Claude doesn't record anything manually.
  • Sensor observation (the awareness layer) is a background script emitting stdout lines when state transitions are worth surfacing. The heavy lifting of turning raw events into discrete observations happens entirely below Claude's token budget.

Most of what happens in an agent-ways session is happening in cheap substrates. Claude only pays tokens for what requires reasoning, and the cheap substrates prepare the ground so that reasoning is aimed at real problems instead of housekeeping.

The visual shape of this separation:

flowchart LR
    classDef cheap fill:#2E7D32,stroke:#1B5E20,color:#fff
    classDef bounded fill:#E65100,stroke:#BF360C,color:#fff
    classDef expensive fill:#C62828,stroke:#B71C1C,color:#fff
    classDef delivery fill:#6A1B9A,stroke:#4A148C,color:#fff

    subgraph CheapSub["<b>cheap substrate</b> — shell, Rust, file I/O — zero reasoning cost"]
        direction LR
        Sensors["sensors<br/>(file, git, context, peers)"]:::cheap
        Attend["attend<br/>(salience, insistence, state)"]:::cheap
        Matcher["ways matcher<br/>(embedding)"]:::cheap
        Gate["refire gate<br/>(ADR-126)"]:::cheap
        Log["event log<br/>(events.jsonl)"]:::cheap
    end

    subgraph Bounded["<b>bounded inference</b> — outside Claude's context"]
        Judge["ways agent → relevance judge<br/>(capped, fails open)"]:::bounded
    end

    subgraph Del["delivery"]
        direction LR
        Monitor["Monitor<br/>(async notifications)"]:::delivery
        Drain["Stop-hook inbox drain<br/>(turn boundary)"]:::delivery
        Hooks["hooks<br/>(sync injections)"]:::delivery
    end

    subgraph Exp["<b>expensive substrate</b> — inference, per-turn cost"]
        Claude["Claude<br/>reasoning + action"]:::expensive
    end

    Sensors --> Attend
    Attend --> Monitor
    Attend --> Drain
    Matcher --> Judge
    Judge --> Gate
    Gate --> Hooks

    Monitor --> Claude
    Drain --> Claude
    Hooks --> Claude

    Gate -->|"fires, near-misses, verdicts"| Log
    Claude -->|"invokes ways show"| Gate

Almost every box on the left runs in a substrate that costs nothing to operate. The judge costs a small, capped amount per prompt and never touches Claude's context. Claude occupies one box on the right. That ratio is the whole design.

Ways for steering

Ways are the reactive guidance layer. A way is a markdown file with YAML frontmatter and a prose body. The frontmatter declares when the way should fire (on a user prompt pattern, on a tool call, on a session event, on a context threshold), and the body is the guidance Claude reads when the way fires.

Ways are triggered by hook events that Claude Code emits on its own loop: SessionStart (including after compaction), UserPromptSubmit, PreToolUse, PostToolUse, PostToolUseFailure, SubagentStart and Stop. A hook fires, a shell script runs ways hook <event>, the binary decides which ways match the current context, and the matched way bodies are injected as additional context. Session conditions such as context usage are state triggers, evaluated by the state hook on session start and on every prompt.

Ways are premises, not rules. They give Claude reasoning to work from, not instructions to follow. The distinction matters because rules scale poorly — a rule that says "always X" becomes a rule you have to remember and a rule Claude has to apply even when X doesn't make sense. A premise says "here's what you should know to reason well about this situation," and Claude's reasoning does the rest. Premises compose because reasoning composes; rules don't, because rules collide.

A typical way looks like:

---
description: git commit messages, branch naming, conventional commits, atomic changes
vocabulary: commit message branch conventional feat fix refactor scope atomic squash amend stash rebase cherry
pattern: push.{0,30}(remote|origin|upstream)
commands: git\ commit
refire: 0.15
scope: agent, subagent
---
# Git Commits Way

## Conventional Commit Format
...

This is the frontmatter of the shipped softwaredev/delivery/commits way. description and vocabulary drive semantic matching, pattern matches prompt text, commands matches the Bash command, refire sets how much of the context window must pass before it can re-disclose, and scope says which agents get it.

When Claude is about to run git commit, the PreToolUse hook fires, the matcher finds this way, and the body is injected into Claude's context right before the tool call. Claude reads the premises, writes a good message, runs the commit. No one had to encode rules; the way provided the reasoning.

See hooks-and-ways/rationale.md for the why, hooks-and-ways/matching.md for how triggers are chosen, and architecture.md for the visual documentation of the hook flow.

Progressive disclosure: why you don't dump

The naive approach would be: at session start, inject every way that's possibly relevant. That fails for three reasons.

  1. Token budget. Hundreds of ways times hundreds of tokens each would consume most of the context window before any work started.
  2. Attention dilution. Claude reads what's nearest the current conversation most carefully. Guidance injected at startup becomes guidance buried under forty turns of later content. See hooks-and-ways/context-decay.md for the formal model.
  3. Habituation. If every way fires on every possible trigger, Claude's context becomes a soup of guidance that doesn't map to what's happening right now.

Progressive disclosure (ADR-105) and window-relative re-disclosure (ADR-126) address this together. The rules:

  • Ways only fire when their triggers match the current situation, never speculatively
  • Once a way has fired, it is marked as "disclosed" and will not fire again until its re-disclosure cooldown expires
  • The cooldown is measured in tokens of context consumed since the disclosure, as a fraction of the model's context window — not wall-clock, not turns, tokens
  • Each agent keeps its own cooldowns: a subagent's disclosures do not silence the main agent's, or the reverse
  • When the cooldown has passed and the trigger fires again, the way re-surfaces fresh

This is habituation in the biological sense. The first mention of something is fully attended to. Repeated mentions become background. If the signal disappears for long enough and returns, it becomes fresh again. Claude's attention budget is managed by the cheap substrate (the refire gate) so that Claude's reasoning gets each signal at the right level of prominence.

The mental model: ways are a rate-limited stream of premises. Claude does not read them all at once; Claude reads them as the situation calls for them, with the understanding that anything already disclosed is still in context unless compaction has happened.

Empirical tuning: letting telemetry revise the thresholds

The thresholds, refire fractions and vocabularies above were set by authorial judgment. ADR-134 feeds the matcher's own record back into them: fires, near-misses and the judge's verdicts accumulate in $XDG_STATE/agent-ways/events.jsonl, and ways tune precision reports which ways keep firing into the wrong kind of session. Vocabulary is never auto-applied, and the threshold auto-tune is deferred. The loop is described in hooks-and-ways.md and the fire rule in engine-reference.md.

Memory across sessions: the event log and repo artifacts

Ways handle the current session. Two things persist across sessions: the event log, which records what fired, and the repository's own artifacts, which record what was understood.

The event log ($XDG_STATE/agent-ways/events.jsonl) is the durable record of ways activity. A SessionStart hook appends a session_start line when a session begins, the ways binary appends a line at each fire, near-miss, suppression and judge verdict, and every writer and reader resolves the file through one path (ways events-log-path, ADR-153). It survives compaction and the end of a session. ways session reads it together with the Claude Code session transcripts to answer which ways fired on which turn and why: list enumerates sessions, replay steps through one on screen (or writes its timeline as JSON with --json), live follows the current session as ways fire, dump writes its reconstruction as JSON for an agent, and fires lists its fires with their scores, lowest first (ADR-154). The log records activity, not content: it does not hold what Claude reasoned about.

What was understood persists in the repository, not in a session-side store. Decisions go into ADRs, working knowledge into ways, open work into GitHub issues, and change history into commit messages and PR descriptions. Claude Code's auto-memory (MEMORY.md) loads at every session start, so ways init seeds it with routing guidance that sends project knowledge to those artifacts and keeps memory for short cross-project facts about the user (ADR-128). Those artifacts travel with the repository, pass review and lint, and are read by teammates and CI as well as by later sessions.

Before compaction, the awareness layer prompts the capture. At 90% context, attend's context sensor points Claude at ways show attend context-pressure, which lists four steps for before the window closes: capture decisions and the approaches tried, update tasks, save to memory anything useful across conversations, and commit work. Memory there is read through the routing above: project knowledge goes to repo artifacts.

The forgetting principle is the counterpart. Most of what passes through a session is not worth keeping. What Claude writes into a repo artifact persists; everything else ends with the session. The event log keeps only the record of what fired, not the reasoning around it.

The same principle bounds the raw telemetry. The fire/near-miss log ($XDG_STATE/agent-ways/events.jsonl, the input to the tuning loop above) is append-only, so it needs a ceiling: when it exceeds ~32 MiB, log_event tail-compacts it to the most recent ~24 MiB at a line boundary via an atomic temp-and-rename. The rewrite is rare, lossy on the oldest events only, and invisible to readers — tuning works from recent behavior, so the old tail is the part safe to forget.

Active perception: the awareness layer

The pieces described so far are reactive: they respond to things Claude is doing. Ways fire on hook events. The event log is written at each fire. Memory and repo artifacts are read at session start. All of this happens on Claude's own timeline — tied to events inside Claude's loop.

But some things happen outside Claude's loop. A background build finishes. A peer Claude Code session modifies a file Claude is editing. Context pressure approaches a critical threshold five turns from now. These are events Claude cannot observe without burning reasoning tokens to check, and that the hook system cannot surface because they do not correspond to Claude's own actions.

The awareness layer (ADR-113, ADR-114) closes this gap. It has two components:

  1. attend — a background Rust binary that observes Claude's session state and environment via small sensor scripts, tracks approaching mechanical consequences using turn-based arithmetic, and emits single-line observations when something is worth surfacing.
  2. Monitor — Claude Code's async-notification tool that delivers background-script stdout as notifications in Claude's chat.

Claude invokes Monitor with attend as the command at session start. attend runs for the session's lifetime, writing observations to stdout as they become worth emitting. Each line becomes a notification Claude reads asynchronously between turns. When the session ends, Monitor terminates attend, which flushes state to disk so the next session can restore it.

Peer messages have a second path. While Claude is working, a Stop hook drains pending messages at the end of each turn (attend inbox --drain, ADR-172). While Claude is idle, the Monitor poller wakes it. Both record what was delivered in one shared set, so a message arrives once.

The awareness layer honors substrate separation rigorously:

  • Sensors operate below the token layer. A file-compare sensor hashes current state, compares to prior, and emits only if the state changed. 99.9% of the time it emits nothing. Only transitions become tokens.
  • Salience is computed in shell. The insistence engine is turn-delta arithmetic: given (disclosed_at_turn, current_turn, context_growth_rate, critical_threshold), compute the projected critical turn. No reasoning required.
  • Emissions are informational, not emotional. The format is declarative: "disclosed at turn 47, currently turn 52, projected critical at turn 58." No simulated urgency, no arbitrary escalation — just honest communication of stakes and timing.

Two delivery paths compose cleanly:

  • Monitor notification only — default for most observations. Claude reads the one-line note, integrates it, acts or dismisses. No ways involvement.
  • Monitor notification + affordance → ways show attend/<signal> — for high-salience observations. attend formats the notification with an explicit ways show command Claude can invoke if deeper guidance is warranted. If Claude invokes it, the ways system shows that way's body.

Claude retains agency at every step. attend suggests; Claude decides. The awareness layer informs, never overrides.

The signal flow through the awareness layer looks like this:

sequenceDiagram
    participant S as sensor
    participant A as attend
    participant M as Monitor
    participant C as Claude
    participant W as ways system

    Note over S,W: turn N — attend running, Claude in a conversation

    S->>A: state transition detected
    A->>A: compute salience, check consequence model

    alt below threshold
        A->>A: hold in deferred intent store (silent)
        Note right of A: most observations end here
    else peripheral (informational)
        A->>M: stdout: short declarative line
        M-->>C: notification delivered async
        C->>C: read, integrate, continue
        Note right of C: no tokens spent on<br/>ways invocation
    else high-salience (insistent or critical)
        A->>M: stdout: line + affordance
        M-->>C: notification delivered async
        C->>C: read, recognize stakes, decide to engage
        C->>W: ways show attend/<signal>
        W->>W: render the way, record the fire
        W-->>C: way body injected
        C->>C: integrate guidance, act
    end

The three alternative branches — silent, peripheral, high-salience-with-affordance — are what makes the awareness layer honest about its cost. The overwhelming majority of sensor events fall into the first branch and cost nothing. Some become one-line peripheral notifications. Only a small minority of genuinely important events pay the full cost of a ways invocation.

The composed loop

Put the pieces together and you get a full cognitive loop running at turn cadence:

flowchart TB
    classDef cheap fill:#2E7D32,stroke:#1B5E20,color:#fff
    classDef expensive fill:#C62828,stroke:#B71C1C,color:#fff
    classDef mixed fill:#6A1B9A,stroke:#4A148C,color:#fff

    Start((session<br/>start))

    W["<b>Wake</b><br/>core ways injected<br/>memory seed checked<br/>attend state restored"]:::cheap
    P["<b>Perception</b><br/>sensors observe<br/>environment + self"]:::cheap
    D["<b>Delivery</b><br/>Monitor delivers<br/>async notifications<br/>hooks deliver<br/>sync injections"]:::cheap
    At["<b>Attention</b><br/>relevance judge<br/>refire gate<br/>salience scoring"]:::mixed
    R["<b>Reasoning</b><br/>Claude integrates<br/>observations +<br/>guidance"]:::expensive
    Ac["<b>Action</b><br/>tools, edits,<br/>responses"]:::expensive
    Ca["<b>Capture</b><br/>fires logged<br/>decisions recorded<br/>in repo artifacts"]:::cheap
    Co["<b>Consolidation</b><br/>compaction distills<br/>working context<br/>event log persists"]:::mixed

    Start --> W
    W --> P
    P --> D
    D --> At
    At --> R
    R --> Ac
    Ac --> Ca
    Ca --> Co
    Co --> W

    Note1["Turns are the tick.<br/>Cheap substrate runs<br/>continuously.<br/>Claude wakes into a<br/>prepared context each turn."]

    Note1 -.- W

Stage by stage:

  • Wake. A new session begins. SessionStart hooks inject the core ways, and ways init checks the memory seed. attend is invoked via Monitor at session start and restores its prior state from disk. Claude reads the orientation context and begins working.
  • Perception. attend runs its sensors in the background, watching Claude's context state, workspace files, peer sessions, and approaching consequences. Most observations are silent; only state transitions worth surfacing reach stdout.
  • Delivery. Monitor delivers attend's stdout lines as asynchronous notifications, and a Stop hook drains pending peer messages at the turn boundary. Hooks deliver synchronous way injections at event boundaries (UserPromptSubmit, PreToolUse, PostToolUse, etc.). All of them land on Claude's attention surface.
  • Attention. On the prompt lanes the relevance judge drops candidates that do not fit the turn. The refire gate (ADR-126) holds back ways disclosed too recently for this agent. Fresh signals get full weight. Deterministic code and one bounded judge call decide what reaches Claude's reasoning.
  • Reasoning. Claude integrates observations and guidance into its working model and decides what to do.
  • Action. Claude acts — edits files, runs tools, responds to the user.
  • Capture. Every fire and near-miss this turn is logged to the event log. As context fills, the context-pressure guidance prompts Claude to commit work and record decisions in repo artifacts. The event log is cheap telemetry; offline, that record drives the empirical tuning of thresholds (ADR-134) without touching the loop.
  • Consolidation. When context fills, compaction distills the working window. The event log, the repo artifacts, and attend's state survive the pass. The next turn begins with a compressed but coherent working context.
  • Back to Wake. At the next session, the loop restarts with the updated state as its foundation.

The loop is turn-driven, not time-driven. Each Claude turn is a tick. Between turns, the cheap substrate (sensors, scripts, the event log) keeps running. When the next turn arrives, Claude wakes into a richer context than the turn before — not because time passed, but because observations accumulated and were filtered by the cheap substrates into summary form.

Every stage of the loop has an appropriate substrate. Reasoning and Action use Claude's inference. Attention spends one small judge call per prompt, outside Claude's context, capped in candidates, input and time, and switchable off. Everything else runs in deterministic code: shell scripts, a Rust binary, file I/O, tool invocations. This is what makes the whole system affordable to run continuously for a full workday — the cost is bounded by what Claude actually reasons about plus a fixed per-prompt judge budget, not by what the system observes.

What this isn't

Worth naming explicitly, because the architecture can be misread if these aren't said out loud:

  • Not consciousness. The substrate is text replay through an inference model. The composition is novel; the substrate is not. agent-ways does not claim or produce sentience. Any language about "presence" or "continuity" in the design note refers to structural properties of the composition, not metaphysical claims about the substrate.
  • Not surveillance. The awareness layer's scope of observation never exceeds the session that owns it. All observations are local. Sensors emit metadata, not content (a presence sensor might emit "user at desk," never a camera frame). The person observed and the person the observations serve are the same person — mirror, not camera.
  • Not required. attend is opt-in. Ways with trigger.type: attend are dormant when attend is not running. The baseline Claude Code experience is unchanged if you do not install the awareness layer. ways itself is additive too — Claude Code works without it. agent-ways is a composition you can opt into at whatever depth makes sense for your workflow.
  • Not C2. Despite superficial resemblance to command-and-control patterns, the architecture points inward, not outward. One session, one user, one machine. No inter-instance protocols. No central servers. ADR-101 and ADR-102 tried outward-facing designs and were abandoned for good reasons; the awareness layer points the other direction.
  • Not automatic guidance injection at the awareness layer. Even at critical salience, attend does not inject ways directly. It suggests affordances; Claude decides whether to invoke them. The "Claude retains agency" invariant is load-bearing.

Where to dig deeper

Ordered roughly by how specific the topic is to your interest:

If you want the same pipeline observed from a running session: - How ways works — what you can watch happen, and how to read the session data

If you want the theoretical framing: - Design note: cognitive loop and awareness layer — reads the system as an active-inference loop and names the invariants the ADRs preserve - hooks-and-ways/rationale.md — the rationale for the ways system - hooks-and-ways/context-decay.md — the attention-decay model underlying progressive disclosure

If you want to understand specific decisions: - ADR-126 — re-disclosure as a fraction of the context window - ADR-160 — late-interaction matching - ADR-196 and ADR-502 — the relevance judge and the ways agent that runs it - ADR-105 — progressive disclosure for way trees - ADR-108 — embedding-based way matching - ADR-128 — memory routing: project knowledge belongs in repo artifacts - ADR-113 — the attend binary - ADR-114 — the way trigger schema for attend signals - ADR-134 — empirical auto-tuning from fire and near-miss telemetry - ADR-153 — the session-introspection substrate over the event log - ADR-154 — ways session

If you want to build ways or operate the system: - hooks-and-ways/README.md — start here for way authoring - hooks-and-ways/matching.md — choosing a trigger strategy - hooks-and-ways/extending.md — adding new ways - hooks-and-ways.md — reference for the hook lifecycle and system mechanics - architecture.md — visual documentation of the ways system's internals (sequence diagrams, state machines, scoring pipeline)

If you are auditing for safety, privacy, or governance: - The "What this isn't" section above - Hard invariants in ADR-113 (session-scoped observation, no C2 topology, informational not enforceable, consequence-anchored, metadata-only for content-bearing sensors, additive never required) - hooks-and-ways/provenance.md — governance traceability

If you are new to the whole thing and want the shortest reading path: 1. This document 2. hooks-and-ways/README.md 3. architecture/practice/ADR-600-cognitive-loop-and-the-awareness-layer.md

Everything else is there when you need it.