ADR-197: Cap the candidates the relevance gate judges per request¶
Summary¶
- Decided: A judge request carries at most a profile-set number of candidates, 8 by default, taken in the matcher's order. Candidates past the cap are not judged and pass as the matcher decided. Each request that hits the cap is logged with the ways it left unjudged.
- Trades away: On a prompt with more candidates than the cap, the lowest-ranked ways reach the session unfiltered. The alternative, a deadline that grows with candidate count, would judge them all and make those prompts wait.
- One-way? No. The cap is a profile field; raising it to the request's whole candidate list restores the ADR-196 behaviour.
- Probes: Confident (latency): you would rather a few low-ranked ways pass unjudged on a busy prompt than wait past 2 s for a verdict on every one. Not confident (overflow): if the overflow turns out to be mostly noise, would you rather block it than pass it?
- Inversion: Between judging every candidate however long it takes and judging none. The answer sits nearer the first end, with a bound set by the latency the operator accepts.
Context¶
ADR-196 sends every candidate of a prompt in one judge request with a 2 s deadline, and fails open past it. Measured on the deployed gate (ADR-195, second addendum), a call takes about 0.6 s plus 0.1 s per candidate, because the answer grows by about 20 output tokens per candidate and output dominates the call. Every deadline fallback in three benchmark runs was a prompt with 10 or more candidates. Those prompts are the ones with the most ways to filter, and on them the gate filtered nothing.
The operator ruled out spending more latency or cost per prompt for a better gate.
Decision¶
- The cap. A profile field,
max_candidates, bounds the candidates in one judge request. The shipped profiles set 8, which the measurement puts near 1.4 s, inside the 2 s deadline. A user layer may override it like any other profile field. - Which candidates. The gate takes the first
max_candidatesof the ways it would judge, in the matcher's admission order: hits from an explicit trigger (files:,commands:, a keywordpattern:) before scored hits, each tree whole before the next, parents before children, siblings by their best score. Ways withpattern_strictare still not judged and do not count against the cap. - The rest. Candidates past the cap get no verdict and pass, as they did before the gate existed, except a way whose ancestor the judge blocked: a way's guidance presumes its parent's, so it is blocked with the parent. This amends ADR-196 §1 and §2: one request per prompt carries up to the cap rather than every candidate.
- Logging. A request that hits the cap logs one
gate_cappedevent with the number judged, the number left unjudged and their ids.ways introspect dumpcounts unjudged ways in its gate summary beside the verdicts and fallbacks.
Consequences¶
Positive¶
- Prompts with many candidates get verdicts on their strongest matches instead of a fallback that judges none.
- Gate latency stays bounded by the cap, whatever the matcher proposes.
Negative¶
- The lowest-ranked candidates on a busy prompt pass unfiltered. The
gate_cappedevents show how often and which ways. - The cap is a second knob beside the deadline. Raising one without the other brings the fallbacks back.
- The admission order fills the cap with explicit-trigger hits first, so on a prompt with many of them the overflow is semantic hits, the kind the gate filters most. Whether the cap should fill with semantic hits first is open, and the
gate_cappedevents are the data for it.
Neutral¶
max_candidatesis a new profile field inways-agent-core, which both thewayshook and the agent parse withdeny_unknown_fields. Both binaries ship together in the release that adds it. An older agent given a user layer that sets the field rejects the file on every request, and the gate fails open withagent_errorlogged. The cap is applied by the client; the agent only reports it.
Alternatives Considered¶
- A deadline that grows with candidate count, about 0.6 s plus 0.1 s per candidate. It judges every candidate, and a 15-candidate prompt waits about 2.5 s. Rejected by the operator on latency.
- Blocking the overflow. The overflow is the matcher's weakest matches and the gate blocks most candidates it judges, so blocking would often be right. It would also block ways no judge has read, which the gate never does elsewhere. Left open as the second probe; the
gate_cappedevents are what would settle it. - Splitting the candidates across parallel requests. Each request stays fast, but verdicts depend on which candidates share a request (ADR-195, second addendum), so splitting changes the verdicts as well as the latency, and it multiplies calls.
- A smaller answer per candidate. Measured as
p_relevantalone: 3 fewer output tokens per candidate, since the JSON keys and ids dominate the answer. Not enough to move the deadline.