perf(signals): index duplicate prompt comparisons - #1425
Merged
wesm merged 3 commits intoAug 16, 2026
Conversation
roborev: Combined Review (
|
Member
|
LGTM |
wesm
added a commit
to salmonumbrella/agentsview
that referenced
this pull request
Aug 16, 2026
…dy-core * origin/main: (25 commits) ci: retry failed Docker image builds (kenn-io#1437) Speed up HTTP sync by processing only changed sessions (kenn-io#1414) chore: remove dead code and unused frontend exports (kenn-io#1434) fix(parser): populate Kimi session cwd (kenn-io#1427) feat(frontend): add raw and formatted tool output display (kenn-io#1424) feat(insights): support OpenAI-compatible endpoints (kenn-io#1430) fix(parser): support legacy Zed thread schemas (kenn-io#1429) perf(signals): index duplicate prompt comparisons (kenn-io#1425) fix(config): honor explicit empty agent directory arrays (kenn-io#1423) feat(parser): add DeepSeek Harness session support (kenn-io#1402) test(sync): wait for archive audit stall state (kenn-io#1422) chore(deps): update github actions dependencies (kenn-io#1421) Tombstone stored Claude sessions a complete full parse no longer emits (kenn-io#1392) fix(sync): report stalled syncs as unhealthy (kenn-io#1419) Reduce idle CPU and let users turn off unused providers (kenn-io#1374) fix(serve): survive dual-stack port collisions at startup (kenn-io#1406) fix(sync): bound startup reconciliation and expose daemon identity (kenn-io#1413) fix(deps): update module golang.org/x/mod to v0.40.0 [security] (kenn-io#1400) Persist Claude and Codex freshness digests across engine restarts (kenn-io#1393) fix(sync): add lifecycle logging (kenn-io#1399) ...
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Duplicate-prompt scoring rebuilt a token set for every comparison against prior substantive prompts. On long sessions, signal recomputation therefore spent most of its time repeating token hashing and garbage collection even when the prompts were distinct.
This change preserves the existing exact and fuzzy scoring semantics while indexing accepted prompts by normalized text and token frequency. Each prompt now evaluates only prior representatives that share tokens. A deterministic reference test compares the indexed implementation with the original pairwise algorithm, and a large-session benchmark is added to the existing PR performance gate. On the same 800-prompt shared-vocabulary fixture, current main takes about 1.42 s and allocates 2.14 GB per operation; this branch takes about 13–19 ms and allocates 16.9 MB.
The postings index grows with retained prompt tokens. A pathological session where most distinct prompts share most of their vocabulary can still require many candidate comparisons, but the implementation no longer rebuilds or scans complete token structures for every prompt pair. The main review point is equivalence with the existing asymmetric score, where current tokens are treated as a set and previous tokens remain a multiset.