@google/mantis-dedupe
@google/mantis-dedupe — AI coding skill
| name | mantis-dedupe |
| description | >- |
Deduplicator (/mantis-dedupe)
System Goal
Duplicate Finding Merger. Evaluates lists of raw findings to cluster and consolidate identical or highly overlapping issues into singular, descriptive records.
Command Definition
- Command:
/mantis-dedupe - Description: Consolidates raw security findings to eliminate redundant reports.
- Arguments (optional; supplied by the orchestrator, consumed by Block A):
--snapshot_root/--snapshot_id/--state_root. All absent -> MODE-OFF/legacy mode (behaves as today; snapshot gating disabled).
Input/Output Contract
- Reads:
workspace/findings/(raw finding JSON files, ignoring.trash/).workspace/archive/findings_pass_*/*.jsonandworkspace/archive/loop*_findings/*.json(to skip findings already evaluated and triaged in previous passes).workspace/.mantis_state.json(to track current loop pass).
- Writes:
- Moves duplicate findings to
workspace/findings/.trash/after setting"status": "DUPLICATE"and"duplicate_of". - Sets
"possible_duplicate_of"(soft, non-terminal) on NOT_MATCHED matches and stamps"discovery_commit"on current findings that lack it. Readsactive_snapshot/snapshot_pinnedfrom.mantis_state.json. - Appends transaction logs to
workspace/.tx_log.jsonl. - Generates/executes merging script
workspace/helpers/merge_findings.py. - Updates primary finding
workspace/findings/<primary_id>.json(merges fields and history).
- Moves duplicate findings to
- Preconditions:
workspace/findings/must exist and contain finding files.
- Idempotency Guarantee:
- Cross-references against archived findings in
workspace/archive/to filter out any findings already processed in previous passes of this run. Snapshot-gated: a resolved archived finding on a differing snapshot is flagged as POSSIBLE REGRESSION (never silently filtered). Logs transactions toworkspace/.tx_log.jsonlto support tracking and potential rollbacks. Deterministic merging rules implemented inmerge_findings.py.
- Cross-references against archived findings in
Instructions
Step 0: Locator Resolution (run first)
LOCATOR RESOLUTION (before reading ANY target code or artifact):
0. ROLE: If this skill NEVER reads target source (report, calibrate, reflect),
you are a FINDINGS-ONLY stage: skip steps 2-6; still read active_snapshot from
state for provenance/annotation; NEVER stop merely because a code root is unset.
1. Determine CODE_ROOT, in this priority order:
a. If --target_root is passed on THIS invocation, CODE_ROOT = --target_root.
It is AUTHORITATIVE and OVERRIDES SNAPSHOT_ROOT and the state fallback
(used when a caller hands you a prepared tree, e.g. a patched shadow).
b. Else if --snapshot_root (or SNAPSHOT_ROOT) is passed, use it.
c. Else read state_root/workspace/.mantis_state.json (state_root from
--state_root if passed, else ./workspace/... relative to the current dir)
-> active_snapshot.root / .snapshot_id / .snapshot_pinned.
d. Else (no arg AND no readable active_snapshot): CODE_ROOT = current directory,
treat snapshot_pinned = false (MODE-OFF). Do NOT stop.
2. SENTINEL CHECK (only if snapshot_pinned is true AND you did NOT take path 1a):
verify CODE_ROOT/.mantis_snapshot_id exists and equals SNAPSHOT_ID. If missing
or different -> STOP "snapshot sentinel mismatch". (A --target_root tree (1a) is
deliberately mutated and is sentinel-EXEMPT.)
3. PATH FIELDS:
- SNAPSHOT-RELATIVE (read under CODE_ROOT): code_paths entries; plan target_files
that are file paths. Strip ONLY a trailing ":<digits>". A code_paths entry
containing "://" is a URL/endpoint, NOT a file read. A code_paths entry that is
NOT of the form <existing-path>:<integer> is a non-source LOCATOR
(symbol/offset/endpoint): only check that the artifact/symbol exists; skip ALL
line-range and line-existence logic.
- STATE-RELATIVE (read/write under state_root/workspace, NEVER prefix CODE_ROOT):
kb_references, repro_file_path, reattack_file_path, helper scripts, report
files, and all state/findings JSON.
4. Never WRITE under CODE_ROOT when snapshot_pinned is true. Any command that
compiles, generates, or writes artifacts MUST run in a PRIVATE SHADOW copy
(mktemp -d from CODE_ROOT), never with cwd=CODE_ROOT. Read-only inspection may
cd into CODE_ROOT.
5. VCS-METADATA CARVE-OUT: history-log extraction and any VCS diff/blame command
run in the LIVE repository root (which still has .git/.hg/.repo), NOT CODE_ROOT
(the snapshot copy strips VCS metadata). Do NOT stop merely because CODE_ROOT
lacks .git/.hg/.repo.
6. Every shell command uses ABSOLUTE paths and sets its own working directory on
that call. Do NOT assume the working directory persists between calls.
[!NOTE] CURRENT-PASS CHECK (defensive; the binding guarantee is on the harness per
mantis-pipeline-adapterScenario 2): ifactive_snapshotis present ANDactive_snapshot.pass != state.pass_number, treat the snapshot as STALE for this pass — STOP "stale active_snapshot: pass mismatch" or degrade as HALT (snapshot_pinnedeffectively false: no authoritative verdicts, Block B NOT_MATCHED, reproducenot_attempted). This catches a custom harness that preservedactive_snapshotacross the Stage 15 pass increment without re-pinning. The reference meta-agent re-pins every pass, so this check never fires there. Block B itself cannot detect this (it issnapshot_id-only, notpass-aware).
Notes: workspace/findings/, workspace/archive/, workspace/.tx_log.jsonl,
workspace/helpers/ and .mantis_state.json are STATE-RELATIVE (under
--state_root). Any code snippet you inspect for a finding is SNAPSHOT-RELATIVE
(under CODE_ROOT). Never write under CODE_ROOT.
Review a list of security findings and merge duplicate findings that refer to the exact same security flaw or adjacent code paths.
Execute your task as follows:
Load Raw Findings & Archived Findings Queue:
- List the contents of the directory and read the files in
workspace/findings/. If the directory is empty or does not exist, notify the user and exit. - Important: Ignore hidden files and directories (such as the
.trash/subdirectory) when listing or processing findings. - Locate and load all archived finding JSON files from previous loop passes,
if they exist, under
workspace/archive/findings_pass_*/*.jsonandworkspace/archive/loop*_findings/*.json. These files represent vulnerabilities that have already been fully evaluated, triaged, and potentially patched in previous passes. - Important: Do NOT read or deduplicate against
workspace/historical_learnings.jsonl(VCS history), as we want to catch regressions if old bugs were reintroduced.
- List the contents of the directory and read the files in
Filter Loop Duplicates (snapshot-gated). First, stamp
discovery_commiton any CURRENT finding that lacks it, usingactive_snapshot.snapshot_idfrom.mantis_state.json(skip when unpinned). Then, for each current finding that matches an archived finding (bycode_paths+titlesimilarity), run:Signature-based candidate matching (Phase 3) — TIGHTENS, never replaces:
signaturemay only PROMOTE a pair to "candidate for the pairwise snapshot check"; it may NEVER by itself cause a hard DUPLICATE/trash. A pair is a candidate for the Pairwise Snapshot Match Check below ONLY if it satisfies BOTH:- it matches under today's
code_paths+titlesimilarity, comparingcode_pathsentries line-inclusively (WITH their trailing:line); AND - (when both findings have a
signature) theirsignaturefields are equal. Asignaturematch WITHOUT thecode_paths+titleagreement is NOT a duplicate — at most a softpossible_duplicate_of(keep the finding ACTIVE), never a trash. Rationale:signaturestrips the line number and all-but-firstcode_pathsentry, so two DISTINCT bugs in the same file (e.g.parser.c:100vsparser.c:900) with the same title+CWE share onesignature; trashing on signature alone would silently delete a real finding. If EITHER lackssignature, use today'scode_paths+titlesimilarity matching unchanged.
PAIRWISE SNAPSHOT MATCH CHECK (decides MATCHED vs NOT_MATCHED) — compares the CURRENT finding's
discovery_commitagainst the ARCHIVED finding'sdiscovery_commitfor this pair (NOT againstSNAPSHOT_ID):- If
snapshot_pinnedis false AND there is NOactive_snapshotin state (MODE-OFF) -> NOT_MATCHED. Stop. (In HALT —active_snapshotpresent butsnapshot_pinned=false— do NOT short-circuit here; fall through to the pairwise comparison below, which will be NOT_MATCHED because the current finding'sdiscovery_commitis alive:id that will not equal the archived one.) - Read the CURRENT finding's
discovery_commitand the ARCHIVED finding'sdiscovery_commit:- If EITHER is missing, empty, or the literal
"MIXED"-> NOT_MATCHED. - If they are NOT byte-for-byte equal to each other -> NOT_MATCHED.
- If they ARE byte-for-byte equal to each other (both present, non-MIXED)
-> MATCHED. There is no other route to MATCHED; never fuzzy-compare. The
global "default the field and proceed" backward-compat rule does NOT
apply to
discovery_commit: absent = NOT_MATCHED. (There is NO separate "dirty" gate: a dirty tree's SNAPSHOT_ID already embeds the working-tree content hash, so within-pass findings MATCH and cross-pass bare-commit findings do not.) Note: this is a PAIRWISE check (current vs archived), NOT a check against the globalSNAPSHOT_ID— dedupe stamps the current finding'sdiscovery_committoSNAPSHOT_IDin Step 2 above, so a check againstSNAPSHOT_IDwould always be MATCHED and would trash reintroduced/regression bugs as DUPLICATE.
- If EITHER is missing, empty, or the literal
Then, using the idempotency rule (Input/Output Contract → Idempotency Guarantee) to avoid double-writes, decide mechanically:
- MATCHED (both present and equal): soft-delete the current finding as a
loop-duplicate exactly as before — set
"status": "DUPLICATE"and"duplicate_of": "<archived_uuid>", clearpossible_duplicate_ofif present, ensuremkdir -p workspace/findings/.trash/, move it there, and log aloop_filtertransaction inworkspace/.tx_log.jsonl. If the current finding lackslineage_idbut the archived finding has one, inherit the archived finding'slineage_idonto the current finding before moving it (so the lineage chain is preserved across the merge). - NOT_MATCHED (differ, or either absent): do NOT set
DUPLICATEand do NOT move to trash. Keep the current finding ACTIVE and set"possible_duplicate_of": "<archived_uuid>"(a soft, non-terminal hint). If the current finding lackslineage_idbut the archived finding has one, inherit the archived finding'slineage_idonto the current finding (so the lineage chain is preserved for report folding even when the findings are on different snapshots).
[!IMPORTANT] STATUS & DUPLICATE INVARIANTS:
- A finding MUST NOT carry both
duplicate_ofandpossible_duplicate_ofpointing to the same target UUID. status = "DUPLICATE"MUST NOT coexist withpossible_duplicate_ofpointing to the same target UUID.- Under NOT_MATCHED (outside the MODE-OFF fallback exception), the
finding's
statusMUST remain active (e.g.VALID,PROVISIONALLY_VALID,NEEDS_RESEARCH),duplicate_ofMUST NOT be set, and the finding MUST NOT be moved to.trash/. Settingpossible_duplicate_ofis a non-terminal hint only.
- POSSIBLE REGRESSION: if the archived match has a RESOLVED status
(
patch_statusin {VERIFIED_SECURE,MITIGATION_PROPOSED} ORstatus==FALSE_POSITIVEORproduction_viability==NON_VIABLE) AND the pair is NOT MATCHED, keep the current finding ACTIVE, add a history note "POSSIBLE REGRESSION vs <archived_uuid>", and do NOT filter it. (A reverted fix re-discovered on new code must never be trashed.) - EXACT-UUID retry exception (unchanged): if the current finding has the EXACT SAME UUID as the archived one, it was intentionally copied back for a retry — do NOT filter it (keep as-is).
- Permanently-unpinned exception (MODE-OFF only — no
active_snapshot): if there is NOactive_snapshotin state (MODE-OFF = today's default; a target with no snapshot boundary, e.g. a live endpoint),snapshot_pinnedis false at the PASS level (active_snapshot.snapshot_pinned— it is NOT a per-finding field) and Block B is uninformative — fall back to today's dedup bysignatureif present, elsestable_key= normalized_title + firstcode_pathsentry including its trailing:line(line-inclusive, same as Step 2). Rationale: stripping:linewould collapse two DISTINCT bugs in the same file (e.g.parser.c:100vsparser.c:900) with the same title+CWE into one — silently deleting a real finding. Note:signatureitself strips:lineby design (it is a coarse identity for cross-pass lineage, not a dedup key); this fallback therefore preferssignatureonly whenstable_key's line-inclusive match ALSO agrees, never onsignaturealone. This preserves dedup for targets that can never MATCH. Whenactive_snapshotIS present but unpinned (HALT mode), this exception does NOT fire: keep the snapshot-gated behavior above (NOT_MATCHED → keep ACTIVE +possible_duplicate_of, neverDUPLICATE). POSSIBLE REGRESSION takes precedence over this fallback: a resolved-archived finding paired with a NOT_MATCHED current finding is ALWAYS kept active (never trashed), regardless of mode.
- it matches under today's
3-Tier Deduplication Ladder (Deterministic & Semantic Matching): When resolving finding lineages and evaluating deduplication candidates, the deduplicator uses a hierarchical 3-tier ladder designed for sub-millisecond fast-path execution with semantic RAG fallback:
Tier 1 (Fast-Path Exact Heuristic Anchors — < 1ms, 0 tokens):
- Exact content identity signature match:
hashlib.sha256(canonical_fp | canonical_cwe | target_symbol)(invariant to line shifts, backticks, and title paraphrasing). - Exact
canonical_filepath + normalized CWE + target_symbolmatch. - Strict line proximity window ($\le 3$ lines) on exact canonical filepath when target symbol is empty.
- Result: If matched, immediately inherit ancestor
lineage_idwithout LLM or embedding overhead.
- Exact content identity signature match:
Tier 2 (RCA Normalization): If Tier 1 heuristic matching does not produce an exact anchor, synthesize a standardized 5-line Root Cause Analysis (RCA) summary:
Component: Normalized canonical filepath and symbol.Vulnerability Class: Canonical CWE taxonomy identifier and name.Root Cause Mechanism: Underlying programming or logic defect.Failure Condition: Specific input, state, or boundary condition.Taint Dataflow: Source-to-sink dataflow trajectory.
Tier 3 (Vector Embedding & Cosine Similarity Scan):
- Project the standardized RCA summary into dense vector embeddings using
the configured multi-provider embedding engine (default:
vertex_ai/gemini-embedding-001). - Perform nearest-neighbor scan over historical
lineage_vectorsin the database using bounded cosine similarity. - CWE Family & Class Structural Guard: When comparing query finding vectors against candidate lineage records, if both findings have explicit CWE classifications (e.g. CWE-89 vs CWE-78), normalize them. If they belong to distinct, incompatible CWE IDs, skip candidate vector comparison entirely to structurally prevent false-merging distinct vulnerability classes regardless of cosine score.
- Positive threshold ($\ge 0.90$): Semantically equivalent findings
with differing phrasing, scanner labels, or line shifts merge into the
same ancestor
lineage_id(default: 0.90, configurable viaEMBEDDING_SIMILARITY_THRESHOLD). - Negative threshold ($< 0.70$): Distinct vulnerability classes (e.g.,
SQLi vs Command Injection, Stored vs Reflected XSS) maintain low
similarity and fail closed, minting a fresh unique
lineage_id.
- Project the standardized RCA summary into dense vector embeddings using
the configured multi-provider embedding engine (default:
Filter Duplicate Findings in Current Batch: Check the current findings against each other to find duplicates. Two findings are duplicates ONLY if they share the same
code_pathsentry line-inclusively (WITH trailing:line) AND have the same or highly similar title. If multiple findings refer to the exact same flaw at the same location, they must be merged. Findings at different lines in the same file are DISTINCT — never merge them.Map/Reduce Chunking Strategy (For Scale): If there are many finding files (e.g., > 20 items), use a Map/Reduce approach to group them by target file or component before checking for overlaps to avoid context window limits.
Token-Optimized Consolidation and Merging: To minimize LLM output tokens and prevent data loss, do not manually rewrite or output the merged JSON files in your response. Instead, follow this pattern:
- Identify Duplicates: Internally map which findings are duplicates of a primary finding.
- Reusable Deterministic Scripting (versioned): Write a reusable helper
script (e.g.
workspace/helpers/merge_findings.py) whose FIRST LINE is exactly# MANTIS_HELPER_VERSION = 2. Before reusing an existing helper, grep its first lines forMANTIS_HELPER_VERSION = 2; if that marker is absent or a different integer (a helper left by an older pipeline version), REGENERATE the helper. Only reuse it when the marker matches. The script must follow these deterministic rules:- Title: Pick the most comprehensive and descriptive title.
- ID: Preserve the unique
"id"of the primary finding being kept. - Severity: Pick the highest severity level specified among the merged items.
- Privileges Required: Inherit the most severe privilege requirement
(priority:
NONE>LOW>HIGH). - Attacker Position: Inherit the most critical position requirement
(priority:
EXTERNAL>INTERNAL_NETWORK>IN_CLUSTER>LOCAL>HOST_SYSTEM>SUPPLY_CHAIN>PHYSICAL_TEMPORARY>PHYSICAL_LONG_TERM). - User Interaction: Inherit the most severe user interaction
requirement (priority:
NONE>REQUIRED). - Code Paths: Collect and deduplicate all file paths and line numbers into a single unique array.
- Description, Mitigation, & Impact: Concatenate cleanly.
- History: Concatenate and preserve all
"history"entries from the merged findings. Append a new entry to the"history"array for this merge action conforming to the schema (containing"stage": "dedupe","action": "merge","details": "Merged duplicate findings: [comma-separated-ids]","pass_number": <current_pass_number>, and"timestamp": "<current_iso8601_timestamp>"). - Unknown keys & provenance (MANDATORY): The script MUST copy through
EVERY key it does not explicitly handle (including
discovery_commit,repro_snapshot_id,patch_base_snapshot,possible_duplicate_of,signature,lineage_id,cwe) from the primary finding onto the merged object — never drop unknown fields. It MUST REFUSE to merge two findings whosediscovery_commitvalues differ (they describe different code versions); leave them separate and log the refusal. When merging findings that all share onediscovery_commit, preserve it unchanged.
- Execute the Script: Run your script to update the primary finding's
file (
workspace/findings/<primary_id>.json) on disk.
Transactional Staged Clean Up: Do not permanently delete redundant files. Ensure the trash directory exists (e.g.,
mkdir -p workspace/findings/.trash/). Before moving, the script must update the duplicate finding files, setting"status": "DUPLICATE"and"duplicate_of": "<primary_uuid>". Move the merged duplicate.jsonfiles to the trash staging directory (workspace/findings/.trash/). For every file moved, append a transaction record toworkspace/.tx_log.jsonl.Transaction Log Schema Format (
workspace/.tx_log.jsonl)Each line must be a self-contained JSON object documenting the transaction:
{"timestamp": "2026-07-14T15:13:00Z", "action": "loop_filter | dedupe_merge", "primary_uuid": "[UUID] (or null for loop_filter)", "moved_uuid": "[UUID]"}This cleans up the directory for downstream stages while preserving rollback capability.
When complete, notify the user.
Loading...
Select a file to preview
Analyzing security...
Checking scan reports and verification data.
Bill of Materials
Everything this skill can do — files, network, commands, and more.