@paulnsorensen/hard-cheese
>-
| name | hard-cheese |
| description | Checks whether an author can explain a code change before review. Use when the user requests `/hard-cheese`, `/cheese --hard`, or an understanding check. Use it before a pull request or through the `--hard` pipeline flag. Do not use it for reviews, test hardening, or fixes. |
| license | MIT |
| metadata | {dispatches-agents: true} |
/hard-cheese
The gate reduces epistemic debt. This debt exists when code passes checks, but the author cannot explain it.
Inputs
/hard-cheese [<slug>] [--socratic-cap N=3] [--passing-score N=3] [--no-judge]
Arguments:
<slug>identifies the artifact at.cheese/hard-cheese/<slug>.md. This argument is optional. Without it, use the short SHA ofHEAD. An explicit slug overrides the SHA.--socratic-cap Nsets the maximum number of retries. The gate then marks the artifactFAILEDand returns a non-zero status. The default is3. Vibecheck has no limit, but easy-cheese prevents infinite loops.--passing-score Nsets the minimum SOLO score for PASS. Use a value from1through5. The default is3. The gate treats a previous PASS below this value as stale.--no-judgeenables log-only mode. Record the user's explanation withstatus: LOGGED. Do not start the judge sub-agent. This mode is the easy-cheese equivalent of the optional JSONL telemetry mode in vibecheck. It retains more content. See## Divergence from the paper.
Invocation modes
| Mode | How the gate runs | Where the gate sits |
|---|---|---|
| standalone | The user runs /hard-cheese <slug> before a pull request. |
Outside the pipeline. No upstream skill is required. |
| propagated | /plate --hard runs /hard-cheese <slug> after the final writes and before publication. |
At the verified-artifacts to share-for-review boundary. |
--hard passes through /cheese → /mold → /cook → /press → /age → /cure → /plate. Only /plate runs /hard-cheese.
See ../cheese/references/harness-portability.md for portability requirements. It covers helper resolution, sub-agent dispatch, GitHub operations, and handoff transitions.
Use the bundle or repository helper first. Use ${CLAUDE_SKILL_DIR} only as an optional host fallback.
The handoff blocks define the portable contract because slash commands are host renderings, not the control model.
Flow
Resolve scope.
- Set
diff_base = origin/mainanddiff_head = <short-sha of HEAD>. - Load
.cheese/specs/<slug>.mdas the optional intent reference when it exists. The diff remains the source of truth. - Use the short SHA of
HEADwhen no slug exists. - If the diff against
origin/mainis empty, return0with"nothing to gate on". Do not write an artifact.
- Set
Freshness check. Run the freshness check before you run the gate:
python3 skills/hard-cheese/scripts/hard-cheese.pyz freshness-check \ --slug <slug> --passing-score <n>Exit
0forpreviously_passed. Print"previously passed"and stop. Continue to step 3 forstaleornew. A stale result has exit status2. A new result has exit status3. A result is stale whenHEADchanges or the last PASS score is too low.Compose the vibecheck prompt. Keep it faithful to Sankaranarayanan 2026. Use "share for review" to keep the gate implementation independent.
Before this is shared for review, explain its causal logic in your own words. How does <feature or fix> work? Why does it produce the desired behavior? What state, control flow, or invariants does it rely on?
Show a diff summary with the prompt. For
/plate, also show the complete final evidence:- the final artifact inventory,
- each
{target, backend, verified}completion row, - the tracked artifact diff,
- the quality gate result.
Stop with a non-zero status when
/plateomits one of these four values. Stop with a non-zero status when a completion row hasverified: false.Record the user's explanation as free text. Do not provide coaching or example answers. The explanation is the artifact under test.
Start the judge sub-agent in a fresh context. Use the same pattern as the
/cookfan pathway.- Use
references/judge-prompt.mdas the system prompt. - Provide the passing score, diff summary, optional spec excerpt, and user's explanation.
- Require this JSON object:
{score, level, pass, feedback, socratic_qs}.
See
references/judge-prompt.mdfor the full system prompt and output shape.Skip this step when the user sets
--no-judge. Mark the attemptstatus: LOGGED, write the artifact, and return0.- Use
Process the judge result.
- Mark the attempt PASS when
score >= <passing-score>. - Mark the attempt FAIL when
score < <passing-score>. Show the Socratic questions. Return to step 4 while retries remain. - Mark the attempt ERROR when the judge fails. Print a warning and return
0. See## Divergence from the paper.
Append the attempt row:
python3 skills/hard-cheese/scripts/hard-cheese.pyz append-attempt \ --slug <slug> --status <PASS|FAIL|ERROR> --score <n> \ --feedback "<judge feedback>" --explanation "<user explanation>"- Mark the attempt PASS when
Process an exhausted limit. Set the artifact
status: FAILED. Print the artifact path and return a non-zero status. Stop downstream chains.
Artifact
.cheese/hard-cheese/<slug>.md contains the audit trail. The .gitignore file excludes .cheese/, so the audit trail remains local.
Each file starts with this YAML frontmatter block:
---
slug: <slug>
attribution: Sankaranarayanan 2026 / vibecheck
rubric: SOLO Taxonomy (1-5), pass threshold = <passing-score>
passing_score: <n>
divergence: fail-open on judge error (vibecheck fails closed)
diff_base: <sha>
diff_head: <short-sha>
status: PASS | FAIL | FAILED | LOGGED | ERROR
attempts: <n>
---
append-attempt writes the attempt log as this six-column markdown table:
| timestamp | head_sha | status | score | feedback | explanation |
| --- | --- | --- | --- | --- | --- |
| 2026-06-25T10:00:00+00:00 | a1b2c3d | FAIL | 2 | "Unistructural: lists steps but no causal link" | <user explanation verbatim> |
| 2026-06-25T10:05:00+00:00 | a1b2c3d | PASS | 4 | "Relational: explains why invariant holds" | <user explanation verbatim> |
Each invocation appends attempts and does not overwrite rows. If HEAD changes, append new rows below the earlier rows.
Sub-agent contract — fresh judge
- Use fresh context for every invocation. The code-writing context can bias the judge.
- Resolve a no-tool or read-only
revieweratpowerfulpower andhigheffort. Use the shared agent resolver. - The shared resolver pins each reviewer to
powerful. Do not lower this value for the judge. - Use a general worker only with no-write enforcement. Set
degraded: true. - Use
references/judge-prompt.mdas the system prompt. - Give the judge the diff summary, the optional spec excerpt, and the explanation. Require a JSON reply. Prohibit repository writes.
- Parse the JSON output. On a parse error, log an
ERRORattempt and fail open.
The gate requires a host sub-agent feature. Without this feature, recommend /hard-cheese --no-judge to record the explanation without a grade.
Attribution
Sankaranarayanan, S. (2026). Mitigating 'Epistemic Debt' in Generative AI-Scaffolded Novice Programming using Metacognitive Scripts. Proceedings of the 13th ACM Conference on Learning at Scale. https://arxiv.org/abs/2602.20206
The implementation uses the open-source VS Code extension from the paper's author:
https://github.com/sreecharansankaranarayanan/vibecheck
This SKILL.md, references/judge-prompt.md, and each artifact include the attribution. Thus, the citation stays with the audit trail.
Divergence from the paper
Hard-cheese has two differences from vibecheck:
1. Judge errors. Vibecheck fails closed. The modal blocks code application until the judge recovers or the user retries.
Hard-cheese fails open. The gate records an ERROR, prints a warning, and returns 0.
This policy prevents API failures from blocking pull request work.
2. Telemetry content. Vibecheck records the length of an explanation. It never records the text of an explanation.
Hard-cheese records the complete text of every explanation in the local artifact. The .gitignore file excludes .cheese/, so this text remains on the author's machine.
Tell the user about this retention before --no-judge records the first explanation.
Add each new difference to this section.
Composition with --auto
--hard and --auto can operate together. Terminal /plate --hard pauses automation once before publication, after /plate verifies the final artifacts.
The user responds to the prompt. PASS permits publication. FAILED stops publication. ERROR uses the documented fail-open behavior.
Commit-only /plate --hard does not run the gate. That path shares nothing. See references/composition.md for new pull requests and non-TTY behavior.
Output
When the gate ends, print:
Hard-cheese artifact: .cheese/hard-cheese/<slug>.md
Status: PASS | FAILED | LOGGED | ERROR
Score: <n>/5 (<SOLO level>, pass ≥ <passing-score>)
Attempts: <n>
The Score line reports the latest judged attempt. Omit this line for LOGGED mode or an ERROR without a scored attempt.
Then print one applicable message:
- On PASS:
Ready to share for review. - On FAILED:
Cap exhausted. Improve understanding of the change before sharing. - On LOGGED:
Telemetry only — judge skipped via --no-judge. - On ERROR: Print one warning that identifies the failure. Include
Fail-open divergence active — gate exited 0; you may share for review at your discretion.
Preferred tools and fallbacks
| Need | Prefer | Fallback |
|---|---|---|
| Diff inspection for the user-facing summary | delta |
git diff --unified=3 |
| Reading the spec (when present) | bounded file read per code-intelligence-routing.md |
host file read |
| Spawning the judge | host sub-agent primitive (Agent() or harness equivalent) |
none — without sub-agent spawn, run --no-judge mode and tell the user the judge is unavailable |
| GitHub / PR context (out of scope here) | n/a | n/a |
Rules
- Run the judge sub-agent in fresh context. Do not use the code-writing context to grade the author's understanding.
- Do not coach the user before the answer. The explanation is the artifact under test.
- Show only the judge's Socratic questions after a FAIL. Do not add hints.
- Pass the user's explanation to the judge unchanged.
- Always run the freshness check. A changed
HEADrequires a new attempt sequence. - Record every ERROR attempt. Show a warning for each judge failure.
- Do not call
/ghor a specific pull request tool. The gate operates before code enters review. - Apply the shared voice rules from
../age/references/voice.md. Report the result. Classify the residual risk ascertain | speculating | don't know. - Do not describe FAILED as
"almost passing".
References
references/judge-prompt.mddefines the SOLO Taxonomy rubric, judge prompt, and JSON output.references/composition.mddefines the complete--hardand--automatrix.references/commands.mdlists the generated bundle commands.
Agent resolution
Resolve the fresh judge through ../cheese/references/agent-resolution.md.
| Work | Preferred types | Permissions/isolation | Minimum power | Effort | Fallback |
|---|---|---|---|---|---|
| Grade the explanation | reviewer | no-tool or read-only, fresh-context | powerful | high | compatible reviewer, then general |
The canonical hard-cheese audit includes the shared agent_resolution block.
Loading...
Select a file to preview
Analyzing security...
Checking scan reports and verification data.
Bill of Materials
Everything this skill can do — files, network, commands, and more.