ratemyharness
RateMyHarness · evidence-backed release review

A demo is not evidence.

Your agent completed a demo. Now prove the runtime deserves real tools, data, and users. RateMyHarness audits the loop, tool dispatch, context, permissions, and recovery around an AI agent — and issues a verdict a skeptical engineer can check: severity-sorted findings, four independent evidence lanes, five hard vetoes.

Audit · R-014 target: public-launch ref: sha256:9f3c…e2a1

Issue list

BLOCKER · H-002 Cancellation does not stop tool calls — the agent keeps spending money after the user presses Stop.

HIGH · H-005 Retry re-executes a non-idempotent transfer — one approval can charge twice.

To verify

UNKNOWN · U-003 Tenant isolation has not been exercised — one user's data may reach another.

Evidence lanes

deterministic-checksFAIL
critical-journey-e2ePASS
probabilistic-evalUNVERIFIED
continuous-evidenceN/A — no deployment yet
Maximum safe target: internal-demo
Blocking gates: runaway-execution, duplicate-irreversible-effect
NOT READY
The RateMy family

One rubric discipline, four artifacts under audit.

Each skill is a standalone plugin with the same evidence contract. Install only what you need.

Why it says no when others say yes

The review is a contract, not a vibe.

Four clauses most AI reviewers can't keep — because keeping them means disagreeing with you.

§ 1

No score averages away a veto.

Hard gates — authority bypass, cross-boundary leakage, runaway execution, duplicate irreversible effects, fabricated completion — block the release regardless of the weighted average. Accepting the risk doesn't turn a gate into a pass.

§ 2

Evidence has states.

Every claim is labeled E3 reproduced, E2 instrumented, E1 static, or E0 unverified. A reachable code defect is real, but its runtime consequence stays labeled as inference until it's exercised.

§ 3

A plausible diff is not a fix.

Findings advance to verified-fixed only on fresh passing evidence from an independent retest context. The pass that wrote the fix never grades its own work.

§ 4

Green CI proves structure, not behavior.

Repository checks and schema validation cannot masquerade as a critical-journey run. Public or privileged targets require fresh runtime evidence on the exact build under review — a mutable latest is not an identity.

How a review runs

Settings first. Evidence second. Verdict third.

1

Resolve the settings gate

Pick a reviewer role (product owner, staff runtime engineer, red-team, SRE, oral-defense professor) and a review level. It never silently defaults to the strictest — or the friendliest.

2

Behavior before internals

Golden path, one realistic failure, one retry of a state-changing action, one authority boundary, one lifecycle edge — captured as traces, logs, and persisted state. First audit is read-only.

3

Deliver the verdict

Every confirmed issue as one plain-language line, severity-sorted, uncapped. Four evidence lanes that can't substitute for each other. A decision with a named maximum safe target.

4

Fix and independent retest

Authorized fixes land in batches with recorded scope; a separate context reruns the original reproduction. Only its fresh passing evidence closes a gate.

Install

One repository, three clients.

Pick one method per client and scope. The first audit runs read-only — your sandbox, permissions, and approvals remain the security boundary.

Claude Code
/plugin marketplace add \
  AmsonntagChow/ratemyharness
/plugin install \
  ratemyharness@amsonntagchow-ratemyharness
Codex
codex plugin marketplace add \
  AmsonntagChow/ratemyharness
codex plugin add \
  ratemyharness@ratemyharness
Any Skills client
npx skills add \
  AmsonntagChow/ratemyharness \
  --skill ratemyharness