Your agent completed a demo. Now prove the runtime deserves real tools, data, and users. RateMyHarness audits the loop, tool dispatch, context, permissions, and recovery around an AI agent — and issues a verdict a skeptical engineer can check: severity-sorted findings, four independent evidence lanes, five hard vetoes.
Issue list
BLOCKER · H-002 Cancellation does not stop tool calls — the agent keeps spending money after the user presses Stop.
HIGH · H-005 Retry re-executes a non-idempotent transfer — one approval can charge twice.
To verify
UNKNOWN · U-003 Tenant isolation has not been exercised — one user's data may reach another.
Evidence lanes
Each skill is a standalone plugin with the same evidence contract. Install only what you need.
Four clauses most AI reviewers can't keep — because keeping them means disagreeing with you.
Hard gates — authority bypass, cross-boundary leakage, runaway execution, duplicate irreversible effects, fabricated completion — block the release regardless of the weighted average. Accepting the risk doesn't turn a gate into a pass.
Every claim is labeled E3 reproduced, E2 instrumented, E1 static, or E0 unverified. A reachable code defect is real, but its runtime consequence stays labeled as inference until it's exercised.
Findings advance to verified-fixed only on fresh passing evidence from an independent retest context. The pass that wrote the fix never grades its own work.
Repository checks and schema validation cannot masquerade as a critical-journey run. Public or privileged targets require fresh runtime evidence on the exact build under review — a mutable latest is not an identity.
Pick a reviewer role (product owner, staff runtime engineer, red-team, SRE, oral-defense professor) and a review level. It never silently defaults to the strictest — or the friendliest.
Golden path, one realistic failure, one retry of a state-changing action, one authority boundary, one lifecycle edge — captured as traces, logs, and persisted state. First audit is read-only.
Every confirmed issue as one plain-language line, severity-sorted, uncapped. Four evidence lanes that can't substitute for each other. A decision with a named maximum safe target.
Authorized fixes land in batches with recorded scope; a separate context reruns the original reproduction. Only its fresh passing evidence closes a gate.
Pick one method per client and scope. The first audit runs read-only — your sandbox, permissions, and approvals remain the security boundary.
/plugin marketplace add \ AmsonntagChow/ratemyharness /plugin install \ ratemyharness@amsonntagchow-ratemyharness
codex plugin marketplace add \ AmsonntagChow/ratemyharness codex plugin add \ ratemyharness@ratemyharness
npx skills add \ AmsonntagChow/ratemyharness \ --skill ratemyharness