ai-pm-skills

Claude Code skills for AI product managers

See the skills work before you install anything.

Three skills built from how I make AI product decisions. Everything below comes from real runs, including the eval that failed and changed one of the skills.

/build-or-not: it said don't build

Most bad builds aren't badly executed. They solve a problem that turns out to be rare, and a handful of real cases can usually show that in an hour.

Prompt"Add a check to mcp-doctor that flags MCP servers whose tool names aren't snake_case. Keep the sample to 6 repos from GitHub (use gh search), and keep the answer to the decision record only."
Terminal replay of the /build-or-not run checking six MCP server repos for non-snake_case tool names
Trimmed replay of the run. Claude Sonnet, 2026-09-22.
Claim checked
Real, actively used MCP servers define at least one tool name that isn't snake_case.
Sample
6 single-purpose MCP server repos from the mcp GitHub topic, in descending star order. googleapis/mcp-toolbox was scored N/A (its tool names come from user config) and replaced.
Bar, set first
Hit = at least one non-snake_case tool name. Build if 3 or more of 6.
#RepoResultEvidence
1DeusData/codebase-memory-mcp (C)Misssrc/mcp/mcp.c:466: index_repository, search_graph, …
2microsoft/playwright-mcp (TS)MissREADME tool list, ~70 tools: browser_click, …
3github/github-mcp-server (Go)Missget_issue, create_pull_request, …
4idosal/git-mcp (TS)MissTool-name generators replace non-alphanumerics with _
5GLips/Figma-Context-MCP (TS)Missget_figma_data, download_figma_images
6wonderwhy-er/DesktopCommanderMCP (TS)Misssrc/server.ts, 24 tools: read_file, edit_block, …
0 / 6 → Don't build

The check would report "clean" on almost every real server, across C, TypeScript and Go.

What would reopen it
Evidence that config-driven servers commonly produce non-snake_case names. mcp-toolbox's own README example is search-hotels-by-name, so the real risk may sit in user config, which this sample didn't test.

That last line is the part that matters most. The skill doesn't just say no. It says where the real question is.

Full, untrimmed decision record →

/eval-plan: gates before results

An eval only protects you if it can fail. Most can't: the test set has only easy cases, the bar gets set after the numbers come in, or one accuracy number hides the error that hurts users.

InputThe PRD for Agent Outcome Trust Score, a tool that scores whether an AI agent can be trusted. The thing being evaluated is itself an evaluator.
Worst error
The scorer marks a transcript "pass" that a human reviewer would call a failure. That is the exact problem the tool exists to fix, so costs lean toward precision on "pass" verdicts, not overall accuracy.
Test set
Two tiers: tasks for the agent, and human-labeled transcripts to check the scorer against. With one labeler, it proposed a blind re-label after a cooling-off period instead of claiming a two-labeler agreement rate.

Launch gates, set before any results

BehaviorGateWhy
Task success scoring≥95% precisionA false "success" is the failure this tool exists to catch
Graceful escalation≥95%"Handed off honestly" must never score the same as "answered wrong"
Auditability≥90%A non-engineer can reconstruct the score from the transcript alone
Consistency≥95% stableOtherwise the scorecard itself is the noisy variable
Cost per task0% driftMechanical, so no reason to tolerate any
Cheapest baseline
Rule-based scoring first. If it fails, how it fails tells you what an LLM judge has to catch.
If it fails
Report it as a finding about the method. Don't loosen the threshold.

Where the method comes from

A practice triage prototype set its bar in the PRD before any code, and included two off-topic tickets on purpose. The cheap baseline scored 60% and failed the gate. Both off-topic tickets were confidently matched to billing. The first three tickets had all looked fine by eye. Only the deliberate negatives caught it.

Full eval plan →

/agent-trust-review: it said not safe to ship

Most agent risk reviews list what the team did, which makes coverage look complete. This one sorts all 17 risk areas into three lists and ends with two coverage numbers.

InputThe ticket-triage practice prototype and its PRD: an assistant that classifies support tickets and drafts replies for a human to send.
Worst failure
A confidently wrong classification produces a plausible draft on the wrong topic, and a rushed agent sends it.
What it found
It ran the prototype's own eval. The "route to a human when unsure" check exists (rag.py:99) but doesn't fire: both off-topic tickets scored above the 0.08 threshold (0.182 and 0.112) and got confidently wrong drafts.
3
covered, with evidence
3
declined on purpose (e.g. no auto-send in v1)
9
genuinely missing, top one: the threshold doesn't abstain
Not safe to ship

"No auto-send" is a real safeguard: it keeps today's failure "annoying, not dangerous." It doesn't make the classifier ready.

Where it's wrong
The full-map coverage number has an arithmetic slip (3/17 should be 3/15), and it credits the escalation check as covered even though its own top gap shows the check fails.

Full review, with what I checked by hand →

The review it's based on: my own MCP tools

12
areas covered, with evidence
9
declined on purpose, each with a reason
2
genuinely missing

That comes to roughly 80–85% of the niche the tools chose to own, and 25–30% of AI agent testing overall. Both numbers are true, and giving only one would mislead.

What the eval showed (final run)

CaseWith skillPlain Claude
Credits only evidenced items as covered1.001.00
Refuses to certify "it's safe" with no evidence1.001.00
Separates a reasoned decline from a gap, gives two numbers1.000.44

On two of the three cases plain Claude already did as well. What the skill adds is telling a reasoned decline apart from an unexplained gap, and giving two coverage numbers. Plain Claude never gave two. It took four runs to measure cleanly, and every fix was to my test cases, not the skill.

Tested, including a failure

Each skill has an eval suite with launch gates committed before the first run. Each case runs 3 times with the plugin and 3 times without, so the results show what the skills add.

Run 1 failed

With no evidence available, /build-or-not still gave a firm "don't build" from market knowledge it recalled. The skill never said what to do when there's no sample, so the model filled the gap.

Fix "Can't decide yet" is now its own outcome, naming the sample that would settle the question. The grader and gates didn't change. Run 2 passed every gate.

BehaviorWith the skillsPlain Claude
States the bar before deciding3 of 3 runs0 of 3
Refuses a verdict when there's no evidence3 of 30 of 3
Plans a rollback trigger for launch3 of 31 of 3
Separates a reasoned decline from an unexplained gap3 of 32 of 3
Gives two coverage numbers3 of 30 of 3

Where they don't help: plain Claude already spots hits a feature can't reach, and already pushes back on a bar set after the results. Limits: 8 cases, 3 runs each, one model under test. That's a smoke test of the key behaviors, not a benchmark.

All eval results, run by run →

Install

In Claude Code:

/plugin marketplace add vishalhabib99/ai-pm-skills
/plugin install ai-pm-skills@ai-pm-skills

Then:

/ai-pm-skills:build-or-not <the feature someone just proposed>
/ai-pm-skills:eval-plan <path to a PRD, or a description of the AI feature>
/ai-pm-skills:agent-trust-review <agent description, PRD, or repo path>

If you try one on a real decision, I'd like to hear where it was wrong. Open an issue.