Behavioral regression tests for Agent Skills
Ship skills that
fire on cue.
Tripwire runs real agent sessions against prompts that should and shouldn’t activate your SKILL.md, then turns those cases into a CI gate.
- Open source
- No account
- Runs locally
- Prompts stay private
Example probe · api-error-handler
Activation record
The intent matched the skill, but the description did not route the agent to it.
Representative output. Your probes run locally through the selected agent CLI.
The silent failure
Valid YAML is not verified behavior.
An agent treats your skill description as routing code. A file can pass every structural check and still miss “help me make this endpoint safer” because the description only mentions “error handling.”
Static preflight
Check the file before you probe it.
Paste or drop a SKILL.md. The exact CLI lint engine runs in your browser, with no upload and no account.
Paste a SKILL.md or open a local file.
The shortest path to confidence
One prompt. Then a contract. Then CI.
Start with the behavior you understand. Expand only after the first real verdict proves the setup works.
- 01
tripwire testVerify one real prompt
Use your existing agent login. No generation key, account, or global install.
- 02
tripwire analyzeMap the routing boundary
Generate positive, adjacent, negative, and paraphrased cases for review.
- 03
tripwire initGate every skill change
Commit the scenarios and one workflow file. Regressions fail the pull request.
GitHub-native handoff
The Action is a file, not another account.
Paste a repository URL in setup. Tripwire prepares GitHub’s file editor with the workflow, ready for you to review and commit. Start with free static checks, then enable behavioral probes after your scenarios exist.
- ✓Checks changed SKILL.md files
- ✓Annotates the exact lines on the diff
- ✓Posts one updated pull request summary
- ✓Keeps provider keys on your runner
name: Tripwire
on: pull_request
jobs:
skills:
runs-on: ubuntu-latest
permissions:
contents: read
pull-requests: write
steps:
- uses: actions/checkout@v5
with: { fetch-depth: 0 }
- uses: bharath31/tripwire@v1Straight answers
Before you run it.
How is this different from a SKILL.md linter?+
A linter checks structure and authoring rules. Tripwire also opens a real agent session and observes whether the skill activates. That catches valid-looking descriptions that miss intended prompts or fire for unrelated work.
Do I install a GitHub App?+
No. A GitHub Action is a workflow file in your repository. The setup guide prepares that file for you to review and commit on GitHub, or you can create it locally with tripwire init.
Does Tripwire upload my skill or prompts?+
No. Browser lint runs in the page. Behavioral probes run through your agent CLI on your machine or GitHub runner. Anonymous telemetry contains only the command, agent, outcome, source, and version, and can be disabled.
Which agents work today?+
Claude Code activation detection is live verified. Codex CLI and Gemini CLI adapters are available but explicitly marked experimental.
No account · about three minutes