Behavioral regression tests for Agent Skills

Ship skills that
fire on cue.

Tripwire runs real agent sessions against prompts that should and shouldn’t activate your SKILL.md, then turns those cases into a CI gate.

  • Open source
  • No account
  • Runs locally
  • Prompts stay private

Example probe · api-error-handler

Activation record

observed event
missed

The intent matched the skill, but the description did not route the agent to it.

Representative output. Your probes run locally through the selected agent CLI.

200public skills scanned
96%fail a best-practice lint
93%lack explicit “Use when” routing
Reproduce the scan ↗

The silent failure

Valid YAML is not verified behavior.

An agent treats your skill description as routing code. A file can pass every structural check and still miss “help me make this endpoint safer” because the description only mentions “error handling.”

01 / missed activationThe right user asks. The skill stays quiet.
02 / false triggerAn unrelated request loads the wrong instructions.

Static preflight

Check the file before you probe it.

Paste or drop a SKILL.md. The exact CLI lint engine runs in your browser, with no upload and no account.

SKILL.mdevaluated locally
∅No skill loaded

Paste a SKILL.md or open a local file.

The shortest path to confidence

One prompt. Then a contract. Then CI.

Start with the behavior you understand. Expand only after the first real verdict proves the setup works.

  1. 01
    tripwire test

    Verify one real prompt

    Use your existing agent login. No generation key, account, or global install.

  2. 02
    tripwire analyze

    Map the routing boundary

    Generate positive, adjacent, negative, and paraphrased cases for review.

  3. 03
    tripwire init

    Gate every skill change

    Commit the scenarios and one workflow file. Regressions fail the pull request.

Build my first command

GitHub-native handoff

The Action is a file, not another account.

Paste a repository URL in setup. Tripwire prepares GitHub’s file editor with the workflow, ready for you to review and commit. Start with free static checks, then enable behavioral probes after your scenarios exist.

  • ✓Checks changed SKILL.md files
  • ✓Annotates the exact lines on the diff
  • ✓Posts one updated pull request summary
  • ✓Keeps provider keys on your runner
Add to GitHub
.github/workflows/tripwire.ymlreview before commit
name: Tripwire
on: pull_request
jobs:
  skills:
    runs-on: ubuntu-latest
    permissions:
      contents: read
      pull-requests: write
    steps:
      - uses: actions/checkout@v5
        with: { fetch-depth: 0 }
      - uses: bharath31/tripwire@v1

Straight answers

Before you run it.

How is this different from a SKILL.md linter?+

A linter checks structure and authoring rules. Tripwire also opens a real agent session and observes whether the skill activates. That catches valid-looking descriptions that miss intended prompts or fire for unrelated work.

Do I install a GitHub App?+

No. A GitHub Action is a workflow file in your repository. The setup guide prepares that file for you to review and commit on GitHub, or you can create it locally with tripwire init.

Does Tripwire upload my skill or prompts?+

No. Browser lint runs in the page. Behavioral probes run through your agent CLI on your machine or GitHub runner. Anonymous telemetry contains only the command, agent, outcome, source, and version, and can be disabled.

Which agents work today?+

Claude Code activation detection is live verified. Codex CLI and Gemini CLI adapters are available but explicitly marked experimental.

No account · about three minutes

Give one prompt
a real pass or fail.

Run your first prompt