Open-source agent skills

Agent skills for Claude Code, Codex & Cursor.

Give your coding agent a real browser, a repeatable workflow, and evidence you can inspect. Use your logged-in Chrome session or an isolated cloud browser.

6
portable skills
4
agent harnesses
120
trigger evals per model

Six skills for browser work.

Choose the job you need done. Every card opens a first-party guide with installation, supported agents, and related workflows.

web-browse

Browse

Open pages, fill forms, extract content, and take screenshots in your local Chrome session or an isolated cloud browser.

Best for: Opening URLs, researching pages, filling forms, and inspecting authenticated apps.

Learn how to use web-browse

ux-audit

Review

Review a product flow, capture screenshots at each step, and report usability problems with a prioritized fix list.

Best for: Onboarding, checkout, mobile layouts, error states, and logged-in product flows.

Learn how to use ux-audit

prove-it

Evidence

Turn testable claims into a pass/fail table with browser screenshots, console output, and network evidence for every target.

Best for: Code review, pre-merge verification, and answering “did the agent actually test this?”

Learn how to use prove-it

before-after

Evidence

Capture a page before and after a UI change at the same viewport size, then arrange the screenshots in a pull request table.

Best for: UI changes, visual fixes, design-system work, and release evidence.

Learn how to use before-after

thinkbrowse-cli

Control

Control the ThinkRun browser from terminal commands or shell scripts, with recovery instructions for local and cloud sessions.

Best for: CLI workflows, shell scripts, diagnostics, and repeatable browser automation.

Learn how to use thinkbrowse-cli

thinkbrowse-mcp

Control

Use ThinkRun through MCP tools from Claude Code, Codex, Cursor, Cline, Windsurf, or another compatible client.

Best for: Browser MCP requests and tool-level control from an MCP client.

Learn how to use thinkbrowse-mcp

Install for Claude Code, Codex, Cursor, or Gemini CLI.

The repository includes identical skill files in the native folders for Claude Code, Codex, Cursor, Gemini CLI.

# Live skills MCP — follows the canonical GitHub repository
{ "mcpServers": { "thinkrun-skills": { "url": "https://mcp.skillsovermcp.com/mcp/dundas/thinkrun" } } }

# Or clone and copy the native skill layout for your agent
git clone https://github.com/dundas/thinkrun
mkdir -p ~/.claude/skills ~/.codex/skills
cp -r thinkrun/.claude/skills/* ~/.claude/skills/
cp -r thinkrun/.codex/skills/* ~/.codex/skills/

Cursor and Gemini CLI users can use the matching .cursor/skills and .gemini/skills directories.

Inspect the trigger tests and their misses.

The benchmark measures whether each model chooses the right skill across 120 test queries—not whether a browser task completed successfully.

ModelSix-skill trigger accuracyPublished evidence
GPT-5.6 Terra120 / 120Per-query results
Claude Fable 5120 / 120Per-query results
Claude Sonnet 5113 / 120Before/after and MCP misses documented
Review the methodology and raw outputs

Questions about agent skills

What are agent skills?

Agent skills are portable folders of instructions, references, and optional scripts that teach an AI agent a repeatable workflow. ThinkRun ships them in the native directory layout for four coding agents.

What is the difference between a skill and an MCP server?

A skill teaches the agent when and how to perform a workflow. An MCP server supplies tools. ThinkRun provides both: skills route the work, while the ThinkRun MCP or CLI controls the browser.

Can a skill use my logged-in browser?

Yes. Local mode attaches to your real Chrome session for apps behind login. Cloud mode uses an isolated browser for public work. You complete any required password entry yourself.

Do these skills work with Codex and Cursor?

Yes. The same six packages are mirrored for Claude Code, Codex, Cursor, and Gemini CLI, with identical trigger fixtures checked in the repository.