# Hodios paste pack: Testing

Everything in Testing from Hodios, the open prompt library by Hermes IDE: 33 entries, catalog 2026.1004.3.

Every entry is dedicated to the public domain under CC0 1.0. Copy, change and share them freely, no attribution needed.

Browse and search the library at https://hermes-ide.com/prompts

## How to use

Find an entry below and copy the text inside its block into ChatGPT, claude.ai or any chat. Replace each [PLACEHOLDER] with your own material. Personas, rules and styles work best as custom instructions or project instructions.

## Contents

- Testing
  - [Add a regression test for a bug](#add-regression-test) (prompt)
  - [Add characterization tests to legacy code](#add-characterization-tests) (prompt)
  - [Build a fake for an external API](#build-fake-for-external-api) (prompt)
  - [Exploratory tester](#exploratory-tester) (persona)
  - [Find and fill the riskiest test gaps](#fill-test-gaps) (prompt)
  - [Fix a flaky test](#fix-flaky-test) (prompt)
  - [Fix a red test suite after an upgrade or merge](#fix-failing-tests-track) (workflow)
  - [Generate API tests from a spec](#generate-api-tests-from-spec) (prompt)
  - [Plan a device and browser test matrix](#plan-device-and-browser-matrix) (prompt)
  - [Raise meaningful test coverage across a codebase](#test-coverage-campaign-track) (workflow)
  - [Refactor tests for clarity without losing coverage](#refactor-test-suite) (prompt)
  - [Review test quality](#review-test-quality) (prompt)
  - [Run mutation testing on a module](#run-mutation-testing) (prompt)
  - [Speed up a slow test suite](#speed-up-test-suite) (prompt)
  - [Test an app under bad network conditions](#test-app-under-bad-network) (prompt)
  - [Test engineer](#test-engineer) (persona)
  - [Test firmware off target](#test-firmware-off-target) (prompt)
  - [Test game mechanics deterministically](#test-game-mechanics-deterministically) (prompt)
  - [Test-writing rules](#test-writing-rules) (rule)
  - [Unit test data transformations](#unit-test-data-transformations) (prompt)
  - [Write a fuzz harness](#write-fuzz-harness) (prompt)
  - [Write a resilient end-to-end test](#write-e2e-test) (prompt)
  - [Write a test plan](#write-test-plan) (prompt)
  - [Write consumer-driven contract tests](#write-contract-tests) (prompt)
  - [Write exploratory test charters](#write-exploratory-test-charters) (prompt)
  - [Write Gherkin scenarios](#write-gherkin-scenarios) (prompt)
  - [Write integration tests with real dependencies](#write-integration-tests) (prompt)
  - [Write manual test cases](#write-manual-test-cases) (prompt)
  - [Write native mobile UI tests](#write-native-ui-tests) (prompt)
  - [Write property-based tests](#write-property-based-tests) (prompt)
  - [Write test data factories](#write-test-data-factories) (prompt)
  - [Write unit tests](#write-unit-tests) (prompt)
  - [Write visual regression tests](#write-visual-regression-tests) (prompt)

---

<a id="add-regression-test"></a>

## Add a regression test for a bug

`add-regression-test` · prompt · Testing · https://hermes-ide.com/prompts/add-regression-test

Writes the smallest test that fails on the buggy code and passes with the fix, and proves both by running it. Use after fixing a bug, or before fixing one, so it cannot return.

````markdown
<context>
A regression test is only worth its place in the suite if it fails without the fix. Many "regression tests" pass on the broken code too, because they test a neighbouring path or assert too little. The proof is running the test against both versions.
</context>

<task>
Add a regression test for: [BUG]
1. State the bug as one triggering input and one expected result.
2. Find the lowest level where the bug can be observed (unit before integration before end-to-end), and the existing test file where a test for that code belongs.
3. Write one focused test with that input and the expected result. Name it after the behaviour, and reference the issue in a comment if there is one.
4. Prove it:
   - On the code without the fix, the test must fail, and fail for the right reason (the assertion on the bug, not an import or setup error). If the fix is already applied, revert it temporarily, for example with `git stash` or by checking out the parent commit of the fix in a separate worktree.
   - On the code with the fix, the test must pass.
   - If the bug is not fixed yet, the test fails now; report that and leave the fix to the user.
5. Run the surrounding test file or suite to confirm nothing else broke, and restore the work tree to the state you found it in.
</task>

<constraints>
- One bug, one test. Add a second test only for a distinct boundary of the same bug, and say why.
- Do not change production code, except to temporarily revert the fix during the proof.
- Never leave the work tree with the fix reverted or with stashed changes the user did not make.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Test
The file path and the test code as a diff.
## Proof
Two results with commands: without the fix (failing, with the assertion message) and with the fix (passing). If the bug is not fixed yet, the failing run only.
## Notes
Anything that limits the test, such as a bug that is only observable end to end, or "None".
</output_format>
````

---

<a id="add-characterization-tests"></a>

## Add characterization tests to legacy code

`add-characterization-tests` · prompt · Testing · https://hermes-ide.com/prompts/add-characterization-tests

Pins down what untested legacy code does today with characterization and golden-master tests, bugs included, so it can be changed safely. Use before refactoring or modifying code with no tests.

````markdown
<context>
A characterization test records what the code actually does, not what it should do. It is a safety net for a later change: if a refactor alters any output, a test fails. That means the tests must pin current behaviour exactly, including odd and probably wrong behaviour, and must fail when the behaviour changes. Tests that only check "no exception" or that assert what the author guessed the code does give false confidence.
</context>

<task>
Write characterization tests for:
[CODE]


1. Find the entry points (from the list above, or from callers in the repository) and test through the highest-level one that is practical to call. Avoid testing private helpers that a refactor will move.
2. Find the seams that make the code nondeterministic or hard to call: current time, randomness, generated ids, environment, file system, network, database, global state. For each, choose the least invasive way to control it: an existing parameter or injection point first, then a test double at the module boundary, then a minimal seam (extract a parameter with the current value as its default). Name any production change you need; keep it behaviour-preserving.
3. Choose inputs that exercise every branch you can see: typical values, boundaries, empty and missing values, error paths, and combinations of flags. Read the conditionals to derive them.
4. Capture current outputs:
   - for small outputs, assert exact values;
   - for large or structured outputs (reports, HTML, JSON, files), write a golden-master or approval test that stores the output in a snapshot file, with scrubbers that normalise timestamps, ids and unordered collections so the snapshot is stable;
   - record side effects too: calls to collaborators, rows written, messages sent, exceptions raised.
   Derive expected values by running the code where you can. If you cannot run it, derive them by tracing the code and mark those tests "traced, confirm on first run".
5. Check the net catches change: for each important branch, describe a small mutation (flip a comparison, drop a line) and confirm a test would fail. Add inputs where none would.
</task>

<constraints>
- Do not fix bugs. Pin the current behaviour and list it under "Suspicious behaviour", with the test name, so a human decides later.
- Do not refactor production code beyond the minimal seams named in step 2.
- Name tests by behaviour (`returns_zero_discount_when_cart_empty`), not by number.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
</constraints>

<output_format>
## Behaviour inventory
Table: Entry point | Input class | Current output or side effect.
## Seams
Bullets: the nondeterminism or dependency, and how the tests control it (including any production change).
## Tests
The complete test file or files, with snapshot files if any.
## Suspicious behaviour
Table: Behaviour | Test that pins it | Why it looks wrong. Or "None".
## Coverage and gaps
Branches covered, branches not covered and why, and the mutations you checked.
</output_format>
````

---

<a id="build-fake-for-external-api"></a>

## Build a fake for an external API

`build-fake-for-external-api` · prompt · Testing · https://hermes-ide.com/prompts/build-fake-for-external-api

Builds a test double for a third-party API, choosing an in-memory fake, stub server or record-and-replay, with contract checks against the real API and error, timeout and rate-limit cases.

````markdown
<context>
The user's code integrates with [API_NAME] and its tests either call the real API (slow, flaky, costly, sometimes impossible in CI) or mock individual methods so tightly that tests pass while the integration is broken. A good fake behaves like the API for the subset the app uses, keeps state where the API does (create then fetch returns the same object), produces the API's real error shapes, and is checked against the real API often enough that it cannot drift.

Choosing the double:
- In-memory fake behind the app's own client interface: fastest, best for unit and service tests, needs an interface seam.
- Stub HTTP server (for example WireMock, MockServer, a small local server, or an HTTP mocking library at the transport layer): tests the real client code, serialisation and headers.
- Record and replay (VCR-style cassettes): cheap to start, but recordings go stale and can capture secrets; use for a few smoke paths, scrub them, and re-record on a schedule.
</context>

<task>
<client_code>
[CLIENT_CODE]
</client_code>

1. List the operations the app actually uses, the fields it reads from responses, and the state those operations imply (objects created, updated, listed).
2. Recommend the double (or a combination, for example an in-memory fake for service tests plus a stub server for client tests), with the reason for this codebase.
3. Write the fake:
   - Same interface as the real client (or the same HTTP routes and payloads for a stub server).
   - Realistic state: ids in the API's format, timestamps from an injected clock, pagination behaving like the API's.
   - Configurable failure injection: per-call errors with the API's real error body and status codes, timeouts or slow responses, rate limiting (429 with the retry header the API uses), partial failures, and webhooks or async callbacks if the app depends on them.
   - Test helpers to inspect what was sent (last request, all requests) without exposing internals to production code.
4. Write example tests that use the fake: the happy path, a retried rate limit, a timeout, an API validation error surfaced to the user, and idempotency on retries if the API supports idempotency keys.
5. Keep the fake honest: a small contract suite that runs the same assertions against the fake and against the real API's sandbox (on a schedule or before releases, not on every pull request), checks response shapes against the API's published schema if there is one, and fails when they diverge. Name what to do when the vendor changes their API.

If the client code does not show the response fields the app uses, ask for them and stop. If no sandbox exists, say what to use instead (recorded production-safe responses, vendor docs) and the risk.
</task>

<constraints>
- Do not invent error codes, headers or payload fields for [API_NAME]; use those in the code or the docs the user gave, and mark others [CHECK DOCS].
- No real credentials, keys or customer data in fakes, fixtures or recordings; scrub recordings.
- The fake must live in test code and never be reachable from production builds.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Choice of double
The recommendation and why, in under 120 words.
## The fake
Code.
## Failure modes
Table: failure | how to trigger it in the fake | real API behaviour it mimics | source (code, docs or [CHECK DOCS]).
## Tests using the fake
Code.
## Keeping it honest
The contract suite, when it runs, and the drift procedure.
</output_format>
````

---

<a id="exploratory-tester"></a>

## Exploratory tester

`exploratory-tester` · persona · Testing · https://hermes-ide.com/prompts/exploratory-tester

Acts as a curious, sceptical exploratory tester who hunts bugs automation misses with tours, heuristics and oracles, and writes bug reports developers can reproduce. Use as a testing partner.

````markdown
From now on, work as this persona: Exploratory tester.

You are an exploratory tester. You treat testing as learning about a product fast enough to find the problems that matter before users do. Scripts and automated checks confirm what someone already expected; you go looking for what nobody expected. You care about risk to real people: lost data, wrong money, locked-out users, confusing errors, and the person on an old phone with a bad connection.

How you work:
- You start by asking what the feature is for, who uses it, what changed, what is already covered by automated tests, and what would hurt most if it broke. Without that you say what you are assuming.
- You test in focused sessions guided by a charter: explore a target, with some resources, to discover a kind of information. You report what you covered and what you did not.
- You model the product with heuristics such as SFDIPOT (structure, function, data, interfaces, platform, operations, time) and pick tours to match: the money tour, the data tour following one record through every screen and export, the back-button and interruption tour, the permissions tour, the "bad neighbour" tour of other features that share the same data.
- You vary what checks forget: boundaries and zero-one-many, empty, very long, Unicode, emoji, right-to-left and pasted text, time zones and daylight saving changes, two tabs or two users at once, slow or dropped networks, session expiry, undo and retry, and switching roles mid-flow.
- You name your oracles: the spec, consistency with the rest of the product, comparable products, user expectations, standards such as WCAG, and the history of past bugs. When no oracle says whether something is wrong, you raise it as a question, not a bug.
- When you are given a running system, screenshots or logs, you work from them. When you are not, you propose the tests and say exactly what to try and what to observe.

What you flag:
- Data loss or corruption, wrong totals, duplicate actions, security and privacy leaks (another user's data, secrets in URLs or logs), and states users cannot get out of.
- Error messages that blame the user, hide the cause or offer no next step.
- Inconsistencies: the same thing named or calculated differently in two places.
- Accessibility barriers: keyboard traps, missing labels, contrast, focus loss.
- Gaps in the spec that the team has not decided, phrased as questions with the options.

Your boundaries:
- You do not test systems you have not been asked to test, run destructive tests against production, or use real people's personal data; you ask for a test environment and test accounts.
- You do not invent bugs or results; anything you did not observe is labelled as a hypothesis with the test that would confirm it.
- For security issues beyond ordinary misuse, you hand over to a security specialist rather than attempting exploitation.

Your habits:
- Bug reports have a specific title (what is wrong, where, under what condition), environment and version, minimal numbered steps, expected and actual results, evidence, frequency, and severity separated from priority.
- You isolate before reporting: you cut the steps down until every remaining one is needed.
- You end a session with a short debrief: covered, not covered, bugs, questions and the next charter you would run.
- You are generous with developers and precise with facts. You never mock a bug or the person who wrote it.
````

---

<a id="fill-test-gaps"></a>

## Find and fill the riskiest test gaps

`fill-test-gaps` · prompt · Testing · https://hermes-ide.com/prompts/fill-test-gaps

Finds untested behaviour that matters most, ranked by risk rather than coverage percentage, and writes tests for the top gaps. Use when a module feels under-tested or before a risky change.

````markdown
<context>
Coverage percentage measures which lines ran, not which behaviours are checked. A module can show 90% coverage while its error handling, money arithmetic and permission checks are never asserted. The useful question is which untested behaviour would hurt most if it broke.
</context>

<task>
Find the riskiest test gaps in [SCOPE] and fill up to 5 of them.
1. Map the behaviours in scope: public functions, endpoints, state transitions, error paths, validations, permission checks.
2. Map the existing tests to those behaviours. A behaviour counts as covered only if a test asserts its result. Lines that merely run do not count.
3. Rank each uncovered behaviour by impact (money, data loss, security, user-visible failure) times likelihood (complex logic, recent churn in `git log`, past bugs, many callers).
4. Write tests for the top 5 gaps, following the project's existing test conventions. Each test must assert a specific result.
5. Run them. A test that fails on current code may have found a bug: keep it, mark it as expected to fail or skipped with a clear reason using the framework's mechanism, and report it. Do not change production code.
</task>

<constraints>
- Rank by risk, not by how easy a test is to write.
- Do not write tests whose only purpose is to raise coverage, such as tests that call code without asserting a result, or tests of trivial getters.
- Cite `path:line` for every gap.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Gaps
A table, highest risk first: # | Behaviour | Where | Why it is risky | Filled (yes or no).
## Tests
The new tests as a diff.
## Run
The command and its result. List any test that exposed a bug, with input, expected and actual.
## Remaining gaps
The gaps you did not fill, one line each, or "None".
</output_format>
````

---

<a id="fix-flaky-test"></a>

## Fix a flaky test

`fix-flaky-test` · prompt · Testing · https://hermes-ide.com/prompts/fix-flaky-test

Finds why a test passes and fails intermittently and fixes the cause instead of adding retries. Use when a test fails only sometimes, locally or in CI.

````markdown
<context>
A flaky test passes and fails on the same code. Retries and longer timeouts hide the defect and teach the team to ignore red builds, so the goal is the cause, not a green run. Sometimes the flakiness is in the product code rather than the test, and then it is a real bug that users can hit.
</context>

<task>
Investigate [TEST].
1. Read the test, its fixtures and setup, and the code it exercises before running anything.
2. List the sources of nondeterminism you can see:
   - time: the current date or time, time zones, timers, timeouts that are too tight;
   - randomness: random data, unseeded generators, generated ids;
   - ordering: unordered collections, query results without ORDER BY, parallel tests, test order;
   - shared state: globals, singletons, caches, databases, files or ports used by other tests;
   - concurrency: unawaited promises, background work, sleeps used for synchronisation;
   - the outside world: network, external services, environment variables, locale.
3. Reproduce the failure: run the test repeatedly, in random order, in parallel, or alongside the tests that run before it in CI. Report how often it fails.
4. Fix the cause: wait on the condition instead of a duration, inject the clock or the seed, isolate the state, sort before comparing. If the race is in the product code, fix it there and say so.
5. Run the test enough times to show the failure is gone, using the same method that reproduced it.
</task>

<constraints>
- Never add retries, sleeps or longer timeouts as the fix.
- Never delete, skip or quarantine the test as the fix. If quarantine is needed while the fix lands, say so separately.
- If you cannot reproduce the failure, say so, report the most likely causes ranked with evidence, and do not claim a fix.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Cause
One paragraph: the nondeterminism and how it makes the test fail. Say whether it is in the test or in the product code.
## Fix
The diff, then one sentence on why it removes the cause.
## Evidence
Runs before and after, with the method used and failure counts (for example "7 of 200 failed before, 0 of 200 after").
</output_format>
````

---

<a id="fix-failing-tests-track"></a>

## Fix a red test suite after an upgrade or merge

`fix-failing-tests-track` · workflow · Testing · https://hermes-ide.com/prompts/fix-failing-tests-track

Takes a red test suite back to green in gated steps, clustering failures, proving each root cause and fixing code or outdated tests with evidence. Use after an upgrade or merge breaks many tests.

````markdown
Gets the suite run by `[TEST_COMMAND]` back to green without cheating. A red suite after an upgrade or merge usually holds a handful of root causes behind dozens of failures, plus a few failures that were already there or are flaky. This track finds those causes, fixes the code where the code is wrong, updates a test only when the intended behaviour really changed (and says why), and reports whatever it could not fix.

Rules for every step:
- Work from real command output only. Never report a test as passing, a cause as proven or a count without having run the command that shows it.
- Fix behaviour, not tests. A test may change only when you can point to the intended behaviour change: an upgrade note, a changelog entry, a commit message, a spec or a decision the user approved.
- Stay inside [SCOPE] when it is given, and stop and report before the run changes more than 20 files in total.
- Write artifacts to the paths listed, outside version control unless the user wants them kept. Commit nothing unless the user asked for commits.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.

## Steps

Work through these steps in order. Do not skip a gate.

1. triage (verify)
2. diagnose (plan)
3. fix (build)
4. verify (verify)

### Step 1: Run and cluster the failures

<recent_change>
[RECENT_CHANGE]
</recent_change>

1. Check the environment before the code: dependencies installed from the lockfile, the runtime version the project pins, caches cleared if the upgrade touched build tooling. Many "test failures" after an upgrade are a stale install.
2. Run `[TEST_COMMAND]` once on the whole suite (or on [SCOPE] when given) and save the raw output. Record total, passed, failed, errored and skipped.
3. Group the failures into clusters that share a cause signature: the same error message or exception type, the same failing import or fixture, the same module under test, the same assertion shape. Name each cluster by its signature, not by a guess at the cause.
4. Rerun each failing test, or one representative per cluster, in isolation three times. A test that passes sometimes is flaky: list it separately and leave it for `fix-flaky-test` style work rather than this run.
5. If the recent change is known, check whether the cluster also fails on the commit before it (for example in a separate worktree checked out at that commit, or with `git bisect` over a small range). Mark clusters that already failed before as pre-existing.

Write the artifact with sections Environment, Baseline counts, Clusters (Cluster | Signature | Tests | Isolated result | New or pre-existing), Flaky. Continue to step 2.

Save this step's result to `fix-failing-tests/01-triage.md`.

### Step 2: Prove root causes and plan the fixes

For each cluster from step 1, largest first:

1. Read the failing test, the code it exercises and the part of the recent change that touches either. Find the line where expected and actual first diverge.
2. Classify the cause, with the evidence that proves it:
   - **Code regression**: the code no longer does what the test rightly expects.
   - **Intended behaviour change**: the upgrade or merge deliberately changed behaviour, and the test still expects the old one. Cite the upgrade note, changelog or commit.
   - **Test infrastructure**: a fixture, mock, config or helper broke (renamed API in the test framework, changed default, removed global).
   - **Environment**: versions, missing services, time zone, locale, file paths.
   - **Unknown**: you could not prove it. Say what experiment would settle it.
3. Propose the smallest fix for each cluster and list the files it touches. Prefer one fix at the shared cause over many edits at the symptoms.
4. Count the files the whole plan touches.

Write the artifact with a table: Cluster | Cause class | Evidence | Proposed fix | Files | Test changes and justification. Stop and wait for approval. Approval is essential when the plan changes any test's expectations, touches more than five files, or exceeds the 20-file budget; mark those rows clearly.

Save this step's result to `fix-failing-tests/02-diagnosis.md`.

**Gate:** stop here and wait for the user's approval before step 3 (fix).

### Step 3: Fix, one cluster at a time

Work through the approved plan in order.

1. Apply the fix for one cluster. Keep the change minimal and in the style of the surrounding code.
2. Run that cluster's tests, then the tests of the touched modules. Record the real result.
3. If the fix does not turn the cluster green, or turns something else red, revert it, go back to diagnosis for that cluster, and do not pile a second guess on top of the first.
4. Change a test only where the approved plan says the intended behaviour changed. Update the expectation to the new intended behaviour, keep the assertion as strict as before, and add a one-line comment or commit message citing the reason. Never loosen an assertion, add a broad try/except, mark a test skip or xfail, or special-case a test input to get green.
5. Keep a running count of files changed. If the next fix would take the run past 20 files, or past the plan's file list by more than a file or two, stop and report instead.

Continue to step 4 when every planned cluster is fixed or explained.

### Step 4: Verify the whole suite and report

1. Run `[TEST_COMMAND]` on the full suite (or [SCOPE]) and compare with the step 1 baseline: no test that passed before may fail now, and the skipped count must not have grown.
2. Run the project's linter or type checker if it has one, since fixes can break them.
3. Write the report:

#### Result
Before and after counts from real runs, and the commands used.

#### Fixed
Table: Cluster | Cause | Fix | Files.

#### Tests changed
Table: Test | Old expectation | New expectation | Justification (cite the source).

#### Still failing
Table: Test or cluster | What is known | Next experiment | Why it was not fixed (unknown cause, out of scope, budget reached, needs a decision).

#### Flaky and pre-existing
The tests from step 1 that were left alone, and why.

#### Follow-ups
One line each for anything noticed but not changed.

Save this step's result to `fix-failing-tests/04-report.md`.
````

---

<a id="generate-api-tests-from-spec"></a>

## Generate API tests from a spec

`generate-api-tests-from-spec` · prompt · Testing · https://hermes-ide.com/prompts/generate-api-tests-from-spec

Generates API tests from an OpenAPI or GraphQL schema with positive, boundary, invalid-input, auth and permission cases, response schema checks and property-based fuzzing where tools allow.

````markdown
<context>
The user is a backend or QA engineer with an API contract and wants tests that exercise it systematically. Test stack: recommend one and say why. Tests generated naively from a spec only hit each operation once with a valid body and check for a 200, which proves almost nothing. The defects live in boundaries (min and max lengths, numeric limits, enum values, nullable fields), malformed input that should get a 4xx and gets a 500, missing object-level authorization (user A reading user B's order), undocumented fields leaking in responses, and responses that drift from the schema.
</context>

<task>
<spec>
[SPEC]
</spec>

1. Inventory the operations (method and path, or query and mutation names), their parameters, request bodies with constraints, documented responses and security requirements.
2. For each operation, derive cases by technique:
   - Positive: one minimal valid request and one with every optional field.
   - Boundaries: for each constrained field, values at, just inside and just outside `minLength`, `maxLength`, `minimum`, `maximum`, `pattern`, `enum` and array `minItems`/`maxItems`.
   - Invalid input: missing required fields, wrong types, null where not nullable, unknown fields if `additionalProperties: false`, malformed JSON, wrong content type. Expect the documented 4xx, never a 5xx.
   - Auth: no credentials (401), valid credentials without the scope or role (403), and object-level access with a second user's resource id (403 or 404, never the data). For GraphQL, check field-level authorization and depth or complexity limits.
   - State: create then read, update a deleted resource, idempotency keys and duplicate submissions where the spec has them, pagination edges (empty page, last page, invalid cursor).
3. Rank cases by risk and keep the matrix lean: every operation gets positive, auth and invalid-input cases; boundaries only for constrained fields.
4. Write the tests in the chosen stack with shared fixtures for base URL, two test users with different roles, and data setup and teardown through the API itself. Read secrets from environment variables.
5. Add response schema validation on every test (validate the body against the spec's response schema, for example with an OpenAPI validator or GraphQL type checks), so drift fails loudly.
6. Where tooling exists, add property-based or schema-driven fuzzing (for example Schemathesis for OpenAPI, or a GraphQL fuzzing tool) with a run budget and the checks it should enforce (no 5xx, schema conformance, status codes documented).
7. List spec gaps you found: undocumented error responses, missing constraints, missing security on an operation, inconsistent naming. These are findings, not tests.

If the spec has no constraints or security section, say how that limits the tests and ask whether to infer constraints from the implementation.
</task>

<constraints>
- Every expected status code must come from the spec; when the spec is silent, mark the expectation [ASSUMED] and list it under Spec gaps.
- Never point tests at production or use real customer data or credentials.
- Do not invent endpoints or fields that are not in the spec.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Coverage matrix
Table: operation | positive | boundaries | invalid input | auth | state | priority.
## Test setup
Fixtures, users and roles, environment variables, run command.
## Tests
Code, grouped by operation.
## Schema conformance and fuzzing
The validator wiring and the fuzzing command with its budget and checks.
## Spec gaps
Bullets: gap, where in the spec, suggested fix.
</output_format>
````

---

<a id="plan-device-and-browser-matrix"></a>

## Plan a device and browser test matrix

`plan-device-and-browser-matrix` · prompt · Testing · https://hermes-ide.com/prompts/plan-device-and-browser-matrix

Chooses which devices, OS versions, browsers and screen sizes to test from usage analytics and risk, split into must-test, sampled and emulator-only tiers. Use before a release cycle.

````markdown
<context>
The user leads QA for a [APP_TYPE] product and must decide where to spend limited device and browser testing. Testing everything is impossible; testing only the team's own new phones and latest Chrome misses the users who struggle most. Good matrices are driven by real usage weighted by risk: older OS versions with different permission or WebView behaviour, low-memory Android devices, small and very large screens, WebKit behind most iOS browsers (Chrome on iOS is not Chrome on Android), tablets and foldables if the layout adapts, and the browsers that business customers are locked into.

Usage data: none provided. Budget: not stated; size a minimal and a recommended option.
</context>

<task>
1. Set a coverage target: the share of active users the must-test tier should represent (a common target is 80-90% of sessions), and the minimum supported versions. If usage data is missing, say which report to pull (browser and OS version, device model, viewport, by sessions or revenue), and build a provisional matrix from public-knowledge defaults labelled [VERIFY WITH YOUR DATA].
2. Group the data into dimensions that change behaviour: rendering engine and major version (Blink, WebKit, Gecko), OS major version, screen size class, device performance class (low, mid, high RAM and CPU), input type (touch, mouse, keyboard, screen reader), and locale or right-to-left if relevant. Merge versions that behave the same.
3. Rank combinations by usage share multiplied by risk; add high-risk low-share items deliberately (oldest supported OS, a low-end Android device, an iPhone SE-size screen, Safari on iOS, a tablet, a screen reader).
4. Split into tiers:
   - Tier 1 must-test on real devices or real browsers every release (about 3-6 configurations).
   - Tier 2 sampled: rotated across releases or covered by automated runs in a cloud device or browser farm.
   - Tier 3 emulator or simulator only, or responsive checks for layout.
   - Unsupported: stated explicitly, with the message users see if any.
   For cross-platform apps, build the tiers per OS (Android and iOS side by side) and add the WebView or embedded browser version if any screens are web content.
5. Map test types to tiers: full regression on tier 1, smoke and automated suites on tier 2, layout checks on tier 3.
6. Size the time and cost: hours per release per tier and device-farm minutes, compared with the budget; give a minimal and a recommended option if the budget is unclear. Do not quote prices; give the quantities to price.
7. Say when to revisit the matrix: each quarter, a new major OS or browser release, a usage shift above a threshold (for example a configuration crossing 5% of sessions), or a crash spike on one device family.

If any of these are missing, do not invent them: mark them [X] in the matrix, state the assumption you sized with, and ask for them in a short "Questions" line at the end of Time and cost: the minimum supported OS or browser versions, the release frequency (to turn monthly farm minutes into per-release capacity), and the devices the team already owns.
</task>

<constraints>
- Never present usage shares or device statistics as fact without the user's data; label defaults [VERIFY WITH YOUR DATA].
- Do not quote device-farm prices; give quantities to price.
- Keep tier 1 small enough to run every release within the budget.
</constraints>

<output_format>
## Coverage target
Two or three sentences with the target share and minimum versions.
## Matrix
Table: tier | platform or browser | version | device or viewport | why it is in | share of usage.
## Why these
Bullets for each deliberate high-risk inclusion and each notable exclusion.
## Time and cost
Table: tier | test type | hours per release | farm minutes per release. Then the comparison with the budget.
## Review triggers
Bullets.
</output_format>
````

---

<a id="test-coverage-campaign-track"></a>

## Raise meaningful test coverage across a codebase

`test-coverage-campaign-track` · workflow · Testing · https://hermes-ide.com/prompts/test-coverage-campaign-track

Raises test coverage where it reduces risk, measuring first, writing behaviour tests for risky untested code and checking them with sampled mutation testing. Use for a coverage push that must count.

````markdown
Runs a coverage campaign that buys real safety. Line coverage is easy to inflate with tests that execute code but assert nothing, and a campaign judged by percent drifts there. This track measures, picks the untested code where a bug would hurt most, writes tests that pin down behaviour, then checks a sample with mutation testing to prove the tests would catch a real change. Target: risk-based.

Rules for every step:
- Every number in an artifact comes from a command actually run: `[COVERAGE_COMMAND]`, git history, or the mutation tool.
- Tests go through public behaviour (inputs, outputs, side effects at boundaries), not private helpers or call counts, unless the boundary itself is the behaviour.
- When a new test exposes a bug, do not write the test to expect the buggy result. Mark it as a known failure in the way the project allows (or leave it out), record the bug, and do not fix production code in this campaign unless the user asks.
- Stop and report before adding or changing more than 15 test files.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.

## Steps

Work through these steps in order. Do not skip a gate.

1. measure (discover)
2. risk-map (plan)
3. write-tests (build)
4. mutation-check (verify)
5. report (verify)

### Step 1: Measure

1. Run `[COVERAGE_COMMAND]` and record overall line and branch coverage, and coverage per file or package. If the suite fails, stop and report: a coverage campaign on a red suite measures nothing.
2. Gather risk signals for each source file with low coverage:
   - Churn: commits touching the file in the last six to twelve months (`git log --since=... --name-only`).
   - Bug history: commits or issues mentioning fix, bug or revert for that file.
   - Complexity: branch count or cyclomatic complexity from an existing tool, or a rough count of conditionals.
   - Criticality: money, auth, permissions, data deletion, external integrations, anything the README or architecture docs call core.
3. Note what the existing tests look like: framework, helpers, fixtures, factories, how external services are faked. New tests must match.

Write the artifact: Baseline (overall and per package), Risk table (File | Coverage | Churn | Bug fixes | Complexity | Criticality), Test conventions. Continue to step 2.

Save this step's result to `coverage-campaign/01-measure.md`.

### Step 2: Choose targets

1. Rank the files by risk: high criticality and churn with low branch coverage first. When risk-based names a path or a percentage, rank within it and compute how many uncovered branches the goal needs.
2. For the top targets, list the specific untested behaviours: the uncovered branches and what each one means in domain terms ("refund larger than the original charge", "expired token with a valid refresh token"), the error paths, and the boundary values.
3. Drop code that is not worth testing here (generated code, trivial getters, dead code, thin wrappers over a library) and say why. Dead code goes on the follow-up list rather than getting tests.
4. Fit the plan inside 15 test files.

Write the artifact: Targets (File | Behaviours to test | Why risky | Test file), Skipped and why, Expected coverage change. Stop and wait for approval.

Save this step's result to `coverage-campaign/02-targets.md`.

**Gate:** stop here and wait for the user's approval before step 3 (write-tests).

### Step 3: Write behaviour tests

For each approved target:

1. Write one test per behaviour, named for the behaviour in the project's style. Arrange the minimum setup with existing factories and fakes; assert on the observable outcome, including error types and messages where callers depend on them.
2. Cover the boundaries listed in step 2, not only the happy path.
3. Avoid assertion-free tests, snapshot tests of large structures that nobody reads, tests that assert mocks were called as a stand-in for outcomes, and sleeps.
4. Run the new tests and the surrounding suite. Each new test must pass for the right reason: temporarily break the behaviour (flip a condition locally, never committed) and confirm the test fails, at least for the riskiest ones.
5. Keep count of test files touched against 15.

Continue to step 4.

### Step 4: Check a sample with mutation testing

1. Use the project's mutation tool if it has one, or the standard tool for the stack (for example Stryker, mutmut, cosmic-ray, PIT, cargo-mutants or go-mutesting). If none can be run, apply five to ten manual mutations per sampled file (negate a condition, change a boundary, drop a statement, return early) and run the tests against each, reverting every one.
2. Limit the run to the target files so it finishes in reasonable time.
3. Triage surviving mutants: equivalent (no behaviour change, ignore), unimportant, or important. Strengthen or add tests for the important survivors and rerun.
4. Rerun `[COVERAGE_COMMAND]` for the final numbers.

Continue to step 5.

### Step 5: Report risk reduced

#### Behaviours now protected
Table: Target | Behaviours covered | Why it mattered.

#### Mutation check
Tool or manual method, files sampled, mutants killed / survived / equivalent, and what was strengthened.

#### Coverage
Before and after, overall and for the targets, from real runs. Present it after the behaviours, as supporting evidence.

#### Bugs found
Table: File and line | Behaviour | How the test shows it | Status.

#### Remaining risk
The next targets from the ranking, and anything skipped.

#### Checks
Commands run and real results.

Save this step's result to `coverage-campaign/05-report.md`.
````

---

<a id="refactor-test-suite"></a>

## Refactor tests for clarity without losing coverage

`refactor-test-suite` · prompt · Testing · https://hermes-ide.com/prompts/refactor-test-suite

Cleans up a test file or suite, fixing unclear names, duplicated setup, over-mocking and assertions on internals, while proving with coverage and mutation checks that nothing stopped being tested.

````markdown
<context>
Tests that are hard to read get skipped in review, copied with their mistakes, or deleted when they break. Cleaning them up is worth it, but a test refactor is uniquely risky: a test that silently stops checking something still passes. Every change here has to keep each test failing for the same bugs it caught before.
</context>

<task>
Refactor the tests in [TESTS].
1. Run the tests and record the result and coverage as the baseline.
2. Read each test and list the problems: names that do not say the behaviour and expected result, several behaviours in one test, long duplicated setup, magic values with no meaning, mocks of the code under test or of simple values, assertions on private details instead of observable behaviour, missing assertions, sleeps, and shared mutable state between tests.
3. Fix them in small steps:
   - name each test after the behaviour and the expected outcome;
   - split tests that check unrelated behaviours;
   - move repeated setup into builders, factories or fixtures that make the important values visible in the test;
   - arrange, act and assert in a clear order;
   - replace mocks of internals with real objects or fakes at the boundary;
   - assert on outcomes the caller can observe.
4. After each step, run the tests. They must still pass.
5. Prove nothing was lost: compare coverage with the baseline, and for the tests you changed most, break the production code on purpose (or run a mutation tool if the project has one) and confirm the refactored test still fails.
</task>

<constraints>
- Do not change production code, except temporarily for the mutation check; revert it afterwards.
- Do not delete a test unless another test provably covers the same behaviour; say which one.
- Do not weaken assertions or widen expected values to make a test pass.
- If a test looks wrong rather than just unclear, report it instead of silently changing what it checks.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Problems found
Table: test, problem, fix applied.
## Diff
The refactor as a diff.
## Coverage check
Baseline and final test results and coverage, and the mutation checks run with their results.
## Left alone
Tests you did not change and why, plus any test that looks wrong.
</output_format>
````

---

<a id="review-test-quality"></a>

## Review test quality

`review-test-quality` · prompt · Testing · https://hermes-ide.com/prompts/review-test-quality

Reviews a test suite or diff for weak assertions, over-mocking, hidden coupling, sleeps, nondeterminism and tests that cannot fail, with a concrete rewrite for each problem. Use when reviewing tests.

````markdown
<context>
A test earns its maintenance cost only if it fails when the behaviour it covers breaks and passes otherwise. Many tests do neither: they assert that a result is "not null", verify that a mock was called with whatever the mock returned, pass because an async assertion never ran, break when an internal method is renamed, depend on the order the suite runs in, or sleep and hope. Coverage numbers do not reveal any of this. The quickest way to judge a test is to ask which plausible bug in the code under test it would catch.
</context>

<task>
Review these tests:

<tests>
[TESTS]
</tests>

1. For each test, state in one line the behaviour it claims to check, judged from its name and body.
2. Look for tests that cannot fail: no assertion; assertions inside callbacks, loops or branches that may never run; un-awaited promises or async assertions; exceptions swallowed by `try`/`catch`; expected values computed with the same logic as the code; and comparisons of a mock's return value with itself.
3. Look for weak assertions: checking only existence, type, length or "truthy"; large snapshots nobody reads; asserting a subset when the whole result matters; and error tests that accept any exception instead of the specific one.
4. Look for over-mocking: mocking the unit under test or its pure collaborators, mocking types the project does not own instead of wrapping them, asserting call sequences instead of outcomes, and mocks whose behaviour differs from the real dependency (say how).
5. Look for hidden coupling: shared mutable fixtures, order dependence, global state, tests of private methods or internal structure, and one test covering several behaviours so a failure does not say what broke.
6. Look for nondeterminism: sleeps and fixed timeouts, real clocks and time zones, randomness without a seed, network or file-system dependence, unordered collections compared as ordered, concurrency without synchronisation, and locale-dependent formatting.
7. Mutation check: for the most important tests, name two or three small, realistic bugs in the code under test (an off-by-one, a flipped condition, a missing null check, a dropped field) and say whether each test would catch them. If the code under test was not provided, say what you infer and mark it as an inference.
8. Rewrite each problem test in the same framework and style, keeping its intent, so that it fails for the bug it should catch.

If the tests are fine, say so plainly and do not invent problems.
</task>

<constraints>
- Every finding cites the test name and line, the smell, the concrete bug it lets through or the false failure it causes, and the fix.
- Do not comment on naming or formatting unless it hides what is tested.
- Rewrites stay in the project's framework, helpers and conventions; no new test libraries unless one is clearly needed, and then say why.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
</constraints>

<output_format>
## Verdict
One line: solid | usable with fixes | gives false confidence. Then the main reason.

## Findings
Numbered, most harmful first. Each: `test name:line` - smell - what it lets through or breaks on - fix.

## Bugs these tests would miss
Table: plausible bug | caught? | by which test, or which test should catch it.

## Rewrites
Code blocks with the corrected tests, one per finding that needs code.
</output_format>
````

---

<a id="run-mutation-testing"></a>

## Run mutation testing on a module

`run-mutation-testing` · prompt · Testing · https://hermes-ide.com/prompts/run-mutation-testing

Sets up mutation testing for one module, triages the surviving mutants and writes tests that kill the ones that matter. Use when coverage looks high but you doubt the tests would catch a real bug.

````markdown
<context>
Line coverage says code ran during a test, not that a test would fail if the code were wrong. Mutation testing changes the code in small ways (flip `<` to `<=`, drop a call, return a constant) and reruns the tests; a mutant that survives is a change no test noticed. Running it over a whole repository on day one produces hours of runtime and thousands of survivors nobody reads. The value comes from a narrow scope, a careful triage, and tests that assert behaviour. Common tools: Stryker (JavaScript, TypeScript, C#), PIT (Java, Kotlin), mutmut or cosmic-ray (Python), cargo-mutants (Rust), Gremlins or go-mutesting (Go), Infection (PHP), mutant (Ruby).
</context>

<task>
Run mutation testing on [MODULE] ([LANGUAGE]).

1. Read the module and its tests. Run the existing tests once; if they fail or are flaky, stop and report, because mutation results on a red or flaky suite are meaningless.
2. Pick the tool for [LANGUAGE], unless the project already has one. Configure it to mutate only the module and to run only the tests that cover it. Enable incremental or per-test coverage mode if the tool has one, and set a timeout multiplier so infinite-loop mutants are classed as timeouts.
3. Run it and record: mutants generated, killed, survived, no coverage, timed out, and the mutation score.
4. Triage every survivor into one of:
   - **Important**: a boundary, a branch of business logic, error handling or a security check whose change would be a real bug.
   - **Weak test**: code is covered but the test asserts too little (no assertion on the return value, only "does not throw").
   - **Equivalent**: the mutant behaves identically (for example a change to an unobservable log message or a redundant condition). Explain why in one line.
   - **Low value**: logging, `toString`, generated code. Suggest excluding it in config.
5. Write tests that kill the Important and Weak-test survivors. Each test asserts observable behaviour at the boundary the mutant changed; name it after the behaviour, not the mutant.
6. Re-run the tool on the module and report the new numbers, listing any mutant still alive and why.
</task>

<constraints>
- Do not change production code to kill a mutant unless the mutant revealed a real bug; if it did, say so separately and ask before fixing.
- Never chase a 100% score. Equivalent mutants exist and are not failures.
- Do not commit tool caches or reports unless the project already does.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Setup
Tool, version, config file contents and the command used.
## Results
Table: generated, killed, survived, no coverage, timeout, score.
## Triage
Table: mutant id, file:line, mutation, verdict, reason.
## New tests
Code blocks with file paths.
## Re-run
The real numbers from the second run, or a plain statement that it was not run.
## Next steps
Where a CI threshold makes sense (incremental, on changed code only) and what to exclude.
</output_format>
````

---

<a id="speed-up-test-suite"></a>

## Speed up a slow test suite

`speed-up-test-suite` · prompt · Testing · https://hermes-ide.com/prompts/speed-up-test-suite

Cuts test suite runtime from timing reports and setup code by fixing slow fixtures, sleeps, database resets and isolation-safe parallelism. Use when the tests themselves, not the CI config, are slow.

````markdown
<context>
The user's test suite takes too long locally or in CI, and the cause lives in the tests: setup, fixtures, waits and data handling. Stack: infer from the output and say what you assumed. (Caching dependencies and splitting CI jobs is a separate job; mention it only if the timings show the suite itself is already fast.)

Suites are usually slow because of a few patterns, not because there are many tests: fixed sleeps; expensive setup repeated per test (app boot, migrations, container start, browser launch); truncating or recreating the database instead of rolling back a transaction; real network calls and retries with backoff; end-to-end tests covering what a unit test could; and a runner that uses one core. The usual trap is making it fast by sharing state, which brings back order-dependent, flaky tests.
</context>

<task>
<timings>
[TIMINGS]
</timings>

1. Analyse the timings: total, the top 10 tests or files and their share, the distribution (a long tail of slow tests or uniformly slow), and setup versus test-body time where visible. Do the arithmetic: say what share of the runtime the top items hold.
2. Classify each hotspot: sleep or polling, per-test expensive setup, database reset strategy, external call, heavy test level, large data generation, slow collection or import time, or serial execution.
3. For each, propose the fix and its expected saving as a range, ranked by seconds saved per effort:
   - Sleeps: replace with condition waits or a fake clock.
   - Setup: widen fixture scope only for read-only or immutable resources (session-scoped container or app, per-test transaction).
   - Database: wrap each test in a transaction rolled back at the end, or use templated databases per worker; avoid truncating all tables per test.
   - External calls: fakes or recorded responses at the boundary; disable retries and backoff in tests.
   - Test level: move logic checks down to unit tests; keep a few end-to-end tests for the wiring.
   - Parallelism: the runner's workers (pytest-xdist, Jest workers, `go test -p`, JUnit parallel, RSpec parallel), with a resource per worker (database, port, temp directory).
   - Selection: run tests affected by changed files locally, keeping the full suite on the main branch.
   - Import or collection time: lazy imports, narrower test paths.
4. Show the code changes for the top three fixes.
5. For every change, state the isolation risk and how to guard it: randomised test order, running each file alone, and a check that no test depends on another's data.
6. Give the measurement plan: the command to time the suite, three runs before and after, and the per-test duration report to keep in CI.

If the timings do not include per-test durations, give the command that produces them for this runner and stop.
</task>

<constraints>
- Never trade isolation for speed silently; each shared resource must be immutable or reset.
- Do not delete or skip tests to save time without listing them and asking.
- Savings are estimates until measured; label them as such.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Where the time goes
Table: test or file | seconds | share of total | cause class.
## Ranked fixes
Table: fix | tests affected | estimated saving | effort | isolation risk.
## Changes
Code for the top three fixes.
## Isolation risks
Bullets with the guard for each.
## How to measure
Commands and what to record.
</output_format>
````

---

<a id="test-app-under-bad-network"></a>

## Test an app under bad network conditions

`test-app-under-bad-network` · prompt · Testing · https://hermes-ide.com/prompts/test-app-under-bad-network

Designs and automates tests for slow, lossy, captive-portal and offline networks, covering timeouts, retries, offline queues, duplicate submissions and what users see in each state.

````markdown
<context>
The user's app (web) is used on trains, in basements, on congested mobile networks and behind hotel Wi-Fi login pages, but is tested on office Wi-Fi. The bugs that appear are predictable: spinners that never end because there is no timeout; retries of non-idempotent requests that create duplicate orders or payments; double taps while a slow request is in flight; an offline queue that replays in the wrong order or never; captive portals returning a 200 HTML page the app tries to parse as JSON; stale data shown as current; and errors that blame the user ("Something went wrong") with no way to retry.
</context>

<task>
<app>
[APP_DESCRIPTION]
</app>

1. Define network profiles with numbers: offline; high latency (about 500-800 ms round trip); slow 3G-like (roughly 400 kbps down, 400 ms latency); lossy (5-10% packet loss); flapping (connection drops for 5-30 s every minute); captive portal (all requests answered with an HTML login page or redirect); DNS failure; and server slow (responses delayed 10-30 s). Say which tool applies each on web: browser DevTools throttling and request blocking, Network Link Conditioner on iOS and macOS, the Android emulator's network settings, a proxy such as Charles, mitmproxy or Toxiproxy for loss and latency injection, and airplane mode on a real device.
2. Rank the app's flows by harm if the network fails mid-request: payments and orders first, then other writes, then reads.
3. For each risky flow and profile, list what to check:
   - Timeouts exist and fit the action (connect and read timeouts, not infinite).
   - Retries: only for idempotent requests or with an idempotency key; exponential backoff with jitter; a cap.
   - Duplicate submission: the button disables or the request is deduplicated; the server rejects a repeated idempotency key.
   - Offline: queued writes persist across app restart, replay in order, and resolve conflicts as designed; the user can see what is pending.
   - What the user sees: loading state within 100-300 ms, a clear offline or slow message, a retry action, no lost form input, and stale data labelled as such.
   - Recovery: when the network returns, the app resumes without a restart and without duplicates.
   - Captive portal and non-JSON responses do not crash or log users out.
4. Automate the highest-value checks: stub or proxy-based tests that inject delay, drop and errors at the network layer (for example route interception in a browser test tool, a fault-injecting proxy in CI, or an injected HTTP client in unit tests), with timeouts and retry counts asserted. Keep real-device manual checks for radio-level behaviour.
5. From the code given, list specific places that look risky (missing timeout, retry on POST, no idempotency key) as findings to verify.

If the description does not say which flows write data or how requests are made, ask for those and stop.
</task>

<constraints>
- Test against test environments and test accounts; never run fault injection against production or real payments.
- Do not invent the app's behaviour; label assumptions [ASSUMED].
- Findings from code are hypotheses until a test confirms them; say so.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Network profiles
Table: profile | parameters | tool on this platform.
## Risky flows
Ranked list with the harm if interrupted.
## Test checklist
Table: flow | profile | check | expected behaviour | manual or automated.
## Automation
Code or config for the top automated checks.
## Findings to check in code
Bullets with file or function, the risk and the test that would confirm it.
</output_format>
````

---

<a id="test-engineer"></a>

## Test engineer

`test-engineer` · persona · Testing · https://hermes-ide.com/prompts/test-engineer

Designs and writes tests that catch real regressions, chooses the cheapest test level that proves a behaviour, and refuses flaky or assertion-free tests. Use as a testing persona or subagent.

````markdown
From now on, work as this persona: Test engineer.

You are a test engineer. You judge a test by one question: would it fail if the behaviour it describes broke? A suite that is green by default proves nothing, so you make sure each test can fail.

How you work:
- You start from behaviour: what the code promises its callers, including errors and limits. You read the code to find the branches, then test through the public interface, not the internals.
- You choose the cheapest level that can prove the behaviour: a unit test before an integration test before an end-to-end test. You go higher only when the risk lives in the wiring.
- You follow the project's existing test conventions, such as framework, layout, naming and fixtures, rather than introducing new ones.
- You watch every new test fail once, by breaking the behaviour or inverting the assertion, before you trust it.
- You treat flakiness as a defect with a cause: time, randomness, ordering, shared state, concurrency or the network.

What you flag:
- Tests that cannot fail: no assertion, assertions on mocks only, `expect(x).toBeTruthy()` where a value is known, snapshots nobody reads.
- Over-mocking: mocks of the code under test or of plain data, and tests that break on every refactor.
- Shared state between tests, order dependence, and real clocks, network or randomness inside unit tests.
- Retries, sleeps and skipped tests used to make a build green.
- Missing boundaries: empty, one, many, maximum, invalid, duplicate, Unicode, time zones, money rounding.

Your habits:
- You name tests after behaviour, so a failure message reads as a sentence about what broke.
- You keep one reason to fail per test and arrange, act and assert in that order.
- You report bugs you find instead of quietly changing production code to make a test pass.
- You report the command you ran and its real result.
````

---

<a id="test-firmware-off-target"></a>

## Test firmware off target

`test-firmware-off-target` · prompt · Testing · https://hermes-ide.com/prompts/test-firmware-off-target

Sets up host-based unit tests for firmware by separating logic from the HAL, faking registers and peripherals, and running in CI. Use when firmware has no automated tests.

````markdown
<context>
The user is an embedded engineer whose firmware is only tested by flashing a board and watching it. Most firmware logic (state machines, protocol parsing, scaling, filtering, retry and timeout rules) does not need hardware at all, and can run as fast unit tests on the build machine, compiled with the host compiler. What blocks that is code that reads and writes registers directly, calls the vendor HAL from inside business logic, uses compiler-specific keywords, or depends on a real tick counter.

Common failures this prompt avoids: mocking every HAL call so tests only restate the implementation; trying to emulate the whole MCU; ignoring host-versus-target differences (int width, endianness, struct packing, `volatile`, alignment) so tests pass on the laptop but not on the chip; and claiming host tests prove timing, interrupts or electrical behaviour.

Toolchain: not given; infer from the code and say what you assumed. Test framework: recommend one that fits the language and build.
</context>

<task>
<firmware_code>
[FIRMWARE_CODE]
</firmware_code>

1. Sort the code into three layers: pure logic (no hardware access), hardware-facing glue (calls into the HAL, drivers, RTOS), and direct register access. Name the functions in each.
2. Propose seams with the least churn: a thin interface (struct of function pointers, link-time substitution of a `.c` file, or a small C++ interface) between logic and hardware. Prefer link-time substitution for C when the code must not change shape; prefer passing a dependency when the module is being touched anyway.
3. Design fakes, not just mocks: a fake GPIO or UART that records writes and lets the test inject reads, a fake register block as a plain struct the code points at in tests, a controllable fake clock or tick source, and a fake for any RTOS call the logic uses (queue, semaphore, delay). Use generated mocks (for example CMock) only to check that a call happened at the boundary.
4. Set up the host build: a separate target compiled with the host compiler, compile flags such as `-Wall -Wextra -Werror` and sanitizers (`-fsanitize=address,undefined`) on host, fixed-width types, a guard for compiler-specific attributes, and the one command to run the tests locally and in CI. Note host-versus-target differences to check for this code.
5. Write the first 4-8 tests for the most valuable logic: boundaries, invalid input, wraparound of counters and ticks, timeouts and error paths. Each test names the behaviour it proves.
6. List what still needs hardware-in-the-loop or on-target tests: interrupt timing and races, DMA, peripheral configuration, power modes, real sensor noise, watchdog and boot behaviour. Suggest the cheapest way to cover each (on-target test runner, logic analyser capture, HIL rig).

If the code is too partial to see where the hardware calls happen, ask for the header or HAL wrapper it uses and stop.
</task>

<constraints>
- Do not change the firmware's behaviour while adding seams; keep refactors minimal and show them as diffs or before-and-after snippets.
- No dynamic allocation added to target code for the sake of tests.
- Do not invent register names, addresses or HAL function signatures; use the ones in the code or mark them [CHECK DATASHEET].
- Never claim a host test proves timing, interrupt safety or electrical behaviour.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Testability assessment
Table: function or module | layer (logic, glue, register) | testable on host now? | blocker.
## Seams and fakes
For each seam: technique, the interface or substitution, and the fake's code.
## Test setup
Directory layout, build target, flags and the commands to run locally and in CI.
## First tests
The test file code, with one comment line per test on what it proves.
## Needs real hardware
Table: behaviour | why host tests cannot prove it | cheapest on-target check.
## Next steps
Up to five, in order.
</output_format>
````

---

<a id="test-game-mechanics-deterministically"></a>

## Test game mechanics deterministically

`test-game-mechanics-deterministically` · prompt · Testing · https://hermes-ide.com/prompts/test-game-mechanics-deterministically

Makes game logic testable with a fixed timestep, seeded randomness and input replays, then writes tests for combat, physics and progression rules. Use when only manual playtests catch regressions.

````markdown
<context>
The user is a game developer whose mechanics are only checked by playing. Engine: infer from the code and say what you assumed. Game logic is hard to test because it is tangled with the frame loop and engine objects, uses variable delta time, calls a global random generator, and reads input devices directly. The same jump then lands in different places on a 30 fps and a 144 fps machine, and a "can't reproduce" crit bug stays unfixed.

The expert approach: pull rules out of engine callbacks into plain code, step simulations with a fixed timestep, inject a seeded random source, feed recorded inputs, and compare results against approved snapshots with tolerances. Common mistakes to avoid: testing exact floats with `==`, snapshotting the whole world so every tweak breaks the tests, testing "feel" values that designers will tune weekly as if they were rules, and running everything as slow play-mode tests when most could be edit-mode or plain unit tests.
</context>

<task>
<game_code>
[GAME_CODE]
</game_code>

1. Separate rules from tuning: list the invariants that must hold whatever the numbers (damage never negative, cooldown cannot be bypassed, a jump always reaches the same apex at any frame rate, loot weights sum to 1) from tuning values designers change. Test invariants and formulas; read tuning values from the same data the game uses.
2. Make it deterministic:
   - Fixed timestep: run simulation in fixed steps (for example 1/60 s) with an accumulator; tests call `Step(dt)` N times instead of waiting for frames.
   - Randomness: one injected, seedable random source per system; no calls to the global generator in game rules.
   - Time and input: an injected clock and an input interface the test can drive, with a recorded input sequence format (frame number, action, value).
3. Show the smallest refactor that achieves this, as before-and-after code, keeping behaviour the same.
4. Write tests at the cheapest level: plain unit tests for formulas and state machines; edit-mode (or headless) tests for components that need engine types; play-mode or scene tests only for behaviour that needs physics or the scene graph. Use the engine's runner (Unity Test Framework, GUT or GdUnit for Godot, Unreal Automation, `cargo test` for Bevy).
5. Cover: boundaries (zero, max level, overflow of stacks), frame-rate independence (same result at 30, 60 and 144 Hz steps within a tolerance), seeded probability (for drop rates, run a large seeded sample and check the observed rate is within a stated tolerance), and state transitions (stun during a dash, death during a cooldown).
6. Add one replay or golden-state test: a recorded input sequence and seed, run for N steps, compare selected state (position, health, score) with an approved snapshot and a float tolerance; explain how to update the snapshot when a change is intended.
7. Say which questions still need humans playing: feel, readability, difficulty, fun.

If the code does not show the formula or the engine callbacks needed, ask for them and stop.
</task>

<constraints>
- Compare floats with an explicit tolerance and say why it was chosen.
- Never call the global random generator or real time in tests.
- Keep snapshots small and named; never snapshot the whole scene.
- Do not invent engine APIs; mark uncertain ones as [CHECK ENGINE DOCS].
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## What to test
Table: rule or invariant | test level (unit, edit-mode, play-mode) | why.
## Making it deterministic
Timestep, random source, clock and input interfaces.
## Refactor
Before-and-after code.
## Tests
Test code, one comment per test on what it proves.
## Replay and golden checks
The replay format, the test, and how to approve an intended change.
## Still needs playtesting
Bullets.
</output_format>
````

---

<a id="test-writing-rules"></a>

## Test-writing rules

`test-writing-rules` · rule · Testing · https://hermes-ide.com/prompts/test-writing-rules

Standing rules for tests an assistant writes, covering behaviour over implementation, no sleeps, deterministic data, mocks only at boundaries and one reason to fail per test.

````markdown
Follow these rules for the rest of this conversation.

Apply these rules to files matching: `**/*.test.*`, `**/*.spec.*`, `**/*_test.*`, `**/test_*.py`.

When you write or change tests in this project:

**What to test**
- Test observable behaviour through the public interface: return values, state others can see, emitted events, HTTP responses, rendered output. Do not assert on private functions, internal call order or intermediate variables.
- Cover the cases that break code: empty input, a single item, boundaries, invalid input, error paths and concurrency where it applies, not just the happy path.
- Every bug fix comes with a test that fails without the fix.

**Shape**
- Each test checks one behaviour and has one reason to fail. Several assertions are fine when they describe the same behaviour.
- Name tests after the behaviour and the condition, such as "returns 404 when the order does not exist", not "test_get_2".
- Structure tests as arrange, act, assert, and set up only the data the test needs, using builders or factories with clear defaults.
- Make assertions specific: exact values, specific error types and messages. Avoid snapshot assertions of large output unless someone reviews the snapshot.

**Determinism**
- Never use sleeps to wait for something. Wait on the condition or event with a timeout, or use the framework's async utilities.
- Control time with a fake clock, randomness with a fixed seed, and time zone and locale explicitly. Never depend on the current date.
- Tests must not depend on execution order or on state left by other tests. Clean up files, records and global state, and give each test its own data.
- No real network calls to third parties in unit tests.

**Test doubles**
- Mock or fake only at the boundaries you do not own or cannot run cheaply: network, clock, file system, third-party services. Do not mock the unit under test or its internal collaborators.
- Prefer simple fakes and stubs to mocks with strict call expectations, which break on harmless refactors.

**Integrity**
- Follow the project's test framework, file layout and helpers. Do not add a new test library without asking.
- Keep unit tests fast, and mark slow or integration tests the way the project does.
- Run the tests you wrote and report the real result. If you could not run them, say so.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
````

---

<a id="unit-test-data-transformations"></a>

## Unit test data transformations

`unit-test-data-transformations` · prompt · Testing · https://hermes-ide.com/prompts/unit-test-data-transformations

Writes small-fixture unit tests for SQL, dbt, pandas or Spark transformations with hand-built rows for nulls, duplicates, late records and boundaries that run in seconds in CI.

````markdown
<context>
The user is a data or analytics engineer whose transformation logic (sql) is only checked by eyeballing dashboards or by data-quality tests on full production tables. Those catch problems after they land and cannot say which rule broke. Unit tests with a handful of hand-written rows can: they pin each business rule, run in seconds, and fail with a readable diff.

What makes these tests worth having: rows chosen to hit one rule each (not a sample of production), explicit expected output rather than re-implementing the logic in the test, and coverage of the traps of data work: NULLs in join keys and aggregates, duplicate keys that fan out joins, late-arriving and out-of-order records, time zone and day boundaries, empty inputs, and type coercion.
</context>

<task>
<transformation_code>
[TRANSFORMATION_CODE]
</transformation_code>

1. State the output grain and list each business rule the code implements (filters, joins, dedup logic, aggregations, window logic, incremental conditions), one line each.
2. For each rule, design the smallest input rows that prove it, plus the edge cases that apply:
   - NULL in a join key, a grouped column and an aggregated value (COUNT(*) versus COUNT(col), SUM of all NULLs).
   - Duplicate keys on either side of a join; ties in window ordering.
   - Late and out-of-order records, and the incremental boundary (a row exactly at the high-water mark).
   - Boundaries: midnight and month ends, time zones and daylight saving changes, zero and negative amounts, empty input.
3. Write the expected output table by hand for each case. If the expected result is ambiguous from the code (for example which duplicate wins), list it under Questions instead of guessing.
4. Write the tests with the engine's tooling:
   - sql: a test that loads fixture rows into temporary tables or CTEs on the same database engine (or DuckDB when the SQL is portable, saying so) and compares the result.
   - dbt: dbt unit tests (`unit_tests:` with `given` and `expect`) in YAML; mention the minimum dbt version they need.
   - pandas: pytest with small DataFrames built inline and `pandas.testing.assert_frame_equal` with explicit dtypes and sorted rows.
   - spark: pytest with a local SparkSession fixture (session scope, few shuffle partitions) and a DataFrame equality helper that ignores row order.
   - other: the nearest equivalent, stated.
5. Make comparisons exact about what matters: column order, dtypes, row order (sort or compare as sets), float tolerance for computed ratios.
6. Give the CI setup: the command, how long it should take (target under a minute), and no dependency on shared warehouses or production data.
</task>

<constraints>
- Fixture rows are invented, small and obviously synthetic; never copy real customer data.
- One rule per test; name each test after the rule and condition.
- Do not re-implement the transformation in the test to compute expected output.
- If table schemas are missing, ask for them or mark assumed columns as [ASSUMED].
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Rules under test
Numbered list.
## Edge cases
Table: rule # | edge case | why it can break.
## Fixtures and expected output
For each test: input rows and expected output as small tables.
## Tests
Code or YAML.
## Running in CI
Command, runtime target and setup.
## Questions
Ambiguous rules the owner must decide, or "None".
</output_format>
````

---

<a id="write-fuzz-harness"></a>

## Write a fuzz harness

`write-fuzz-harness` · prompt · Testing · https://hermes-ide.com/prompts/write-fuzz-harness

Writes a coverage-guided fuzz harness for a parser or decoder with a seed corpus, dictionary, sanitizers, a CI time budget and crash triage. Use for code that reads untrusted input.

````markdown
<context>
The user wants to fuzz code that handles untrusted or complex input. Coverage-guided fuzzing finds crashes, hangs, memory errors and logic bugs that hand-written tests miss, but only with a good harness: one that is fast (thousands of executions per second), deterministic, free of global state between runs, reaches deep code instead of failing at the first checksum or length check, and turns silent bugs into crashes with sanitizers and assertions.

Engines by language ([LANGUAGE]): libFuzzer or AFL++ for C and C++ (with AddressSanitizer and UndefinedBehaviorSanitizer); cargo-fuzz for Rust; native `go test -fuzz` for Go; Atheris for Python; Jazzer for Java. For other languages, name the closest maintained option and say how mature it is.
</context>

<task>
<target_code>
[TARGET_CODE]
</target_code>

1. Define the target and invariants: the entry function, the input it takes, and what must always hold besides "does not crash" (round-trip: decode(encode(x)) == x; parse never returns success with an inconsistent object; output length bounds; two implementations agree).
2. Write the harness:
   - Take the fuzzer's bytes and feed them to the target with minimal setup; for structured input, use the engine's structured helper (for example a data provider or `arbitrary`) rather than hand-parsing bytes.
   - Reset or avoid global state; no file, network or clock access; no randomness the fuzzer does not control.
   - Bound input size and recursion so hangs and out-of-memory reports are meaningful; set a per-input timeout.
   - Assert the invariants so logic bugs crash.
   - Work around blockers that stop deep coverage (checksums, magic numbers, signatures) with a fuzzing build flag, and say so.
3. Build a seed corpus: small valid inputs from tests and examples, one per feature of the format, plus a few edge files (empty, one byte, maximum nesting). Write a dictionary of format tokens (magic bytes, keywords, delimiters).
4. Give the build and run commands with sanitizers, the flags for timeout, memory limit and maximum input length, and how to read the coverage and executions-per-second output.
5. Set a CI budget: a short run (for example 5-10 minutes) on pull requests that touch the target, longer runs on a schedule, corpus kept as an artefact between runs, and regressions: every crash input added to the corpus and to a normal unit test.
6. Explain triage: reproduce with the crash file, minimise it (the engine's minimise mode), deduplicate by stack, classify (out-of-bounds, use-after-free, overflow, assertion, timeout, out-of-memory), fix, and add the regression test. Mention continuous fuzzing services suitable for open-source projects as an option.

If the target code's input contract is unclear, ask what a valid input looks like and stop.
</task>

<constraints>
- The harness must be deterministic and must not write files or open sockets.
- Do not invent APIs of the user's code; mark assumptions [ASSUMED].
- Never fuzz production services or third-party systems; harnesses run locally or in CI.
- Do not claim the code is safe because a short run found nothing; state what coverage and duration were reached.
</constraints>

<output_format>
## Target and invariants
Bullets.
## Harness
Code.
## Corpus and dictionary
Seed file list with what each exercises, and the dictionary file.
## Build and run
Commands with flags explained in one line each.
## CI budget
Pull-request and scheduled jobs, durations, corpus storage.
## Triage
Numbered steps.
</output_format>
````

---

<a id="write-e2e-test"></a>

## Write a resilient end-to-end test

`write-e2e-test` · prompt · Testing · https://hermes-ide.com/prompts/write-e2e-test

Writes an end-to-end browser test for a user flow with role-based locators, auto-waiting assertions and isolated test data, never fixed sleeps. Use when adding UI coverage for a critical path.

````markdown
<context>
End-to-end tests are the most expensive tests to keep green. They become flaky when they locate elements by CSS structure or generated class names, wait with fixed sleeps, share data between runs, or assert on things a user never sees. A resilient test finds elements the way a user or assistive technology does (role and accessible name, label, visible text), waits on conditions instead of time, owns its data, and checks the outcome the user cares about.
</context>

<task>
Write a playwright test for this flow:
[FLOW]

1. Restate the flow as numbered user actions, each with the observable outcome that proves it worked. If a step's expected outcome is not stated, ask for it or mark your assumption.
2. If you have the repository, read the relevant pages or components and any existing e2e setup (config, fixtures, page objects, auth helpers, test-data factories) and reuse them. Match the existing style. If the project already uses a different end-to-end framework than playwright, say so and ask which to use before writing.
3. Locators, in this order of preference:
   - Playwright: `getByRole` with name, then `getByLabel`, `getByPlaceholder`, `getByText`, then `getByTestId` as a last resort.
   - Cypress: Testing Library queries (`findByRole`, `findByLabelText`) if the project has them, otherwise `cy.contains` scoped to a container, then `data-testid`/`data-cy`.
   - Selenium: accessible attributes, labels and visible text via stable XPath or CSS on `data-testid`; never absolute XPath.
   Never use generated class names, nth-child chains or DOM position.
4. Waiting: use auto-retrying, web-first assertions (Playwright `expect(locator).toBeVisible()`/`toHaveText()`, Cypress `should`, Selenium `WebDriverWait` with expected conditions). Wait for a specific network response or UI state when an action triggers one. No `waitForTimeout`, `cy.wait(<ms>)` or `Thread.sleep`.
5. Isolation: create the data the test needs through an API, fixture or seed helper, with unique values per run, and clean it up or make it disposable. Log in through a stored session or API helper rather than the login form, unless login is the flow under test.
6. Assert the user-visible outcome at each checkpoint, plus one durable side effect if it matters (the saved record, the confirmation email stub), not implementation details.
</task>

<constraints>
- Do not invent selectors, routes or accessible names you have not seen. When the page source is not available, write the most likely role and name, and list each one under "Assumptions to verify".
- One flow per test. Keep the test independent of test order.
- If a step depends on a third-party service (payments, email, maps), stub it at the network layer and say so.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
</constraints>

<output_format>
## Test plan
Numbered steps: action, then expected outcome.
## Test
The complete test file in one code block, including setup and teardown helpers it needs.
## Assumptions to verify
Bullets: each selector, route or data assumption you could not confirm. Or "None".
## How to run
The command to run this one test headed and headless, and how to see the trace or screenshots on failure.
</output_format>
````

---

<a id="write-test-plan"></a>

## Write a test plan

`write-test-plan` · prompt · Testing · https://hermes-ide.com/prompts/write-test-plan

Writes a risk-based test plan for a feature or release covering scope, risks, test levels, environments, data, manual checks automation misses and exit criteria. Use before testing a release.

````markdown
<context>
A test plan is useful when it tells a team where to spend limited testing time and when to stop. Plans that list every possible test case get skimmed and ignored; plans with no risk analysis spread effort evenly, so the payment edge case gets the same attention as a label change. A good plan ranks risks, picks the cheapest test level that addresses each one, names what automation will not catch (usability, unusual data, real devices, integrations with real third parties, migration of existing data), and defines exit criteria that someone can actually check on release day.
</context>

<task>
Write a test plan for:
[FEATURE]

1. Define the scope: what is being tested (functions, platforms, user types, integrations) and what is explicitly out of scope, with the reason.
2. Identify risks: combine the known worries with what the feature implies (money, permissions, data migration, concurrency, third parties, performance, accessibility, localisation, backward compatibility, feature-flag states). Rate each by likelihood and impact, and rank them.
3. For each top risk, choose the test level that addresses it most cheaply (unit, integration, contract, end-to-end, manual exploratory, non-functional), say whether existing automation already covers it, and what new tests are needed. Name the gaps automation will not close.
4. Write exploratory charters for the manual work, in the form "Explore <area> with <resources or data> to discover <kind of problem>", each time-boxed, covering what scripted tests miss: unexpected sequences, interrupted flows, odd data, permissions, different devices and assistive technology.
5. Specify environments and test data: which environment, which configuration and feature-flag states, accounts and roles needed, data volume and edge records, third-party sandboxes, and how data is created and reset. No real personal data.
6. Define entry criteria (what must be true before testing starts) and exit criteria that are checkable: no open critical or high defects, the named risks covered, automated suites green, performance within stated limits, and an explicit decision on known issues. Include a rollback or flag-off check if the release can be reverted.
7. Lay out the schedule against the release date, with owners as roles, and say what to cut first if time runs short (lowest-ranked risks), so the trade-off is visible rather than accidental.

If the feature description is too thin to identify risks (no behaviour, users or integrations), ask for the spec or acceptance criteria and stop.
</task>

<constraints>
- Rank everything by risk. Do not list low-value test cases to look thorough.
- Do not duplicate what existing automation already covers; reference it instead.
- Exit criteria must be measurable or a named decision, never "sufficient testing done".
- Do not invent dates, people or metrics; use roles and placeholders.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Scope
In scope and out of scope, as two short lists.

## Risks
Table: # | risk | likelihood | impact | priority.

## Test approach
Table: risk # | test level | covered by existing automation? | new tests needed.

## Environments and data
Bullets.

## Manual and exploratory testing
Numbered charters with time boxes, plus any must-do manual checks.

## Entry and exit criteria
Two checklists.

## Schedule and owners
Table: activity | owner (role) | when. Then "If time runs short, cut:" in priority order.

## Open questions
Only the ones that change the plan.
</output_format>
````

---

<a id="write-contract-tests"></a>

## Write consumer-driven contract tests

`write-contract-tests` · prompt · Testing · https://hermes-ide.com/prompts/write-contract-tests

Writes consumer-driven contract tests between two services and the CI gate that runs them, so a breaking API change fails before deploy. Use when services that call each other ship independently.

````markdown
<context>
Contract tests catch the integration bug where each service passes its own tests but the provider renames a field, tightens validation or changes a status code that a consumer depends on. In consumer-driven contracts the consumer records only what it actually sends and reads, so the provider is free to change everything else. Contracts that copy whole responses with exact values are brittle and block harmless changes; contracts that are never verified against the real provider, or never gate a deploy, catch nothing.
</context>

<task>
Write contract tests using the pact approach.

Consumer:
[CONSUMER]

Provider API:
[PROVIDER_API]

1. List each interaction the consumer really uses: method, path, query, headers that matter, request body fields, the response status and only the response fields the consumer reads. Trace field reads in the consumer code; do not include fields it ignores. Include the error responses the consumer handles (404, 409, 422 and so on).
2. For each interaction, name the provider state it needs ("order 42 exists and is paid").
3. Consumer side: write tests that exercise the real client code against the contract mock and check the client's own parsing, not just the mock. Use type and format matchers (like-type, regex, each-like with a minimum) instead of literal values, except where the exact value is the contract (enums, status codes).
4. Provider side: write the verification that replays the contract against the running provider, with a state handler per provider state that sets up data through the provider's own code or test fixtures.
5. CI: show how the contract is published (with consumer version and branch), how the provider verifies on every build, and the pre-deploy check that blocks a deploy when the deployed counterpart's contract is not verified (for Pact, a broker or PactFlow with `can-i-deploy --to-environment`, `record-deployment` after each deploy so the broker knows what runs where, and a contract-changed webhook that triggers provider verification). For openapi-schema, validate consumer mocks against the spec and provider responses against the spec, and fail on spec drift. For custom, store fixtures in one place both builds read, and version them.
6. Show one concrete breaking change (a renamed field, say) and which check fails.
</task>

<constraints>
- Do not invent endpoints, fields or status codes that are not in the inputs. If the provider API and the consumer disagree, report the mismatch as a finding instead of choosing one.
- Contracts are not functional tests: do not assert business rules of the provider beyond the shape and semantics the consumer relies on.
- Use the client library and test runner the consumer already uses. Name any package to install with its ecosystem.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
</constraints>

<output_format>
## Interactions
Table: Interaction | Request | Response fields used | Provider state.
## Consumer tests
Complete test file(s) in code blocks.
## Provider verification
Complete verification test with state handlers.
## CI gate
The pipeline steps (as config or a numbered list) for publish, verify and the pre-deploy check, and the breaking-change example.
## What this does not catch
Bullets: behaviour contract tests miss here (performance, auth flows, data semantics) and what covers it instead. Mismatches found between consumer and provider go first, if any.
</output_format>
````

---

<a id="write-exploratory-test-charters"></a>

## Write exploratory test charters

`write-exploratory-test-charters` · prompt · Testing · https://hermes-ide.com/prompts/write-exploratory-test-charters

Writes session-based exploratory testing charters for a feature with time boxes, heuristics, a session sheet and a debrief agenda, aimed at risks automation misses. Use before a release.

````markdown
<context>
The user wants structured exploratory testing for one feature: time-boxed sessions, each guided by a charter, with notes good enough to debrief and report bugs. Session-based exploratory testing works because it is focused but not scripted: a charter says where to look and what kind of problem to look for, and the tester follows what they learn.

Weak charters are either test cases in disguise ("verify the button saves") or so broad they guide nothing ("test the checkout"). Good ones use the form "Explore <target> with <resources> to discover <information>", aim at risks automated checks are poor at (odd sequences, interruptions, real data shapes, permissions, concurrency, error recovery, usability on real devices), and fit a 45-90 minute session.

Time available: about 4 hours of one tester.
</context>

<task>
<feature>
[FEATURE_DESCRIPTION]
</feature>

1. Map the product with the SFDIPOT heuristic (Structure, Function, Data, Interfaces, Platform, Operations, Time): note for each dimension what this feature has and what could go wrong. Combine with the known risks and rank the top areas by impact and likelihood. Skip what automation already covers well.
2. Write 4-8 charters, highest risk first. Each has:
   - The charter line: "Explore <target> with <resources: data, accounts, devices, tools> to discover <kind of problem>".
   - Time box (short 45, normal 60 or long 90 minutes).
   - Setup needed (accounts, data, feature flags, environment).
   - Two to four heuristics or tours to try, chosen for the target: boundaries and zero-one-many, CRUD on each object, interruptions (back button, network loss, app backgrounded, session timeout), concurrency (two tabs, two users), data variety (long, Unicode, right-to-left, emoji, empty, pasted), undo and recovery, permissions and roles, the "follow the data" tour across screens and exports.
   - Oracles: how the tester will recognise a problem (spec, comparable product, consistency with the rest of the app, user expectations, error messages, data in the database or export).
3. Fit the charters into the time available; say which to drop first if time runs short.
4. Provide a session sheet template with: charter, tester, start time, duration, percentage split of time on testing, bug investigation and setup, test notes, bugs (title, steps, expected, actual, evidence), issues and questions, and coverage notes.
5. Provide a short debrief agenda (10-15 minutes per session): what was covered, what was not and why, bugs and their severity, new risks found, and whether a follow-up charter is needed.

If the feature description lacks who uses it or what it does, ask for that and stop.
</task>

<constraints>
- Charters guide; they never contain step-by-step scripts or expected results per step.
- Do not repeat checks the user says automation covers; reference them instead.
- Use synthetic test data and test accounts; never real personal data.
- Do not invent features; mark assumptions as [ASSUMED].
</constraints>

<output_format>
## Risk map
Table: SFDIPOT dimension | what this feature has | what could go wrong | priority.
## Charters
Numbered charters with the fields from step 2.
## Session schedule
Table: session | charter | tester (role) | time box. Then "If time runs short, drop:".
## Session sheet
The template as a fenced Markdown block, ready to copy.
## Debrief agenda
Bullets with minutes.
</output_format>
````

---

<a id="write-gherkin-scenarios"></a>

## Write Gherkin scenarios

`write-gherkin-scenarios` · prompt · Testing · https://hermes-ide.com/prompts/write-gherkin-scenarios

Turns acceptance criteria into declarative Given/When/Then scenarios with one behaviour each, scenario outlines for data variants and business language that survives UI changes. Use in BDD teams.

````markdown
<context>
The user works in a team that uses Gherkin (Cucumber, SpecFlow or Reqnroll, Behave, Behat or similar) and wants scenarios that serve as living documentation and automated acceptance tests. Scenarios go wrong in familiar ways: imperative UI scripts ("When I click the 'Submit' button") that break with every redesign; several behaviours in one scenario; incidental detail that hides the rule; Given steps that perform actions; Then steps that check implementation details; and invented rules nobody agreed.

Good scenarios are declarative ("When the member renews with an expired card"), use the business's own words, show one rule with a concrete example each, and expose gaps in the criteria as questions instead of filling them silently.
</context>

<task>
<acceptance_criteria>
[ACCEPTANCE_CRITERIA]
</acceptance_criteria>

1. Extract the business rules (one line each) from the criteria, and for each rule the examples that illustrate it: the main example, boundary examples and the counter-example where the rule does not apply.
2. Write one `Feature` with a short description of the value (As a / I want / So that only if it adds meaning). Group scenarios under `Rule:` keywords, one per business rule.
3. For each example, write a scenario:
   - Title states the behaviour and condition ("Renewal is refused when the card has expired").
   - Given: state only, in past or present tense, no UI actions. When: one business action. Then: an observable business outcome, not database rows or HTTP codes, unless the audience is an API consumer.
   - Three to seven steps; use `And` sparingly; no conjunction steps ("When I log in and add an item").
   - Include only the data the rule depends on; push the rest into step definitions or defaults.
4. Use `Scenario Outline` with `Examples` only when the same behaviour varies by data (for example price bands); keep tables narrow with column names in business terms. Use a `Background` only for Given steps shared by every scenario in the feature, at most three lines.
5. Reuse existing step phrases from the domain terms where they fit; list the steps the team must implement, with parameter types.
6. List questions where the criteria are silent or contradictory (what happens at exactly the limit, which role can do this, what the user sees on failure). Do not encode an answer for them; mark the affected scenario with a `@question` tag.
</task>

<constraints>
- No UI element names, CSS selectors, URLs or waits in steps.
- One behaviour per scenario; no scenario longer than seven steps.
- Do not invent business rules; anything not in the criteria becomes a question.
- Valid Gherkin syntax that the common runners parse.
</constraints>

<output_format>
## Rules found
Numbered list.
## Feature file
One fenced `gherkin` block.
## Step vocabulary
Table: step phrase | type (Given, When, Then) | parameters | new or existing.
## Questions for the three amigos
Bullets, each naming the rule and scenario it affects.
</output_format>
````

---

<a id="write-integration-tests"></a>

## Write integration tests with real dependencies

`write-integration-tests` · prompt · Testing · https://hermes-ide.com/prompts/write-integration-tests

Writes integration tests that run against real dependencies such as databases and queues in containers, with fixtures, isolation between tests and cleanup. Use when mocks hide bugs at the boundary.

````markdown
<context>
Integration tests exist to catch what mocks cannot: SQL that only fails on the real engine, transaction and locking behaviour, migrations, serialisation across a queue, unique constraints, time zones and encodings. They become a burden when they share state and fail in random order, sleep instead of waiting, start a fresh container per test and take twenty minutes, or test the dependency rather than the code. Good integration tests start each dependency once per run, give every test its own data, wait on conditions, and assert on observable outcomes.
</context>

<task>
Write integration tests for:
<code>
[CODE]
</code>

1. Read the code and the project's existing test setup (framework, runner, folders, helpers, migrations, CI config). Follow what exists. If you cannot see the code or the dependency versions, ask once for what is missing and stop.
2. Write a short test plan: the behaviours that cross a real boundary (queries with filtering and ordering, constraint violations, transactions and rollbacks, concurrent updates, message publish and consume, retries and dead-lettering, cache expiry), each with the outcome to assert. Leave pure logic to unit tests.
3. Set up dependencies in containers, preferring the Testcontainers library for the language, or a compose file the test run starts. Pin image versions to match production. Start each container once per test run or suite, not per test. Apply the real schema migrations, not a hand-written schema.
4. Isolate tests. Pick the cheapest strategy that is correct and say why: a transaction per test rolled back at the end (not valid when the code under test commits or uses several connections), a unique schema, database, queue or key prefix per test or worker, or truncating tables between tests. Make tests safe to run in parallel or mark them serial.
5. Build data with small factories or builders that set only the fields a test cares about. No shared mutable fixtures.
6. Wait on conditions with a timeout (poll until the message is consumed, up to a few seconds); never fixed sleeps.
7. Clean up containers, connections and temporary resources even when a test fails.
8. Run the tests and report the real result. If you cannot run them (no container runtime), say so plainly.
</task>

<constraints>
- Use the real dependency for the behaviour under test; mock only external third parties you do not control, and say which.
- Never point tests at a shared or production environment, and never read real credentials. Use container-generated connection settings.
- Assert on outcomes (rows, messages, responses), not on which internal functions were called.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Test plan
A table: behaviour, dependency, assertion.
## Setup
The container or compose setup and shared fixtures, as code blocks with file paths.
## Tests
The test files, as code blocks with file paths.
## How to run
Commands for local runs and the CI job change, plus the result of running them.
## Notes
Isolation strategy chosen and why, expected runtime, and anything that could make the tests flaky.
</output_format>
````

---

<a id="write-manual-test-cases"></a>

## Write manual test cases

`write-manual-test-cases` · prompt · Testing · https://hermes-ide.com/prompts/write-manual-test-cases

Writes manual test cases a non-developer can run, with preconditions, numbered steps, test data, expected results and priority, grouped into smoke and full passes. Use for UAT or before automation.

````markdown
<context>
The cases will be run by people who did not build the feature: manual testers, support staff or product owners doing user acceptance testing. They need to follow each case without guessing and to know for certain whether it passed. Common failures: steps that say "check it works", expected results that are vague ("page loads correctly"), several checks hidden in one case so a failure is hard to report, missing test data, and fifty equal-priority cases when the team only has time for ten.

Layout requested: table.
</context>

<task>
<feature>
[FEATURE_DESCRIPTION]
</feature>

1. Identify user roles, the main flows and the rules in the acceptance criteria. List assumptions you had to make.
2. Derive cases: each acceptance criterion's main path; then invalid input and error messages; boundaries (limits, dates, amounts, empty and maximum lengths); permissions per role; cancel, back and retry; and what happens to existing data. Merge cases that would test the same thing.
3. Write each case with:
   - ID (for example TC-01), a short title starting with a verb ("Reject a discount code that has expired").
   - Priority: P1 (must pass to release), P2 (important), P3 (nice to check).
   - Preconditions: account, role, starting page, data that must exist.
   - Steps: numbered, one action each, in plain words naming the exact button or field as it appears on screen.
   - Test data: exact values to enter.
   - Expected result: what the tester should see, specific enough to be true or false (the exact message, the new total, the status shown).
   - Result column left blank (Pass, Fail, Blocked) and a Notes column.
4. Group into a smoke pass (P1 only, about 15-30 minutes, run on every build) and a full pass (everything, ordered so setup is reused).
5. Describe the test data to prepare: accounts per role, records in particular states, and how to reset them. Synthetic only.
6. Lay out the cases in the requested layout: table uses the columns ID, Title, Priority, Preconditions, Steps, Test data, Expected result, Result, Notes, with steps as a numbered list inside the cell (`1. ... <br> 2. ...`); markdown-list gives each case a short heading and the same fields as labelled bullets; csv uses the same columns in that order, one fenced `csv` block per pass, with every cell quoted and steps separated by line breaks inside the quoted cell.
7. Keep it runnable: aim for 8-30 cases in total and a smoke pass of no more than 10. If the feature needs more, cover the highest-risk areas and list the areas left out under Scope and assumptions.

If the description does not say what the feature does or who uses it (for example only a feature name), do not write cases: ask for what it does, the user roles, the acceptance criteria and any messages or limits, and stop. If it says what the feature does but lacks acceptance criteria, write the cases you can, mark unclear expected results as [CONFIRM], and list the questions.
</task>

<constraints>
- One check per case; split cases that verify unrelated outcomes.
- No developer jargon in steps; name what the tester sees.
- Expected results must be observable on screen or in an email, export or report the tester can access.
- Do not invent rules, messages or limits; mark them [CONFIRM].
</constraints>

<output_format>
## Scope and assumptions
Bullets.
## Test data
Table: data item | state | how to create or reset.
## Smoke pass
The P1 cases in the requested layout.
## Full pass
All remaining cases in the requested layout.
## Questions
Bullets for anything marked [CONFIRM], or "None".
</output_format>
````

---

<a id="write-native-ui-tests"></a>

## Write native mobile UI tests

`write-native-ui-tests` · prompt · Testing · https://hermes-ide.com/prompts/write-native-ui-tests

Writes UI tests for an iOS, Android, React Native or Flutter screen with stable identifiers, condition waits instead of sleeps, seeded launch state, network stubs and screen objects.

````markdown
<context>
The user is a mobile engineer or QA engineer adding UI tests for one screen and flow on [PLATFORM]. Mobile UI suites rot for predictable reasons: selectors bound to visible text or view hierarchy that change with copy and layout, fixed sleeps that are too short on CI emulators and too long everywhere else, tests that depend on a real backend and on state left by the previous test, and system dialogs (permissions, notifications, keyboard) that appear on one device and not another.

Use the platform's own tools: XCUITest for iOS; Espresso (with Compose testing APIs for Compose) for Android; Detox or Maestro for React Native; `integration_test` with `flutter_test` finders for Flutter. Name another option only if the code shows the team already uses it.
</context>

<task>
<screen_code>
[SCREEN_CODE]
</screen_code>

<flow>
[FLOW]
</flow>

1. Plan the cases: the main path of the flow, then each state the screen can show (loading, empty, error, offline, success), and one input validation case. Keep it to the 3-7 tests that would catch real regressions.
2. Add stable identifiers: `accessibilityIdentifier` (iOS), `testTag` or resource IDs (Android), `testID` (React Native), `Key` values (Flutter). Do not select by display text unless the text itself is the behaviour under test; never by index or deep hierarchy.
3. Replace every wait with a condition: XCTest expectations or `waitForExistence(timeout:)`, Espresso idling resources or Compose `waitUntil`, Detox `waitFor().toBeVisible().withTimeout()`, Flutter `pumpAndSettle` or pumping until a finder matches (with a cap). No `sleep`, `Thread.sleep` or fixed delays.
4. Add test-only hooks: launch arguments or environment (`launchArguments`, instrumentation arguments, build flavour, `--dart-define`) that reset storage, log in a seeded user, set locale and time zone, disable animations and point the network layer at a stub (local stub server or injected fake client) with fixtures per state. Keep hooks out of release builds.
5. Handle system interruptions explicitly: pre-grant permissions where the tool allows it, otherwise an interruption handler; dismiss the keyboard deliberately.
6. Write a thin screen-object layer: one object per screen exposing user actions and assertions in domain words (`login.submit(email:password:)`), hiding identifiers.
7. Write the tests: one behaviour each, independent and runnable in any order, each starting from a known launch state.
8. Give the CI command and settings: device or emulator image and OS version, animations off, retries off by default, screenshots or recordings and logs kept on failure.

If the code does not show how the screen gets its data, ask how the network layer is created (so it can be stubbed) and stop.
</task>

<constraints>
- No fixed sleeps anywhere. No real production backend.
- Every identifier the tests use must be added in the screen code shown; list each change.
- Tests must not depend on order or on another test's state.
- Do not invent APIs of the user's app; mark assumed names as [ASSUMED].
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Test plan
Table: test | state or path | why it matters.
## Identifiers to add
Code changes to the screen, as a diff or snippets.
## Test hooks
Launch arguments, stub setup and fixtures, and how they stay out of release builds.
## Screen objects
Code.
## Tests
Code.
## Running in CI
Command, device or emulator, settings and failure artefacts.
</output_format>
````

---

<a id="write-property-based-tests"></a>

## Write property-based tests

`write-property-based-tests` · prompt · Testing · https://hermes-ide.com/prompts/write-property-based-tests

Finds the invariants a function must keep and writes property-based tests with generators that shrink well. Use when example-based tests miss edge cases in parsers, encoders or pure logic.

````markdown
<context>
Property-based tests state a rule that must hold for every valid input and let a generator search for a counterexample, then shrink it to the smallest failing case. They find the bugs example tests miss, but only when the property is genuinely true of the specification (not a restatement of the implementation) and the generators produce valid, varied, shrinkable inputs. A property that re-implements the function proves nothing; a generator that filters away 90% of its draws is slow and shrinks badly.
</context>

<task>
Write property-based tests for:
[CODE]

Library:  If no library is named, detect it from the project's manifests and existing tests (Hypothesis for Python, fast-check for JavaScript and TypeScript, proptest for Rust, jqwik for Java, FsCheck for .NET, rapid for Go; the standard library's testing/quick is frozen and shrinks nothing). If none is installed, pick the standard one for the language and say how to add it.

1. Read the code and state its contract: valid inputs, outputs, errors it may raise, and side effects. If the contract is ambiguous (for example, what happens on empty input), ask or state the assumption you test against.
2. Find candidate properties, preferring these patterns:
   - round-trip: decode(encode(x)) == x, parse(print(x)) == x;
   - invariants: output is sorted, length preserved, total conserved, no duplicates, within bounds;
   - idempotence: f(f(x)) == f(x);
   - oracle or model: agrees with a simpler, obviously correct implementation or an in-memory model of a stateful system;
   - metamorphic: a known change to the input causes a predictable change to the output;
   - algebraic: commutativity, associativity, identity elements where the domain promises them;
   - robustness: never crashes or hangs on any input of the right type, and fails only with documented errors.
   Keep only properties that follow from the contract. Discard any that just mirror the implementation.
3. Build generators from the domain, not from raw types: construct valid values directly (map, compose, build strategies) instead of generating anything and filtering. Include the edge values the type allows: empty, single element, zero, negative, maximum sizes, Unicode beyond ASCII, NaN and infinities for floats where relevant. Bound sizes so a run stays fast.
4. Write the tests in the project's style and test runner. Make failures reproducible: rely on the library's seed reporting and example database or replay, and add any shrunk counterexample you discover as an explicit regression example.
5. If you can run the tests, do so and report the result. If a property fails, report the minimal counterexample and whether the bug is in the code or in your property. Do not change the code under test.
</task>

<constraints>
- Every property must name the contract clause it checks. No property may call the function under test to compute its own expected value.
- Avoid filter or assume calls that reject more than a small fraction of draws; restructure the generator instead.
- Keep default example counts unless there is a reason to change them, and say why if you do.
- Do not fix bugs you find; report them.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
</constraints>

<output_format>
## Properties
Table: Property | Pattern | Contract clause it checks.
## Generators
One line per generator: what it builds and which edge values it covers.
## Tests
The complete test file in one code block, with imports.
## Counterexamples
Shrunk failing inputs with a one-line diagnosis each, or "None found" with the number of examples run. If you could not run the tests, say so.
## How to run
The exact command, including how to replay a failure from its seed.
</output_format>
````

---

<a id="write-test-data-factories"></a>

## Write test data factories

`write-test-data-factories` · prompt · Testing · https://hermes-ide.com/prompts/write-test-data-factories

Writes test data builders or factories that produce valid objects by default and take short, readable overrides. Use when tests are full of copy-pasted fixtures that break on every model change.

````markdown
<context>
Literal fixtures copied between tests fail in three ways: a new required field breaks dozens of tests at once, the reader cannot tell which of twenty fields matters to the test, and unique fields collide when two tests insert the same email. A good factory returns a minimal object that passes every validation and constraint with no arguments, and lets a test state only the fields it cares about. Each language has an idiom for this: factory_boy in Python, factory_bot in Ruby, Fishery or plain builder functions in TypeScript, the Test Data Builder pattern (`aUser().withEmail(...).build()`) in Java, Kotlin and C#, and functional options in Go.
</context>

<task>
Write factories for these models in [LANGUAGE]:
<models>
[MODELS]
</models>

1. If a field's type, validation or relation is unclear and guessing would produce invalid objects, list it under Open questions and use the most restrictive reasonable reading. If you can read the repository, find the existing test helpers and any factory library first, and follow them; do not add a second library.
2. For each model, define defaults that satisfy every validation, NOT NULL and check constraint, and nothing more. Optional fields stay empty by default unless most tests need them.
3. Unique fields use a sequence (`user-1@example.test`, `user-2@...`), never random values. If fake data libraries are used, seed them once so failures reproduce.
4. Named states become traits or named builders (`suspended`, `paid`, `expired`), each setting the full group of fields that state requires, so a test never sets `status` without the matching `cancelled_at`.
5. Required relations build the smallest valid parent automatically and accept an existing parent instead. Never create more related records than the constraint requires.
6. Separate building in memory from persisting (`build` and `create`, or a builder plus a save helper). Building is the default because it is faster.
7. Dates and times are relative to a fixed reference or the test's fake clock, never the real current time.
8. Show two or three existing-style tests rewritten with the factories so the reader sees the override style.
9. Add one test that builds and, where there is persistence, saves every factory and trait with no overrides and checks it is valid. This is what catches a factory that drifted after a model change.
</task>

<constraints>
- Defaults must not encode business assumptions a test might rely on silently. If a test depends on a value, the test sets it.
- No mutable defaults shared between instances (a list or dict default must be built fresh each time).
- Use obvious fake values (`example.test` domains, `555` phone numbers); never real-looking personal data.
- Keep factories next to the tests in the project's existing layout.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Design
Bullets: library or pattern chosen and why, build versus create, how uniqueness and time are handled.
## Factories
Code blocks with file paths.
## Usage
The rewritten tests, each showing only the fields that matter.
## Validity test
Code block with file path.
## Open questions
Numbered, or "None".
</output_format>
````

---

<a id="write-unit-tests"></a>

## Write unit tests

`write-unit-tests` · prompt · Testing · https://hermes-ide.com/prompts/write-unit-tests

Writes unit tests that pin a unit's behaviour, covering boundaries, errors and edge inputs in the project's own test style, and proves each test can fail. Use for new or untested code.

````markdown
<context>
Good unit tests describe what a unit does, not how it does it. They fail when behaviour breaks and keep passing through refactors. Tests that mirror the implementation, mock everything, or assert only that no exception was thrown add maintenance cost without catching bugs.
</context>

<task>
Write unit tests for [TARGET].
1. Read the target and its callers to learn its contract: inputs, outputs, side effects, errors. Read two or three existing test files to learn the project's conventions (framework, file location, naming, fixtures, assertion style) and follow them.
2. List the behaviours to cover before writing any test:
   - the main cases;
   - boundaries: empty, one element, maximum, zero, negative, off-by-one limits;
   - invalid input and every error path the code defines;
   - inputs that often break code: null or missing values, duplicates, Unicode, very large values, time zones and dates, floating-point amounts.
3. Write one test per behaviour, through the unit's public interface. Name each test after the behaviour (`returns empty list when no orders match`), not after the method.
4. Use fakes or mocks only at real boundaries: network, clock, file system, randomness, other services. Do not mock the code under test or plain data objects.
5. Run the tests. For each new test, confirm it can fail: break the behaviour temporarily or invert the assertion, watch it fail, then restore it.
</task>

<constraints>
- Do not change production code. If the code is hard to test, or you find a bug, report it under "Not covered" with the failing input and leave the code alone.
- Each test asserts specific values, not only that something is truthy or that no error was thrown.
- Keep tests independent: no shared mutable state and no dependence on run order.
- No snapshot tests unless the project already uses them for this kind of output.
- Fix the behaviour, not the test. Never special-case test inputs, weaken assertions or skip tests to make a check pass.
- If a test looks wrong, explain why and ask before changing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Behaviours
A table: Behaviour | Test name | Kind (main, boundary, error, edge).
## Tests
The new or changed test files as a diff.
## Run
The command you ran and its result, plus how you confirmed the tests can fail.
## Not covered
Behaviours you did not test and why, and any bugs found (input, expected, actual). Or "None".
</output_format>
````

---

<a id="write-visual-regression-tests"></a>

## Write visual regression tests

`write-visual-regression-tests` · prompt · Testing · https://hermes-ide.com/prompts/write-visual-regression-tests

Writes visual regression tests for UI components with deterministic snapshots, tight diff thresholds, a state matrix and CI setup. Use when CSS changes keep breaking screens nobody rechecked.

````markdown
<context>
Visual tests fail for two reasons: the UI really changed, or the screenshot is not deterministic. The second kind kills suites, because once a team learns to click "update all baselines", real regressions get approved too. Nondeterminism comes from fonts rendering differently across operating systems, animations and carets caught mid-frame, dates, random or remote data, lazy images, scrollbars and viewport size. A suite that lasts captures the states that matter, removes every source of noise before setting a threshold, and makes baseline updates a reviewed change.
</context>

<task>
Write visual regression tests for:
<components>
[COMPONENTS]
</components>


1. If you can read the repository, find how components are rendered in isolation (Storybook, a test harness, routes) and any existing visual setup, and build on it. If no tool is given, recommend one from the stack in one sentence: Playwright `toHaveScreenshot` when Playwright is present, the Storybook test runner or a hosted service when stories already exist.
2. Build a state matrix: each component by its meaningful states, plus viewport widths (one narrow, one wide unless told otherwise) and themes the product supports (light, dark, right-to-left). Cap the matrix at what someone will actually review; explain what you left out.
3. Stabilise before snapshotting:
   - fix viewport and device scale factor;
   - wait for web fonts (`document.fonts.ready`) and images to load;
   - disable animations and transitions and hide the text caret;
   - freeze time and seed or mock data and network responses;
   - mask or hide regions that are legitimately dynamic (avatars from a CDN, timestamps, ads), and say what each mask covers.
4. Prefer component-level screenshots of the element over full pages; take full pages only for layout-level checks.
5. Set the diff threshold last, small and explicit (for example a max diff pixel ratio around 0.01), and explain that a larger threshold hides regressions.
6. Configure CI to render in one pinned environment (the same container image locally and in CI), so font rendering matches, and to upload the diff images as artifacts when a test fails.
7. Describe the baseline workflow: baselines are generated in that same environment, updated only in a commit that reviewers can see, and never updated in bulk to make CI green.
</task>

<constraints>
- Do not snapshot states you cannot make deterministic; list them as manual checks instead.
- Do not fold functional assertions into visual tests; keep behaviour checks in the existing unit or end-to-end suites.
- Name screenshots after component, state, viewport and theme so a failing diff explains itself.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Approach
Tool, where tests live and why, in three or four bullets.
## State matrix
Table: component, states, viewports, themes.
## Tests
Code blocks with file paths.
## Stabilisation
Bullets: each noise source and how it is removed.
## CI
The CI job or config, with the pinned image.
## Baseline workflow
Numbered steps for creating, reviewing and updating baselines.
</output_format>
````
