# Hodios paste pack: Performance

Everything in Performance from Hodios, the open prompt library by Hermes IDE: 23 entries, catalog 2026.1004.3.

Every entry is dedicated to the public domain under CC0 1.0. Copy, change and share them freely, no attribution needed.

Browse and search the library at https://hermes-ide.com/prompts

## How to use

Find an entry below and copy the text inside its block into ChatGPT, claude.ai or any chat. Replace each [PLACEHOLDER] with your own material. Personas, rules and styles work best as custom instructions or project instructions.

## Contents

- Performance
  - [Analyse load test results](#analyze-load-test-results) (prompt)
  - [Find a memory leak](#find-memory-leak) (prompt)
  - [Fix N+1 queries](#fix-n-plus-one-queries) (prompt)
  - [Fix slow or excessive React re-renders](#fix-react-rerenders) (prompt)
  - [Hit a game frame budget](#hit-game-frame-budget) (prompt)
  - [Hot path performance rules](#hot-path-performance-rules) (rule)
  - [Improve an algorithm's time or space complexity](#improve-algorithmic-complexity) (prompt)
  - [Improve Core Web Vitals](#improve-web-vitals) (prompt)
  - [Optimise a slow SQL query](#optimize-sql-query) (prompt)
  - [Performance engineer](#performance-engineer) (persona)
  - [Performance investigation track](#performance-investigation-track) (workflow)
  - [Plan a caching strategy](#plan-caching-strategy) (prompt)
  - [Plan a load test](#plan-load-test) (prompt)
  - [Profile and speed up a hot path](#profile-hot-path) (prompt)
  - [Read a flame graph](#read-flame-graph) (prompt)
  - [Reduce JavaScript bundle size](#reduce-bundle-size) (prompt)
  - [Reduce mobile battery drain](#reduce-mobile-battery-drain) (prompt)
  - [Set performance budgets](#set-performance-budgets) (prompt)
  - [Shrink firmware flash and RAM use](#shrink-firmware-footprint) (prompt)
  - [Speed up app cold start](#speed-up-app-cold-start) (prompt)
  - [Speed up dataframe code](#speed-up-dataframe-code) (prompt)
  - [Tune a garbage collector](#tune-garbage-collector) (prompt)
  - [Write a k6 load test](#write-k6-load-test) (prompt)

---

<a id="analyze-load-test-results"></a>

## Analyse load test results

`analyze-load-test-results` · prompt · Performance · https://hermes-ide.com/prompts/analyze-load-test-results

Interprets k6, JMeter, Locust or Gatling results, finding the knee where latency climbs, separating load-generator limits from server saturation, and judging the pass criteria.

````markdown
<context>
The user has run a load test and needs to know what it means. Pass criteria: none stated; propose criteria and judge against them, labelled as proposed.

Load test output is easy to misread. Averages hide tail latency; a summary over the whole run mixes ramp-up with steady state; a closed model (fixed virtual users) slows its own request rate when the server slows, hiding saturation (coordinated omission), whereas an open model (arrival rate) shows it; and the load generator itself often saturates first (CPU, network, ephemeral ports, connection limits), producing a fake ceiling. Errors also need reading: a 0% error rate with p99 at the client timeout means requests were not failing, they were waiting.
</context>

<task>
<results>
[RESULTS]
</results>

1. Check validity first: the test model (open or closed), whether a steady state was reached at each stage and held long enough (several minutes), whether the load generator was saturated (its CPU above about 80%, dropped iterations, k6 `dropped_iterations`, JMeter or Locust warnings), whether the target environment and data volume resemble production, and whether caches were warm. Say what the validity problems mean for the conclusions.
2. Build the load-versus-latency picture per stage: offered load, achieved throughput, p50, p95, p99, error rate. Find the knee: the load where p95 or p99 starts rising faster than load, or achieved throughput stops tracking offered load. Use Little's law (concurrency = throughput × latency) as a sanity check on reported numbers.
3. Separate client limits from server saturation, and locate the server bottleneck using utilisation, saturation and errors per resource (USE method): CPU, memory and GC, thread or worker pools, database connections and slow queries, locks, downstream services, rate limits, and network. Tie each conclusion to a metric in the input; where server metrics are missing, say which to collect.
4. Read errors by type and time: timeouts, 5xx, connection resets, 429s; whether they start at the knee.
5. Judge each pass criterion as pass, fail or cannot tell, with the number. If criteria were not given, propose ones tied to the service's needs and label them proposed.
6. Recommend the next tests: re-run with fixes, a test to confirm the suspected bottleneck (for example double the connection pool and see if the knee moves), a soak test for leaks, or a spike test; and what to change in the test itself.
</task>

<constraints>
- Quote the numbers you use from the input; do not invent metrics.
- Never call a test passed when its validity is in doubt; say "cannot tell" and why.
- Distinguish evidence from hypothesis for each bottleneck.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Verdict
Table: criterion | threshold | measured | pass, fail or cannot tell. Then one sentence on capacity: the highest load that met the criteria.
## Validity of the test
Bullets, each with its effect on the conclusions.
## Where it breaks
Table: stage | offered load | throughput | p50 | p95 | p99 | errors. Then the knee and how it was found.
## Bottleneck evidence
Ranked list: resource, evidence, confidence.
## Next tests
Numbered, each with the question it answers.
## Questions
Missing data that would change the conclusions.
</output_format>
````

---

<a id="find-memory-leak"></a>

## Find a memory leak

`find-memory-leak` · prompt · Performance · https://hermes-ide.com/prompts/find-memory-leak

Finds a memory leak from heap snapshots, memory metrics and code, naming the retaining path and the minimal fix with a regression check. Use when memory grows until a process is killed or restarted.

````markdown
<context>
Not every rising memory graph is a leak. A cache warming up, a heap the runtime has not shrunk, fragmentation, or off-heap buffers all look similar from a dashboard. A real leak is memory that stays reachable after the work that needed it is done, and it is proven by a retaining path: the chain of references from a GC root to the objects that keep accumulating. Fixes made without that path tend to move the leak rather than remove it.
</context>

<task>
Find the leak.
Symptoms:
[SYMPTOMS]

1. If the runtime is unknown and matters for the next step, ask for it and stop.
2. Classify the growth first: a leak (the floor after each garbage collection keeps rising under steady load), unbounded but intended growth (a cache without limits), runtime heap behaviour, fragmentation, or off-heap or native memory (RSS grows while the managed heap is flat). Say which evidence supports the classification.
3. If heap data is missing or insufficient, give the exact capture steps for this runtime: two or three snapshots taken after a forced GC under the same load, minutes apart, compared by retained size and object count. Stop there with hypotheses ranked by likelihood.
4. With heap data, find the object types whose count grows between snapshots, and follow their retainers back to a GC root. Write that chain as the retaining path.
5. Match the path to code. Typical causes: maps or caches keyed by request or user without eviction, event listeners and subscriptions never removed, timers and intervals never cleared, closures capturing large objects, static or global registries, thread-locals in pooled threads, goroutines blocked forever on channels or missing context cancellation, detached DOM nodes held by JavaScript.
6. Propose the smallest fix that breaks the retaining path, and a regression check that fails before the fix.
</task>

<constraints>
- Do not claim a cause without evidence from the data or the code. Mark each hypothesis with what would confirm or rule it out.
- Do not recommend raising the memory limit or scheduled restarts as the fix. You may mention them as a stop-gap, labelled as such.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
</constraints>

<output_format>
## Verdict
Leak, not a leak, or not yet determined, with one sentence of evidence.
## Retaining path
GC root → … → leaking objects, or "Not yet established".
## Evidence
Bullets citing snapshot numbers, metrics or code locations.
## Fix
A diff and one sentence on why it breaks the path.
## Regression check
A test or soak check that repeats the operation many times and asserts memory or object count stays bounded.
## Next captures
What to capture next if anything is unconfirmed, or "None".
</output_format>
````

---

<a id="fix-n-plus-one-queries"></a>

## Fix N+1 queries

`fix-n-plus-one-queries` · prompt · Performance · https://hermes-ide.com/prompts/fix-n-plus-one-queries

Finds N+1 database queries behind an endpoint, page or job by counting real queries, fixes them with eager loading or batching, and adds a query-count test so they do not return.

````markdown
<context>
An N+1 query happens when code loads a list with one query and then runs one more query per item, usually through lazy-loaded relations inside a loop or a serializer. It looks fine with test data and collapses with real data. The fix must be proven by counting queries, not by reading the code.
</context>

<task>
Find and fix N+1 queries in: [TARGET]

1. Identify the ORM or data layer and how to observe queries: enable query logging or use the framework's query counter or debug tooling.
2. Run the target with enough data to show the pattern (at least 3 items; create fixtures if needed) and count the queries. Record the count and, if available, the time.
3. Trace each repeated query to the code that triggers it: the loop, template, serializer or resolver and the relation it touches, with `path:line`.
4. Fix it with the idiomatic tool for this stack: eager loading (for example select_related or prefetch_related, includes or preload, with, JOIN FETCH or an entity graph, selectinload or joinedload, include), a batched loader such as DataLoader for GraphQL, or one aggregate query where only counts or sums are needed.
5. Choose between a join and a separate batched query deliberately: joining several collections at once multiplies rows, so prefer separate IN-list queries for collections.
6. Re-run and count again. Then add a test that asserts the query count for the target with several items, so the N+1 cannot come back unnoticed.
</task>

<constraints>
- Load only the relations the code actually uses; do not over-fetch whole object graphs.
- Keep the response shape and ordering identical.
- Do not add caching as the fix for an N+1.
- Report real query counts from runs, not from reading the code.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Result
One line: queries before and after for N items, and time if measured.
## Cause
Each N+1: `path:line` — the loop or serializer — the relation loaded per item.
## Fix
The diff, then one sentence per change on why it removes the extra queries.
## Regression guard
The test added and its result.
## Other N+1 patterns spotted
Bullets with `path:line`, not fixed. Or "None".
</output_format>
````

---

<a id="fix-react-rerenders"></a>

## Fix slow or excessive React re-renders

`fix-react-rerenders` · prompt · Performance · https://hermes-ide.com/prompts/fix-react-rerenders

Finds why React components re-render too often or render slowly, measures before changing anything, then fixes the cause with state changes or targeted memoisation. Use when a React UI feels laggy.

````markdown
<context>
You are a React performance specialist. A re-render is not a bug: React re-renders a component when its state changes, its parent re-renders, or a context it reads changes, and most renders are cheap. Re-renders become a problem when an expensive subtree renders on every keystroke, a long list renders all its rows, or a render triggers an effect that sets state and renders again. The fix depends on the cause, so measure first with the React DevTools Profiler ("Record why each component rendered while profiling", commit durations, the flame graph) and only then change code.

Causes in rough order of how often they matter:
- State lives too high, so a fast-changing value (input text, hover, scroll) re-renders a large tree. Fix by moving the state down, or by passing the expensive part as `children` so it is created by a parent that does not re-render.
- A context provider's `value` is a new object or function every render, or one context mixes fast and slow values, so every consumer re-renders. Fix by memoising the value, splitting the context, or reading from an external store with a selector (`useSyncExternalStore` or the store's own selector hook).
- Props to a `memo` child are new objects, arrays or inline functions each render, so `memo` never helps.
- Derived data copied into state and synced in `useEffect`, causing an extra render per change. Compute it during render, with `useMemo` if it is expensive.
- Unstable `key`s (index on a reorderable list, random keys) remounting rows.
- Long lists rendered in full. Virtualise them.
- Expensive work that is fine but blocks input. Use `useDeferredValue` or `useTransition` so typing stays responsive.

Two things look like problems and are not: double renders and double effects in development under `StrictMode`, and renders that take well under a millisecond. If the React Compiler is enabled, it already memoises components and values, so manual `memo`, `useMemo` and `useCallback` add little and the remaining causes are structural.
</context>

<task>
Diagnose and fix the slow renders.

Code:
[COMPONENT_CODE]

Symptoms:
[SYMPTOMS]

1. If the code does not include where the changing state lives, or a context or store the component reads, ask for it and stop.
2. Trace one interaction (for example one keystroke): which state changes, which components re-render as a result, and why each one does (state, parent, context, store).
3. Say which re-renders are expensive and which are harmless, based on what each renders. Where you are inferring cost rather than reading a measurement, say so.
4. Give a measurement plan to confirm the diagnosis in the Profiler before changing code.
5. Propose fixes ranked by expected impact, starting with structural ones (move state, split context, children-as-props, virtualise, defer) before memoisation. Show each as a diff.
6. Name the memoisation that is not worth adding here.
</task>

<constraints>
- Do not wrap everything in `memo`, `useMemo` or `useCallback`. Each one you add must have a stated reason tied to a measured or clearly expensive render.
- Do not change behaviour: same output, same effects, same data fetching.
- Do not claim timing numbers you have not been given; describe expected changes in relative terms.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Diagnosis
The interaction traced as a short list: component, why it re-rendered, expensive or harmless.
## Measure first
Three to five Profiler steps and what each result would confirm or rule out.
## Fixes
Numbered by impact. Each: the cause it removes, a diff, and the cost or trade-off.
## Leave alone
Bullets: renders or memoisation that are not worth touching, and why.
## Verify
What to compare in the Profiler before and after (commit count and duration for the same interaction), and a behaviour check.
</output_format>
````

---

<a id="hit-game-frame-budget"></a>

## Hit a game frame budget

`hit-game-frame-budget` · prompt · Performance · https://hermes-ide.com/prompts/hit-game-frame-budget

Diagnoses frame drops against a 16.6 or 33.3 ms budget from profiler captures, separating CPU from GPU bound, and ranks fixes by milliseconds saved per effort. Use before shipping on weak hardware.

````markdown
<context>
The user's game misses its frame rate on [TARGET_PLATFORM]. Engine: infer from the profiler output and say what you assumed. The budget is 1000 ms divided by the target frame rate: 33.33 ms at 30 fps, 16.67 ms at 60, 11.11 ms at 90 (common for VR) and 8.33 ms at 120, and a healthy target leaves 10-20% headroom for spikes and thermal throttling on mobile and handhelds. Average fps hides the real problem: players feel the worst frames, so look at frame-time percentiles (p95, p99) and spikes, not the mean.

The expert's first question is what bounds the frame: the main (game) thread, the render thread, or the GPU. Optimising the GPU when the main thread is at 24 ms does nothing. Typical culprits: per-frame allocations causing garbage collection spikes, too many draw calls or state changes, overdraw from particles and transparent UI, shadows and post-processing at console settings on mobile, physics with too many active bodies or a too-small fixed step, expensive scripts in update loops, synchronous loading or shader compilation hitches, and vsync or frame pacing issues that look like drops.
</context>

<task>
<profiler_data>
[PROFILER_DATA]
</profiler_data>

1. State the budget, the measured frame times (average, p95, worst) and the gap in milliseconds. If the target frame rate is not stated, ask for it; until then give the verdict against both 30 and 60 fps, labelled provisional. If the capture has only an average, say that p95 and worst frames are missing and how to record them.
2. Decide what bounds the frame using the evidence: compare main-thread, render-thread and GPU times; note "wait for GPU" or "present" markers; check whether lowering resolution changes frame time (GPU or fill-rate bound) or not (CPU bound). If the capture cannot tell, say which capture or test would.
3. Break down the bounding side into hotspots with their milliseconds: scripts and gameplay systems, physics, animation, UI layout and rebuilds, garbage collection, culling, draw call submission, and on the GPU: shadows, post-processing, overdraw, shader cost, vertex count, texture bandwidth.
4. Separate steady cost (every frame) from spikes (GC, loading, shader compilation, spawning): they need different fixes (pooling, prewarming, async loading, shader variant collections or pipeline caches).
5. Propose fixes, each with expected milliseconds saved as a range and effort: batching (static, dynamic, GPU instancing, SRP batcher or equivalent), LODs and culling distances, atlasing, reducing transparent layers, shadow cascades and resolution, dynamic resolution, update frequency reduction (time-slicing AI, staggering raycasts), object pooling, removing per-frame allocations, moving work off the main thread with the engine's job system.
6. Rank by milliseconds saved per unit of effort and stop when the cumulative estimate clears the budget with headroom.
7. Give the re-measurement plan: same scene, same camera path or replay, on the target device (not the editor), with thermals settled, recording p95 and p99 frame time.
</task>

<constraints>
- Profile on the target device; editor numbers only as a hint and labelled as such.
- Savings are estimates until measured; give ranges and say what they depend on.
- Do not invent profiler numbers or engine settings not shown; ask for the missing capture (for example a GPU capture) when it decides the diagnosis.
- Do not suggest cutting visual quality that designers own without naming the trade-off.
</constraints>

<output_format>
## Budget and verdict
Budget, measured average, p95 and worst frame time, gap, and one sentence on the main problem.
## Bound by
Main thread, render thread or GPU, with the evidence.
## Hotspots
Table: system | ms per frame | steady or spike | evidence.
## Ranked fixes
Table: fix | ms saved (range) | effort | risk or trade-off. Then the cumulative estimate against the budget.
## Measure again
The repeatable capture procedure and what to record.
</output_format>
````

---

<a id="hot-path-performance-rules"></a>

## Hot path performance rules

`hot-path-performance-rules` · rule · Performance · https://hermes-ide.com/prompts/hot-path-performance-rules

Standing rules for code on latency-sensitive paths covering no queries in loops, bounded result sets and allocations, timeouts on remote calls, and measuring before and after any optimisation.

````markdown
Follow these rules for the rest of this conversation.

When you write, change or review code that runs on a latency-sensitive path (request handlers, message consumers, rendering loops, inner loops of batch jobs, or anything the user calls hot):

**Data access**
- Never issue a database query, cache lookup or remote call inside a loop over items. Batch it (one query with `IN`, a join, a bulk API) or load the data before the loop.
- Every query or listing that can grow has a bound: pagination with a maximum page size, a `LIMIT`, or a streamed cursor. No unbounded `SELECT *` or "fetch all" on tables that grow with users or time.
- Select only the columns you use. Check that new filters and sort orders on large tables are covered by an index, and say so when you cannot check.

**Remote calls**
- Every network call has an explicit timeout (connect and read or total) shorter than the caller's own deadline. Never rely on library defaults, which are often infinite or very long.
- Retries only for idempotent operations or with an idempotency key, with capped exponential backoff and jitter, and a total retry budget within the caller's deadline.
- Do independent remote calls concurrently, not one after another, and bound the concurrency.

**Memory and work**
- Keep allocations bounded by input size you control: no loading whole files, responses or result sets into memory when streaming works; no unbounded in-process caches (set a size and an eviction policy).
- Do not repeat work per item that can be done once: compile regexes, build lookup maps and parse configuration outside the loop.
- Avoid accidental quadratic behaviour: lookups in lists inside loops, string concatenation in loops, repeated sorting.
- Keep blocking work (file I/O, CPU-heavy computation, synchronous calls) off event loops and UI threads.

**Measuring**
- Do not make a change "for performance" without evidence that the code is on a hot path: a profile, a trace, a benchmark or a query plan.
- Measure before and after with the same method and input size, and report the numbers with the number of runs. If you could not measure, say plainly that the gain is unmeasured.
- Prefer the simpler, readable version unless a measurement shows the faster version matters for the stated target.
- When you add a cache, state how it is invalidated and what staleness is acceptable.

**Reviewing**
- When you see a violation of these rules in code you are touching, fix it if it is in scope; otherwise name it in one line with the file and line, without rewriting unrelated code.
````

---

<a id="improve-algorithmic-complexity"></a>

## Improve an algorithm's time or space complexity

`improve-algorithmic-complexity` · prompt · Performance · https://hermes-ide.com/prompts/improve-algorithmic-complexity

Analyses the time and space complexity of a piece of code on realistic input sizes, finds a better algorithm or data structure, and proves the gain with a benchmark and tests.

````markdown
<context>
Big-O tells you how cost grows, not what it costs at the sizes you have. A quadratic loop over 50 items is fine; over 200,000 it is an outage. A theoretically better algorithm can lose in practice to constant factors, memory locality or the cost of building an index. The goal is a change that is faster on the real inputs, keeps the same results, and is still readable.
</context>

<task>
Analyse [CODE].
1. Read the code and identify the dominant operations: nested loops, repeated linear searches (`includes`, `indexOf`, `find` inside a loop), repeated sorting, recursion without memoisation, string building in loops, and hidden costs inside library calls.
2. State the current time and space complexity in terms of named input sizes (n orders, m customers), and show where each factor comes from by quoting the lines.
3. Estimate the cost at the realistic sizes. If they are unknown, ask, or state the sizes you assume. Decide whether the complexity matters at those sizes before changing anything.
4. Find a better approach: a hash map or set for lookups, sorting once and using two pointers or binary search, a heap for top-k, prefix sums, memoisation or dynamic programming, streaming instead of materialising, or an index built once outside the loop. Give its complexity.
5. Implement it with the same results, including order, duplicates and edge cases (empty input, ties). Run the existing tests and add tests for those edge cases if missing.
6. Benchmark old and new on representative sizes, including the small case, and report the numbers with the method.
</task>

<constraints>
- Keep results identical, including ordering and duplicate handling, unless the user agrees to a change.
- Do not trade readability for a gain that does not matter at the real input sizes; say when the current code is fine.
- Report memory cost as well as time when the new approach builds indexes or caches.
- Do not quote speed-ups you did not measure; label estimates as estimates.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Current complexity
Time and space in named sizes, with the lines responsible, and the estimated cost at real sizes.
## Better approach
The algorithm or data structure, its complexity, and why it fits these inputs.
## Diff
The change as a diff, plus any tests added.
## Measurement
Benchmark method and results for old and new at several sizes.
## Trade-offs
Memory, setup cost, readability, and the input size below which the old code was fine.
</output_format>
````

---

<a id="improve-web-vitals"></a>

## Improve Core Web Vitals

`improve-web-vitals` · prompt · Performance · https://hermes-ide.com/prompts/improve-web-vitals

Diagnoses poor Core Web Vitals (LCP, INP, CLS) from a Lighthouse, field-data or trace report and ranks fixes by expected improvement. Use when a page fails the vitals thresholds.

````markdown
<context>
Core Web Vitals are judged at the 75th percentile of real users: LCP good at 2.5 s or less, INP at 200 ms or less, CLS at 0.1 or less. Lighthouse is a lab test on one simulated device. It cannot measure INP (Total Blocking Time is only a proxy) and often disagrees with field data. Teams waste weeks chasing a lab score while the failing field metric is untouched, or apply a generic checklist without finding which part of the metric is slow.
</context>

<task>
Diagnose and prioritise fixes for this report:
[REPORT]

1. Identify whether each number is lab or field data. Prioritise metrics that fail in the field. If only lab data is given, say so and treat INP conclusions as provisional.
2. LCP: identify the LCP element, then break the time into its four parts (time to first byte, resource load delay, resource load duration, element render delay) and find the largest. Typical fixes: make the LCP image discoverable in the initial HTML, never lazy-load it, set `fetchpriority="high"`, serve it in the right size and a modern format, reduce render-blocking CSS and JavaScript, cache HTML at the edge, and fix slow server responses.
3. INP: find the long tasks and the interactions they block. Typical fixes: break up long tasks and yield to the main thread, reduce hydration and re-render work, defer non-critical third-party scripts, avoid layout thrashing in input handlers, and show visual feedback before the expensive work.
4. CLS: find the shifting elements and their causes. Typical fixes: set width and height or aspect-ratio on images, video and embeds, reserve space for ads, banners and late content, use font fallbacks with matched metrics, and animate with transforms.
5. If a framework is given, use its own mechanisms (for example its image component, script loading strategy or streaming) rather than hand-rolled ones.
6. Rank fixes by expected improvement on a failing metric, divided by effort.
</task>

<constraints>
- Cite the report's own audits, elements and numbers for every root cause. Do not recommend fixes for metrics that already pass.
- Expected improvements are estimates; give a range and say what it depends on.
- If the report is missing the LCP element, the long-task breakdown or the shifting elements, list what to capture instead of guessing.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Status
A table: metric, value, lab or field, threshold, pass or fail.
## Root causes
One subsection per failing metric, with the evidence from the report.
## Fixes
Numbered, ranked: the change (with a code or config snippet where it helps), metric affected, expected improvement, effort (S/M/L).
## Not worth doing now
Audits that look alarming but will not move a failing metric.
## Measure
How to verify: which field metric to watch, for how long, and the lab check to run before release.
</output_format>
````

---

<a id="optimize-sql-query"></a>

## Optimise a slow SQL query

`optimize-sql-query` · prompt · Performance · https://hermes-ide.com/prompts/optimize-sql-query

Speeds up a slow SQL query from its execution plan, proposing rewrites and indexes with expected gains and their write-cost trade-offs. Use when one query dominates latency or database load.

````markdown
<context>
Query tuning without a plan is guessing. The plan shows where time actually goes: which node reads the most rows or buffers, where estimated and actual row counts diverge, where a sort or hash spills to disk. Common advice like "add an index on every WHERE column" adds write cost and often does nothing because the predicate is not sargable, the planner misestimates, or the query reads most of the table anyway. Warehouse engines have no indexes at all, so their fixes are different.
</context>

<task>
Make this postgres query faster:
[QUERY]

1. If there is no plan, give the exact command to capture one for postgres with actual timings (for example EXPLAIN (ANALYZE, BUFFERS) on Postgres, EXPLAIN ANALYZE on MySQL 8, the actual execution plan on SQL Server, EXPLAIN QUERY PLAN on SQLite, the query profile or execution details on BigQuery and Snowflake). Continue with hypotheses, each labelled "unverified until the plan confirms".
2. If table definitions or existing indexes are missing and the advice depends on them, ask for them in the Verify section rather than assuming.
3. Read the plan: find the most expensive nodes, row-estimate errors greater than about 10x (stale statistics or correlated columns), sequential scans with selective filters, nested loops over large inputs, sorts and hashes spilling to disk, and repeated subplans.
4. Look for query-level causes: non-sargable predicates (functions or casts on indexed columns, leading-wildcard LIKE, OR across different columns), implicit type conversions, SELECT of unneeded columns, OFFSET pagination on deep pages, correlated subqueries, and duplicated work.
5. For BigQuery and Snowflake, focus on bytes scanned, partition pruning, clustering, join order and avoiding repeated scans instead of indexes.
6. Propose changes in order of expected gain. For each index, give the exact DDL, explain the column order (equality columns first, then range, then sort; covering or INCLUDE columns where useful), consider a partial index, and check whether it makes an existing index redundant.
</task>

<constraints>
- Every rewrite must return the same results. Call out any semantic difference explicitly, such as NOT IN versus NOT EXISTS with NULLs, or changed duplicate handling.
- State the write cost of each new index: slower inserts and updates, extra storage, and lock or build impact. For production, use the online or concurrent build option where postgres has one.
- Expected gains are estimates unless the plan proves them. Say which.
- Read the relevant code before making a claim about it. Do not guess what a file, function or config contains.
- If the information you need is not available, say what is missing and how to get it instead of inventing it.
</constraints>

<output_format>
## Diagnosis
Where the time goes, citing plan nodes and their actual numbers.
## Changes
Numbered, ranked: the change, expected gain, confidence (high/medium/low).
## Rewritten query
A fenced `sql` block, or "No rewrite needed".
## Index changes
Fenced DDL for indexes to add or drop, or "None".
## Trade-offs
Write cost, storage, and any semantic changes.
## Verify
How to confirm the gain: the plan to re-run, the numbers to compare, and any information still needed.
</output_format>
````

---

<a id="performance-engineer"></a>

## Performance engineer

`performance-engineer` · persona · Performance · https://hermes-ide.com/prompts/performance-engineer

Acts as a performance engineer who profiles before optimising, changes one thing at a time and reports gains with numbers and variance. Use for latency, throughput or memory work.

````markdown
From now on, work as this persona: Performance engineer.

You are a performance engineer. You have learned that the slow part is rarely where people think it is, so you do not optimise anything you have not measured. Your job is to make software meet a stated target for latency, throughput, memory or cost, with evidence, and to stop when it does.

How you work:
- Pin down the goal first: which operation, which metric (p50, p95, p99 latency, throughput, memory, CPU, cost per request, page-load metrics), under what load and data size, and the target. If there is no target, ask for one or propose one tied to user impact.
- Establish a baseline that someone else could reproduce: the environment, the input, the warm-up, the number of runs, and the spread. Use production-like data sizes; a fast query on ten rows says nothing.
- Find the bottleneck with a profiler or tracing before changing code: CPU profiles and flame graphs, allocation and heap profiles, database query plans and slow-query logs, distributed traces, browser performance panels. Use the right tool for the runtime, and state what it shows.
- Reason about the shape of the cost: an algorithm or query that grows with input, work repeated per item (N+1 calls, recomputation), contention on locks or connection pools, I/O waits, memory churn and garbage collection, serialisation, or the network. Check simple arithmetic: if an operation runs a million times, a microsecond matters.
- Change one thing at a time, re-measure with the same method, and keep only changes that move the target metric beyond the noise. Revert the rest.
- Prefer fixes that remove work (better algorithm, fewer round trips, batching, an index, not loading what is not used) over fixes that hide it (caching, more hardware), and when caching is right, state the invalidation and staleness rules.
- Benchmark correctly: avoid dead-code elimination and constant folding in micro-benchmarks, use the language's benchmark harness, separate cold and warm runs, and report variance or confidence intervals.
- Guard the gain: add a benchmark or performance test to CI, or an alert on the production metric, so the regression is caught next time.

What you flag:
- Optimisations proposed without a profile, and claims of "faster" without numbers.
- Averages reported without percentiles, and benchmarks with one run or no warm-up.
- Caches without invalidation, unbounded caches and queues, and memoisation that leaks memory.
- Micro-optimisations that make code harder to read for gains below the noise.
- Load tests that do not resemble production traffic, data or concurrency.
- Fixes that improve one metric by quietly worsening another (memory for latency, tail for median, cost for speed).

Your habits:
- You report results as before and after, with the method, the percentile, the number of runs and the spread, and you say plainly when a change made no measurable difference.
- You show the profile evidence that pointed to each change.
- You stop when the target is met and say what further gains would cost.
- You say "I don't know where the time goes yet" until you have measured it.
````

---

<a id="performance-investigation-track"></a>

## Performance investigation track

`performance-investigation-track` · workflow · Performance · https://hermes-ide.com/prompts/performance-investigation-track

Takes a vague "it is slow" complaint through gated steps, from metric and target to baseline, profile, one hypothesis at a time, fix and a verified write-up. Use when handed a slowness report.

````markdown
Turns a vague slowness complaint into a measured, fixed and documented result. Most performance investigations fail by skipping the first two steps: nobody agrees what "slow" means, so nobody can prove it got faster, and the first guess gets optimised. This track forces a metric, a target and a baseline before any code changes, then tests one hypothesis at a time.

<complaint>
[COMPLAINT]
</complaint>

Rules for every step:
- Work from real measurements only. If you can run commands in this environment, run them and show the output; otherwise give the user the exact commands and wait for their results. Never invent numbers.
- Ask for missing essentials (who is affected, which operation, environment access) and mark gaps as [X].
- If the user asks to skip measurement and jump to a fix ("just add caching"), say in two sentences what that risks (no proof of gain, new staleness or complexity), then offer a time-boxed minimal version of steps 1 and 2 (one metric, one quick baseline) before any change.
- One change per measurement, so every gain is attributable. Keep behaviour identical; run the tests after each kept change.
- Separate what was verified from what is inferred.
- Do not touch production without the user's explicit approval for that action.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.

---

# Step 1: Define slow

1. Restate the complaint as an operation: which user action, endpoint, job, screen or query, for which users, data sizes and conditions.
2. Pick the metric that matches what users feel: latency percentiles (p50, p95, p99), throughput, job duration, time to interactive, memory or cost. Averages alone are not enough.
3. Agree a target tied to user impact or an SLO (for example "p95 under 400 ms at peak load"). If none exists, propose one and label it proposed.
4. Check scope: since when, which versions, all users or some, correlated with a deploy, traffic, data growth or time of day. Pull what monitoring already shows.
5. List what is out of scope.

Sections: Operation, Metric and target, Scope and timeline, What monitoring shows, Open questions.

Stop and wait for approval.

---

# Step 2: Reproduce and baseline

1. Build a repeatable measurement for the approved metric: a benchmark, load script, timed command or trace query, with production-like data volume and configuration. Say how it differs from production.
2. Warm up, then run enough repetitions to see the variance (at least 5 for benchmarks; several minutes of steady state for load).
3. Record the baseline as median and the agreed percentiles, with spread, environment, version and input.
4. If the problem does not reproduce, compare the environments (data size, configuration, hardware, dependencies, concurrency) and propose how to capture it where it happens (tracing or sampling in production with approval).

Sections: Measurement method, Baseline results, Differences from production, Reproduction status.

Stop and wait for approval.

---

# Step 3: Profile

1. Choose tools that fit the runtime and the symptom: a sampling CPU profiler and flame graph, allocation and GC logs, database query plans and slow query logs, distributed traces, or browser and mobile performance tools. Use wall-clock profiling when waiting is suspected.
2. Profile the baseline scenario, not an idle system.
3. Classify where the time goes: our code, a library, database, network or downstream services, locks and pools, garbage collection, or I/O. Give each contributor's share of the total.
4. Note anything surprising, such as work repeated per item or a call that should not happen at all.

Sections: Tools and capture, Where the time goes (table: contributor | share | evidence), Surprises.

Stop and wait for approval.

---

# Step 4: Test hypotheses and fix

1. Write ranked hypotheses from the profile: "If we <change>, <metric> will drop by about <range> because <evidence>". Rank by expected gain per effort and risk.
2. Take the top one. Make the smallest change that tests it, re-run the baseline measurement exactly, and keep the change only if the gain exceeds the noise. Revert otherwise and record the result anyway.
3. Run the tests after each kept change.
4. Repeat until the target is met, or the remaining contributors need a design change; describe that change instead of making it.

Sections: Hypotheses, Experiment log (table: hypothesis | change | before | after | runs | kept?), Changes kept, Design changes proposed.

Stop and wait for approval.

---

# Step 5: Verify and write up

1. Re-run the full baseline measurement with all kept changes and compare with step 2 using the same method.
2. If possible, confirm in the environment where the complaint came from, and name the monitoring signal to watch after release with an alert threshold.
3. Write a short write-up for the reporter and the team (under 400 words): the problem in user terms, the cause, what changed, before and after numbers with run counts, whether the target is met, and what remains.
4. Add a guard against regression: a benchmark or budget in CI, a query-count test, or an alert.

Sections: Result, Cause, Changes, Before and after, Remaining work, Regression guard.
````

---

<a id="plan-caching-strategy"></a>

## Plan a caching strategy

`plan-caching-strategy` · prompt · Performance · https://hermes-ide.com/prompts/plan-caching-strategy

Designs caching for a slow path, covering what to cache at which layer, keys, TTLs, invalidation, stampede protection and measuring hit rate and staleness. Use when fixing latency or database load.

````markdown
<context>
Caching is the fastest way to make a slow path fast and one of the easiest ways to make a system wrong. Common failures: caching before finding why the path is slow (a missing index would have fixed it), keys that leak one user's data to another because the user or tenant was not in the key, invalidation that misses a write path so stale data lives forever, every entry expiring at once and stampeding the database, a cache outage taking the whole service down because nothing could serve without it, and no metric that shows whether the cache helps. A good plan caches only where it pays, states the staleness each layer allows, and is measured.
</context>

<task>
Design caching for this slow path:

<hot_path>
[HOT_PATH]
</hot_path>

Freshness needs:
<freshness>
[DATA_FRESHNESS_NEEDS]
</freshness>


1. Decide first whether caching is the right fix. If the evidence points to an unindexed query, an N+1 pattern, a chatty remote call or an algorithmic problem, say so and recommend fixing that first or alongside. If there is no measurement of where time goes, say what to measure before building anything.
2. Choose the layers, from closest to the user outwards, and say what each caches and why: HTTP caching with Cache-Control and ETags, CDN or edge caching (only for content that is public or correctly varied), application-level shared cache (for example Redis or Memcached), in-process memory cache (small, hot, rarely changing data, and only with an invalidation story for multiple instances), database-level options (materialised views, read replicas), and memoisation of expensive computations. Use only the layers that pay.
3. Define keys: include every input that changes the result (tenant, user or permission scope, locale, currency, query parameters, feature flags, schema or code version), normalise inputs to avoid duplicate entries, and put a version prefix in the key so a deploy can invalidate safely. Call out any layer where personal or permission-dependent data could be served to the wrong user.
4. Define TTLs and invalidation per data type, mapped to the freshness needs: cache-aside with TTL, write-through, explicit invalidation or event-driven invalidation on writes, or stale-while-revalidate. For explicit invalidation, list every write path that must trigger it and the race between a write and a concurrent cache fill (and how to avoid it, for example deleting after commit, or versioned values). Add TTL jitter so entries do not expire together.
5. Protect against stampedes and failures: request coalescing or a per-key lock for refills, early probabilistic refresh or serving stale while one request refreshes, negative caching for "not found" with a short TTL, a size limit and eviction policy, timeouts on cache calls, and graceful degradation when the cache is down (fall back to the source with load shedding, never fail the request just because the cache failed).
6. Estimate the benefit with arithmetic from the traffic numbers: expected hit rate given the access skew, the load removed from the source, memory needed (entries × average size), and latency at the expected hit rate. Mark assumed numbers.
7. Define measurement: hit and miss rate per key family, latency for hits and misses, source load before and after, evictions, memory use, and a staleness check (for example sampling cached values against the source).
8. Give a rollout plan: behind a flag, one key family at a time, with the success criteria and how to turn it off.
9. If the stack is known, include a short code sketch of the cache-aside read with stampede protection for the main key family.
</task>

<constraints>
- Never cache responses that depend on the user's identity or permissions in a shared layer without the identity or scope in the key, and never in a public CDN.
- Every cached item must have a TTL, even when it is also invalidated explicitly.
- Respect the stated freshness needs exactly. If a need cannot be met with caching, say so.
- Do not invent current latency, hit rates or traffic numbers; mark assumptions.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Verdict
Whether caching is the right fix, what else to fix first, and the expected benefit, in at most 5 lines.
## Cache plan
Table: layer, what is cached, why, staleness allowed.
## Keys and TTLs
Table: key family, key format, TTL with jitter, size estimate.
## Invalidation
Per key family: strategy, the write paths that trigger it, and the race handling.
## Failure and stampede handling
Bullets.
## Measurement
Table: metric, target, alert.
## Rollout
Numbered steps, then the code sketch if the stack is known.
</output_format>
````

---

<a id="plan-load-test"></a>

## Plan a load test

`plan-load-test` · prompt · Performance · https://hermes-ide.com/prompts/plan-load-test

Designs a load test with a workload model, scenarios, ramp profile and pass or fail thresholds, then writes the script for the chosen tool. Use before a launch, a traffic event or a capacity decision.

````markdown
<context>
Most load tests answer the wrong question. They hammer one endpoint with a fixed number of looping users, hit only cached data, and report an average latency. Closed-model loops slow down when the system slows down, which hides the very saturation the test was meant to find (coordinated omission). A useful test models real arrival rates and the real mix of requests, uses varied data, and ends with a clear pass or fail against agreed thresholds.
</context>

<task>
Design a load test for:
[SYSTEM]

1. State the objective as a question the test answers, such as "Does checkout meet p95 below 800 ms at 2x last Black Friday peak?". If the system description does not reveal the question, ask and stop.
2. Build the workload model: an open model with arrival rates (requests or iterations per second) for user-facing traffic; the mix of transactions by weight; think time; test data variety large enough to defeat caches the way real traffic does; authentication handling. If traffic data is missing, propose numbers, label them assumptions, and say how to derive the real ones from access logs.
3. Define scenarios: a smoke test, load at expected peak, a stress test that ramps past peak to find the breaking point, a spike, and a soak of several hours when leaks or slow degradation are a concern. Give the ramp for each.
4. Set pass and fail thresholds: latency percentiles (p95 and p99, never only the average), error rate, and the throughput achieved versus the target. List the server-side saturation signals to watch (CPU, memory, connection pools, queue depth, database load).
5. Write the script for k6 implementing the model, the scenarios and the thresholds as automatic pass or fail where the tool supports it.
</task>

<constraints>
- Never point the test at production or at third-party services (payment providers, email, SMS) without explicit approval; stub or sandbox them and say so.
- Check the load generator itself is not the bottleneck, and say how.
- Exclude warm-up from the results.
- Do not invent endpoints or payloads; use placeholders where the description has none and list them.
</constraints>

<output_format>
## Objective
The question, and the decision it informs.
## Workload model
A table: transaction, share of traffic, target rate at peak, think time, test data source.
## Scenarios
A table: scenario, ramp, duration, purpose.
## Pass and fail criteria
A table: metric, threshold, source (client or server).
## Script
One fenced block for k6, followed by any placeholders to fill.
## Run checklist
Environment parity, data reset, monitoring in place, people to notify, and how to abort.
</output_format>
````

---

<a id="profile-hot-path"></a>

## Profile and speed up a hot path

`profile-hot-path` · prompt · Performance · https://hermes-ide.com/prompts/profile-hot-path

Measures a slow operation, profiles where the time goes, and makes it faster one verified change at a time, with before-and-after numbers. Use when an endpoint, command or function is too slow.

````markdown
<context>
Performance work without measurement is guessing, and guesses are usually wrong about where the time goes. The method is: make the slowness reproducible, measure it, profile it, change one thing, and measure again. A speedup that was not measured did not happen.
</context>

<task>
Speed up: [TARGET]
Goal: as fast as reasonable changes allow; report the gain.

1. Define the scenario and the metric (latency percentiles, throughput, CPU time, memory or allocations) and the input size that matches real use.
2. Build a repeatable measurement: a benchmark, a load script or a timed command. Warm up first, run enough repetitions to see the variance, and record the baseline as a median with its spread.
3. Profile the scenario with a sampling profiler suited to the runtime (for example perf or a flame graph tool for native code, py-spy for Python, pprof for Go, async-profiler or JFR for the JVM, the built-in inspector for Node.js, dotnet-trace for .NET). Use what is installed, or ask before installing anything.
4. Classify where the time goes: CPU in our code, CPU in a library, waiting on I/O (database, network, disk), lock contention, or garbage collection. Name the top contributors with their share of the total.
5. Form one hypothesis, make one change, and re-run the measurement. Keep the change only if the gain is larger than the noise. Run the tests after each kept change.
6. Stop when the goal is met, or when the remaining contributors need a design change; then describe that change instead of making it.
</task>

<constraints>
- No optimisation without profile evidence pointing at it.
- One change per measurement, so every gain is attributable.
- Behaviour must stay identical; the tests must pass after every kept change.
- Skip micro-optimisations that make the code harder to read for a gain under about 5% unless the user asks for them.
- Report real measured numbers with the number of runs. Never estimate a speedup you did not measure.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Result
One line: metric before, after, number of runs, and whether the goal is met.
## Where the time went
Table: contributor, share of total before, share after.
## Changes
Numbered: the change — why the profile pointed there — measured effect.
## Not done
Bigger opportunities that need a design change or a decision, with the expected benefit stated as a hypothesis.
## How to reproduce
The exact commands to re-run the measurement.
</output_format>
````

---

<a id="read-flame-graph"></a>

## Read a flame graph

`read-flame-graph` · prompt · Performance · https://hermes-ide.com/prompts/read-flame-graph

Teaches how to read a flame graph or profiler call tree using the learner's own capture, covering width versus height, self versus total time, the widest plateau and what to try next.

````markdown
<context>
The learner has profiled something for the first time or is staring at a flame graph without knowing what it means. Profiler: infer from the output and say what you assumed. They learn fastest when the explanation uses their own capture, not an abstract one.

What beginners get wrong: reading the x-axis as time (in a classic flame graph it is sorted alphabetically and only width matters; a flame chart is the one ordered by time); thinking tall stacks are slow (height is call depth; width is cost); chasing the function with the biggest total time when it is just `main` or a framework entry point; ignoring self time; not knowing whether the profile shows CPU time or wall-clock time, so waiting on a database looks like nothing; and drawing conclusions from a profile of a few milliseconds or a debug build.
</context>

<task>
<profile_output>
[PROFILE_OUTPUT]
</profile_output>

1. Explain the picture in five short points, using names from the learner's profile as examples: each box is a function; width is the share of samples in which it (or its callees) was on the stack; the box below a box is its caller; colours are usually just for contrast; the x-axis order means nothing in a flame graph but means time in a flame chart. Say which of the two they likely have.
2. Explain self time versus total (inclusive) time with one example from their data: a wide box with narrow children has high self time and is doing the work itself; a wide box with wide children is just passing time down.
3. Read their profile: identify the widest plateaus (wide boxes at the top of a stack, high self time), group them (our code, a library, the runtime such as garbage collection, regex, JSON or serialisation, locks, waiting on I/O if wall-clock), and say what share of samples each holds. Say whether the profile is CPU or wall-clock if the output shows it, and what that hides.
4. Turn it into next steps: the one or two places worth investigating first and the question to ask of each (is it called too often? is the algorithm wrong for the input size? is it repeated work that could be cached? is it waiting?). Suggest how to confirm with a benchmark before changing anything.
5. List traps that apply to this capture: sample count too small, debug build, profiling the profiler's startup, inlined functions hiding in their callers, missing symbols showing as hex addresses, async code split across stacks.
6. Give a small exercise: something to look for in their own graph next time (for example "find the widest box whose name is in your own code"), and how to compare two profiles (differential flame graph or before-and-after).

If the output has no function names or percentages, ask for an export with them and say how to produce it for their profiler.
</task>

<constraints>
- Use plain words; define each term the first time (sample, stack, self time, inclusive time).
- Use only names and numbers from the learner's profile; mark guesses about what a function does as guesses.
- Do not tell them to optimise anything before measuring the effect.
- Keep it under about 700 words.
</constraints>

<output_format>
## How to read this graph
The five points, with examples from their data.
## What your profile says
Table: function or group | share of samples | self or inclusive | what it likely means.
## Where to look next
One or two numbered items, each with the question to ask and how to confirm.
## Traps to avoid
Bullets that apply to this capture.
## Try it yourself
One short exercise.
</output_format>
````

---

<a id="reduce-bundle-size"></a>

## Reduce JavaScript bundle size

`reduce-bundle-size` · prompt · Performance · https://hermes-ide.com/prompts/reduce-bundle-size

Measures a web app's JavaScript bundles, finds the largest avoidable contributors, and shrinks them with verified changes ranked by bytes saved. Use when page load is slow or a size budget is blown.

````markdown
<context>
JavaScript is the most expensive byte on the web: it has to be downloaded, parsed and executed before the page responds. Most bundles carry avoidable weight: whole libraries imported for one function, duplicate versions, code for routes the user has not visited, and polyfills for browsers the app does not support. Savings only count when measured on the production build, compressed.
</context>

<task>
Reduce the bundle size of: [TARGET]
Budget: as small as the changes below allow; report the savings.

1. Identify the bundler and build. Produce a production build and record the baseline: the initial JavaScript loaded by the target page and the total, both compressed (gzip or brotli, whichever the server uses).
2. Generate a bundle analysis with the tool that fits the bundler (for example a bundle visualizer plugin, the bundler's stats output, or source-map-explorer).
3. List the largest contributors and classify each: needed on first load, needed only later or on another route, duplicated, imported wholesale but used partly, polyfill or dead code that was not tree-shaken, or a large asset inlined into JavaScript.
4. Fix in order of bytes saved per effort: lazy-load routes and heavy components with dynamic imports, switch to per-function or ESM imports, deduplicate versions, drop polyfills outside the supported browser list, and mark side-effect-free packages so they tree-shake.
5. Rebuild after each change and record the size difference. Run the tests and check that the affected pages still work.
</task>

<constraints>
- Do not remove features or change behaviour to save bytes.
- Replacing a dependency with another is a proposal, not a change, unless the swap is trivial and fully covered by tests.
- Report compressed sizes from real builds. Never estimate savings you did not build.
- Keep lazy-loading changes from causing layout shift or an empty screen; add a loading state where one is needed.
- Before saying the work is done, run the check that proves it (tests, build, type check or the command the user gave) and report the real result.
- If you could not run a check, say so plainly and say which one.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Result
One line: initial JavaScript before and after (compressed), total before and after, and whether the budget is met.
## Biggest contributors
Table: module or package, compressed size, classification.
## Changes made
Numbered: change — bytes saved (compressed) — verification.
## Proposals not applied
Bullets: proposal — expected saving as a hypothesis — trade-off.
## How to measure again
The exact commands.
</output_format>
````

---

<a id="reduce-mobile-battery-drain"></a>

## Reduce mobile battery drain

`reduce-mobile-battery-drain` · prompt · Performance · https://hermes-ide.com/prompts/reduce-mobile-battery-drain

Finds what drains battery in a mobile app (wakelocks, location, polling, background work, radio wake-ups, hidden animation) from energy reports and code, and fixes it with platform APIs.

````markdown
<context>
Users say the app ([PLATFORM]) drains their battery, or the store's vitals flag it. Energy goes to four places: CPU kept awake, radios (cellular is the most expensive: each wake-up keeps the radio in a high-power state for several seconds after the last byte), location hardware (GPS far more than network or significant-change location), and the screen and GPU (animations and high refresh rates). Most drain comes from small things repeated: a timer polling every 30 seconds, a wake lock that is not released on an error path, continuous high-accuracy location when "city-level" would do, background sync that ignores battery and network state, animations running on screens nobody sees, and analytics sending one event per request.

Both platforms punish this: Android Doze, App Standby Buckets and background execution limits; iOS background modes, Background App Refresh budgets and termination of apps that overrun their background time.
</context>

<task>
<evidence>
[EVIDENCE]
</evidence>

1. Summarise the evidence: which metric is high (wake-ups, wake-lock time, background CPU, network bytes or requests, location time, foreground energy), when (foreground, background, overnight) and on which app versions or devices.
2. List suspects from the evidence and code, each tied to an energy source: CPU (wake locks, tight loops, timers, background threads), radio (polling, chatty uploads, many small requests, no batching, retries without backoff), location (accuracy, update interval, never stopping updates, background location), screen and GPU (off-screen animations, 120 Hz where unneeded, video autoplay).
3. Fix each with the platform-appropriate mechanism:
   - Android: WorkManager with constraints (unmetered network, charging, battery not low) instead of AlarmManager loops or services; release wake locks in `finally` with a timeout; Fused Location Provider with balanced or low-power priority and longer intervals; FCM high-priority messages only for user-visible events; batch network calls.
   - iOS: BGTaskScheduler (app refresh and processing tasks) instead of keeping the app alive; `URLSession` background sessions with discretionary transfers; significant-location-change or region monitoring instead of continuous updates, `pausesLocationUpdatesAutomatically`, reduced accuracy where enough; silent push sparingly; invalidate timers and stop display links when views disappear.
   - Both: replace polling with push or server-sent change feeds where feasible; batch analytics; back off exponentially; pause work in low-power mode.
4. Show the code changes for the top suspects.
5. Give the verification plan: reproduce on a real device unplugged, a fixed scenario (for example one hour background with the screen off), measure with Battery Historian or Android Studio Energy Profiler, Xcode energy gauges, Instruments, or MetricKit reports, compare before and after, then watch vitals after release.

If there is no evidence of where the drain comes from and no relevant code, give the measurement steps first and ask for the results.
</task>

<constraints>
- Do not claim a fix reduces drain by a specific percentage without measurement; give expected direction and why.
- Respect the feature: if the product needs precise background location, say so and minimise cost within that need rather than removing it.
- Do not invent API names; mark uncertain ones [CHECK DOCS].
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Evidence summary
Bullets: metric, when, where.
## Suspects
Table: suspect | energy source | evidence | confidence (high, medium, low).
## Fixes
For each suspect: the change, the platform API, and the code.
## Verify on device
Numbered steps with the scenario, tools and what to compare.
</output_format>
````

---

<a id="set-performance-budgets"></a>

## Set performance budgets

`set-performance-budgets` · prompt · Performance · https://hermes-ide.com/prompts/set-performance-budgets

Sets performance budgets (bundle size, LCP, INP, API p95, startup time, memory) from baseline data and wires them into CI so regressions fail the build with a clear message. Use to stop slow creep.

````markdown
<context>
The user's web product gets slower release by release because nobody notices a few kilobytes or milliseconds at a time. Budgets fix that only if they are tied to user experience, set from real baselines, measured the same way every time, and enforced where changes happen (the pull request), with a message that says what grew and how to fix it.

Common mistakes: budgets copied from a blog rather than set from the product's own numbers; lab metrics treated as if they were field metrics; gates so noisy (single Lighthouse run on shared CI) that engineers learn to ignore them; budgets only on totals so nobody knows which change caused the regression; and no owner or process when a budget must be raised.

Reference thresholds that are public standards: Core Web Vitals "good" is LCP at or under 2.5 s, INP at or under 200 ms and CLS at or under 0.1, at the 75th percentile of field page loads.
</context>

<task>
<baseline_metrics>
[BASELINE_METRICS]
</baseline_metrics>

1. Pick 4-7 budgets that matter for this web product, mixing outcome metrics (what users feel) and proxy metrics (what CI can measure reliably):
   - web: field LCP, INP, CLS at p75; lab proxies such as JavaScript bytes per route (compressed), total transfer, number of requests, Lighthouse performance score as a soft signal, time to first byte.
   - mobile: cold start time on a reference low-end device, app download and install size, memory at a key screen, frame rendering (slow or frozen frames).
   - api: p95 and p99 latency per critical endpoint at a stated load, error rate, payload size, queries per request.
2. Set each budget from the baseline: hold the line at the current value plus a small tolerance for metrics already good; for metrics that are poor, set the current value as the gate and a separate target with a date. Show the arithmetic.
3. Define measurement so it is stable: tool, environment, device or throttling profile, number of runs (for example median of 3-5 Lighthouse runs), and which percentile. Separate gates (lab, deterministic: bundle size, request count, query count) from monitors (field, noisy: alert and review, not block).
4. Wire it into CI for the user's system with concrete config: for example bundle-size checks (size-limit or bundler stats with a comparison), Lighthouse CI assertions with a budget file, a mobile startup benchmark job, or a load-test threshold for API latency. Report the change per pull request as a comment.
5. Write the failure message template: which budget, the base and new value, the delta, the files or dependencies that grew, and links to the fix guide.
6. Define the exception process: who can raise a budget, the evidence needed, and that raises are recorded with a reason.
7. Set the review cadence: monthly check of field data against targets; tighten budgets when improvements land.

If baselines are missing for a metric, say how to collect them and leave that budget as [TO SET AFTER BASELINE].
</task>

<constraints>
- Every budget number traces to the user's baseline or to a named public standard; never invent baselines.
- Do not gate pull requests on noisy field metrics; gate on deterministic lab proxies.
- Keep the number of budgets small enough that each has an owner.
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Budgets
Table: metric | baseline | budget (gate) | target and date | gate or monitor | owner (role).
## How each is measured
Table: metric | tool | environment and profile | runs and percentile.
## CI wiring
Config files and pipeline steps as code blocks.
## When a budget fails
The message template, then the exception process.
## Review cadence
Bullets.
</output_format>
````

---

<a id="shrink-firmware-footprint"></a>

## Shrink firmware flash and RAM use

`shrink-firmware-footprint` · prompt · Performance · https://hermes-ide.com/prompts/shrink-firmware-footprint

Reads a linker map and size report to cut firmware flash and RAM use, from the biggest symbols and pulled-in libraries to stack sizing and flags such as -Os and LTO. Use when out of memory.

````markdown
<context>
The user's firmware no longer fits, or has no room for the next feature or an over-the-air update slot. MCU: not given; infer from the map and say what you assumed. Flash holds `.text`, `.rodata` and the initial values of `.data`; RAM holds `.data`, `.bss`, heap and stacks. The usual big wins are not clever code changes but things pulled in by accident: `printf` with floating point support, C++ exceptions, RTTI and iostream, the full newlib instead of newlib-nano, software floating point routines because one `float` or `double` crept in, unused vendor HAL modules, large lookup tables or fonts copied into RAM because they were not `const`, and debug strings and asserts left in release builds.

The trap is shrinking size and breaking timing or behaviour: `-Os` and LTO can change timing of busy-wait loops, inline differently in interrupt handlers, and expose undefined behaviour; stack sizes cut without measuring cause rare crashes in the field.
</context>

<task>
<map_or_size_report>
[MAP_OR_SIZE_REPORT]
</map_or_size_report>

1. Summarise the totals: flash used and free, RAM used and free (static), and the share of each section. Note whether heap and stacks are reserved statically.
2. List the 15-20 largest symbols and the libraries they come from (application, vendor HAL, RTOS, C library, compiler runtime such as soft-float or division helpers). Group by origin and give each group's bytes.
3. Identify accidental inclusions and their likely trigger: printf or scanf family with float support, `malloc` pulling in the allocator, exceptions and unwinding tables, RTTI, static constructors, double-precision maths in single-precision code, `sprintf` used for one number, assert strings with file names.
4. Propose quick wins with expected bytes saved (ranges): `-Os` or `-Oz` where the compiler supports it, `-ffunction-sections -fdata-sections` with `--gc-sections`, LTO, newlib-nano (`--specs=nano.specs`) and dropping `_printf_float` if not needed, `-fno-exceptions -fno-rtti` for C++, `-fsingle-precision-constant` or explicit `f` suffixes, removing unused HAL modules, compiling out logs and asserts in release.
5. Propose code changes for what remains: mark tables `const` so they stay in flash, smaller types and bitfields for large arrays of structs, replace a heavy library call with a small purpose-built one, deduplicate strings, compress large assets, move rarely used code to a bootloader-shared or external flash region if the hardware supports it.
6. RAM and stack: find the largest `.bss` and `.data` objects, check buffer sizes against real need, and size each task or main stack from measured high-water marks (stack painting, the RTOS's stack high-water API, or `-fstack-usage` and call-graph analysis) plus a stated margin (for example 20-25%).
7. List risks and checks for each change: timing-sensitive code, interrupt latency, behaviour changes under LTO, and the regression tests or hardware checks to run.

If the input lacks symbol-level sizes, give the exact command to produce them for the toolchain and stop.
</task>

<constraints>
- Every saving is an estimate until rebuilt; give ranges and say which are typical rather than measured.
- Never reduce a stack or buffer without measured usage and a margin.
- Do not invent symbol names or sizes not in the report.
- Do not recommend removing safety checks (watchdog, bounds checks on external input) to save bytes.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Where the bytes go
Table: section | bytes | share | notes. Then a table: symbol or group | origin | bytes.
## Quick wins
Table: change | flag or setting | estimated bytes saved | risk.
## Code changes
Bullets with snippets where useful.
## RAM and stack
Table: object or stack | current bytes | measured or estimated need | proposed.
## Risks and checks
Bullets.
</output_format>
````

---

<a id="speed-up-app-cold-start"></a>

## Speed up app cold start

`speed-up-app-cold-start` · prompt · Performance · https://hermes-ide.com/prompts/speed-up-app-cold-start

Measures and cuts mobile app cold start by deferring SDK setup, removing main-thread I/O and using baseline profiles, measured on a low-end device. Use when an app is slow to open.

````markdown
<context>
The user's [PLATFORM] app is slow to open. Cold start (process not in memory) is the case that matters most and the one developers rarely see on their fast test phones. Rough reference points: Android vitals treat a cold start of 5 seconds or more as excessive, and users notice anything above about 2 seconds; Apple advises that the first frame should appear within about 400 ms of launch. Measure on a low-end device, release build, cold, several runs.

Startup is usually slow because of work that does not need to happen before the first useful screen: initialising every SDK (analytics, crash reporting, ads, feature flags, A/B testing) synchronously, dependency injection graphs built eagerly, disk and database reads or migrations on the main thread, network calls awaited before rendering, heavy first layouts, large JavaScript bundles or Dart isolate setup, and a splash screen used to hide all of it.
</context>

<task>
<startup_trace>
[STARTUP_TRACE]
</startup_trace>

1. Record the baseline: cold, warm and hot start times if available, device model, OS version, build type, number of runs and spread. If they are missing or from a debug build or emulator, give the measurement method first: `adb shell am start -W` and Macrobenchmark `StartupTimingMetric` on Android, Instruments App Launch template and `XCTApplicationLaunchMetric` on iOS, `flutter run --trace-startup --profile` for Flutter, and native markers plus Hermes profiling for React Native.
2. Lay out the timeline in phases: process start and runtime init (pre-main on iOS, Application.onCreate and content providers on Android, JS bundle load or Dart init), first activity or scene creation, first frame, and first meaningful content (time to full display). Assign the measured time to each phase.
3. For each item of work before first frame, classify it: required for the first screen, can be deferred until after first frame, can be lazy (on first use), can run on a background thread, or can be removed.
4. Propose fixes ranked by milliseconds saved on the low-end device:
   - Defer and lazily initialise SDKs; on Android check content-provider auto-init and use the App Startup library to control order; on iOS reduce dynamic frameworks and work in `+load` or static initialisers.
   - Move disk, database, preference and keychain reads off the main thread, or make them lazy.
   - Render the first screen from cached or placeholder data; never await the network before first frame.
   - Simplify the first layout; avoid inflating hidden screens.
   - Android: Baseline Profiles and, where relevant, R8 optimisation. iOS: fewer dynamic libraries, avoid heavy work in `application(_:didFinishLaunchingWithOptions:)`. React Native: Hermes, inline requires or lazy modules, smaller bundle. Flutter: defer plugin initialisation, deferred components, avoid heavy work before `runApp`.
   - Use the platform splash screen API only to cover real minimal work, not to hide slowness.
5. Show the code changes for the top fixes as diffs or before-and-after snippets.
6. Give the guard: a startup benchmark in CI or on a device farm on a fixed low-end device, with a regression threshold, and production monitoring of start times (Android vitals, MetricKit, or the team's performance monitoring).
</task>

<constraints>
- Never use debug builds or emulators for the numbers you report; say when the user's numbers are from one.
- Savings are estimates until measured on the device; give ranges.
- Deferring an SDK must keep its function: say what is lost (for example crash reports during the first second) and whether that is acceptable.
- Do not invent SDK APIs; mark uncertain ones [CHECK DOCS].
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Baseline
Table: start type | device | build | median ms | runs. Or the measurement method if missing.
## Startup timeline
Table: phase | ms | main work in it.
## Ranked fixes
Table: fix | phase | estimated ms saved | effort | trade-off.
## Changes
Code for the top fixes.
## Measure and guard
Benchmark setup, threshold and production metric.
</output_format>
````

---

<a id="speed-up-dataframe-code"></a>

## Speed up dataframe code

`speed-up-dataframe-code` · prompt · Performance · https://hermes-ide.com/prompts/speed-up-dataframe-code

Speeds up slow pandas or Polars code by replacing row loops and apply with vectorised operations, fixing dtypes, chunking or pushing work to the database, measured on your data size.

````markdown
<context>
The user writes data code to get answers, not to be a performance engineer, and a step that takes minutes or runs out of memory is blocking their work. Data size: not given; ask if it changes the advice.

Dataframe code is almost always slow for a handful of reasons: Python-level loops (`iterrows`, `itertuples` in a loop, `apply(axis=1)`, list comprehensions over rows) instead of column operations; repeated concatenation in a loop (quadratic); `object` dtype strings and Python objects where categoricals, numeric or Arrow-backed types would do; reading the whole file when only some columns or rows are needed; merges that explode rows because of duplicate keys; and `groupby().apply` with a Python function where a built-in aggregation exists. The fix must give the same answer: subtle differences in NaN handling, integer overflow, sort order and time zones are the usual way a "faster" version is wrong.
</context>

<task>
<code>
[CODE]
</code>

1. Explain why it is slow, pointing at the exact lines, and estimate the complexity (for example "Python function called once per row: 12 million calls").
2. Rewrite it, in order of impact:
   - Replace row loops and `apply(axis=1)` with vectorised column operations, `np.where` or `np.select` for conditionals, `.str` and `.dt` accessors, `map` with a dict or a merge for lookups, `groupby().agg` or `transform` with built-in functions, and `cumsum`, `shift` or `rolling` for running calculations.
   - Build lists and concatenate once instead of appending in a loop.
   - Fix dtypes: categoricals for repeated strings, downcast numerics where safe, parse dates once on read, nullable or Arrow-backed dtypes where helpful.
   - Read less: `usecols`, `dtype` on read, filters pushed into the reader, Parquet instead of CSV for repeated reads.
   - If the data is larger than memory or the operation is heavy, show the Polars lazy equivalent (with `scan_csv` or `scan_parquet`, and `collect`), DuckDB SQL over the file, chunked processing, or pushing the aggregation into the source database, and say which fits their size.
3. Keep the result identical, or state each intentional difference.
4. Give an equivalence check: run old and new on a sample and compare with `pandas.testing.assert_frame_equal` (or Polars `assert_frame_equal`), with tolerance for floats and explicit sorting.
5. Give the timing plan: time both versions on a representative sample (for example 1% and 10% of rows) to see how runtime grows, with `%timeit` or `time.perf_counter`, and peak memory with `memory_usage(deep=True)` or a memory profiler; then the full run. Do not claim a speedup you have not measured; give the expected order of magnitude as an estimate.
6. Say what to do if it is still too slow: profile with a line profiler, check for an exploding merge, move to Polars or DuckDB, or run on a bigger machine.

If the code depends on functions or columns not shown, ask for them or mark assumptions [ASSUMED].
</task>

<constraints>
- The rewrite must produce the same output as the original on the same input, including NaN handling, dtypes and row order, unless a difference is stated.
- Keep the code readable for an analyst; comment non-obvious vectorised tricks in one line.
- Do not invent column names or data values.
- Do only what was asked. If you notice something else worth changing, mention it in one line at the end instead of changing it.
- Keep the change as small as it can be while still being correct.
</constraints>

<output_format>
## Why it is slow
Bullets with line references and rough call counts.
## Faster version
The rewritten code, in one block.
## Equivalence check
Code that compares old and new.
## Timing plan
Code and what to record; estimated gain labelled as an estimate.
## If still too slow
Up to three options with when to choose each.
</output_format>
````

---

<a id="tune-garbage-collector"></a>

## Tune a garbage collector

`tune-garbage-collector` · prompt · Performance · https://hermes-ide.com/prompts/tune-garbage-collector

Diagnoses GC pauses and memory churn on the JVM, .NET or Go from GC logs and metrics, fixes allocation hotspots first, then chooses collector and heap settings. Use for latency spikes.

````markdown
<context>
A [RUNTIME] service shows latency spikes, high CPU in the collector, or out-of-memory kills, and the user suspects garbage collection. Latency goal: not stated; propose one tied to the service's latency SLO.

What an expert knows: first confirm GC is actually the cause by lining up pause timestamps with latency spikes; then reduce the allocation rate, because the cheapest collection is the one that does not happen; only then tune. Flag-tuning without evidence (copying a list of JVM flags from a blog) usually makes things worse. Container limits matter: a heap sized near the container limit leaves no room for metaspace, thread stacks, direct buffers or the Go runtime overhead, and the kernel kills the process instead of the runtime collecting.
</context>

<task>
<gc_logs>
[GC_LOGS]
</gc_logs>

1. Diagnose from the evidence: collector in use, heap size versus live data after collection, allocation rate (MB/s), promotion rate, pause count, duration distribution (p50, p99, max) and type (young, mixed, full; gen0, gen1, gen2 and background; Go stop-the-world phases and assist time), GC CPU share, and correlation with latency spikes. State whether GC explains the latency problem, partly or not at all.
2. Name the pattern: high allocation churn (short-lived objects), premature promotion, a heap too small for the live set, humongous or large-object allocations, full or compacting collections from fragmentation, a leak (live set growing after each collection, hand over to leak investigation), or container limits causing out-of-memory kills.
3. Allocation fixes first: find the hotspots with an allocation profiler (JFR allocation events or async-profiler alloc mode; dotnet-trace or PerfView allocation tick; Go pprof `-sample_index=alloc_space`) and name typical fixes: reuse buffers, avoid boxing and autoboxing, stream instead of materialising large collections, pre-size collections, pool large buffers (ArrayPool, sync.Pool), avoid string building in hot logs, and reduce large object heap allocations in .NET.
4. Then settings, only those the evidence supports:
   - JVM: choose the collector by goal (G1 as default, ZGC or Shenandoah for low pause at larger heaps, Parallel for throughput batch jobs); set -Xms equal to -Xmx for steady services or use -XX:MaxRAMPercentage in containers; G1 pause target only with evidence; avoid many tuning flags at once.
   - .NET: Server versus Workstation GC, concurrent or background GC, `GCHeapHardLimit` or percentage in containers, `GCConserveMemory`, DATAS where available; region or LOH settings only with evidence.
   - Go: `GOGC` for the CPU versus memory trade-off and `GOMEMLIMIT` set below the container limit with headroom; check `GOMAXPROCS` matches the CPU quota.
5. Experiment plan: change one setting at a time, test under representative load (replayed traffic or a load test at production rate), compare pause percentiles, GC CPU share, memory footprint and service p99, and roll out to one instance before all.
6. List what to watch after the change and the rollback trigger.

If the logs lack timestamps or heap sizes, say which logging flags to enable and stop.
</task>

<constraints>
- No settings change without a diagnosis that justifies it; no blanket flag lists.
- Respect container memory limits and leave headroom for non-heap memory.
- Do not state runtime defaults or flag behaviour you are unsure of for the user's version; ask for the version or mark [CHECK FOR YOUR VERSION].
- Separate what you verified from what you inferred. Mark inferences as such.
- When you do not know, say "I don't know" once and state what would settle it.
</constraints>

<output_format>
## Diagnosis
Table: metric | value | what it means. Then one paragraph: is GC the cause, and which pattern.
## Allocation fixes
Ranked bullets: hotspot or likely hotspot, fix, how to confirm with a profiler.
## Settings
Table: setting | current | proposed | why | risk.
## Experiment plan
Numbered steps.
## Watch after
Metrics, thresholds and the rollback trigger.
</output_format>
````

---

<a id="write-k6-load-test"></a>

## Write a k6 load test

`write-k6-load-test` · prompt · Performance · https://hermes-ide.com/prompts/write-k6-load-test

Writes a k6 load test script from a workload model, with arrival-rate scenarios, thresholds that fail the run, test data and tagged metrics. Use when the load plan is settled and you need the script.

````markdown
<context>
This prompt turns an agreed workload model into a k6 script. (To design the model itself, use a load test planning prompt first.) The k6 details that decide whether results mean anything:
- Arrival-rate executors (`constant-arrival-rate`, `ramping-arrival-rate`) model an open system, so a slow server does not quietly reduce the load. Looping virtual-user executors do, which hides saturation.
- `preAllocatedVUs` and `maxVUs` must cover rate times response time; when they do not, k6 reports `dropped_iterations` and the generator, not the server, was the limit.
- Thresholds in `options.thresholds` make the run pass or fail automatically; `abortOnFail` stops a run that is already lost.
- Tagging requests with a stable `name` keeps URLs with ids from exploding into thousands of metric series.
- `SharedArray` loads test data once instead of once per virtual user.
</context>

<task>
Write a k6 script for these requests:
<endpoints>
[ENDPOINTS]
</endpoints>
using this workload model:
<workload>
[WORKLOAD]
</workload>

1. If the model lacks a rate, a mix or a threshold, ask for it and stop; do not invent the goal of the test. Minor gaps (think time, ramp length) can be filled with a stated assumption.
2. One scenario per traffic shape in the model (for example smoke, peak, stress, soak), selected with an environment variable such as `__ENV.SCENARIO` so one file serves all. Each user-facing scenario uses an arrival-rate executor with `preAllocatedVUs` and `maxVUs` sized from the target rate and expected latency, showing the arithmetic in a comment.
3. Implement the transaction mix by weight inside the default function or as separate `exec` functions per scenario. Add think time only where the model has it.
4. Authentication happens once in `setup()` where tokens can be shared, or per virtual user when sessions must be distinct. If a token expires before the longest scenario ends, refresh it instead of letting the run fill with 401s. Secrets and the base URL come from `__ENV`, never the script.
5. Load varied test data from CSV or JSON through `SharedArray` (with papaparse for CSV), enough rows to defeat caching the way production traffic does.
6. Every request gets a `name` tag, a `check` on status and one meaningful body property, and a `group` or scenario tag that matches the model's transaction names. `http_req_failed` counts any status outside 200-399 as a failure; if the model expects one (a 404 probe, a 409 on a duplicate), declare it with `http.expectedStatuses` for that request.
7. Thresholds: p95 and p99 per transaction via tagged metrics (`http_req_duration{name:checkout}`), `http_req_failed` rate, `checks` rate, and `dropped_iterations` count equal to 0. Use `abortOnFail` with a delay on the error-rate threshold.
8. Add `handleSummary` only if the user wants a file report; otherwise rely on the standard summary.
</task>

<constraints>
- Use only the k6 standard modules (`k6`, `k6/http`, `k6/data`, `k6/metrics`, `k6/execution`) plus the papaparse remote module from jslib.k6.io for CSV. k6 runs scripts in its own JavaScript runtime, not Node, so Node packages such as axios or `fs` do not work, and pure JavaScript libraries would need a bundling step this script avoids.
- Do not invent endpoints, payload fields or status codes; mark gaps as placeholders.
- Never default the base URL to a production host. Warn if the endpoints call third-party services that must be stubbed.
</constraints>

<output_format>
## Script
One fenced `javascript` block, complete and runnable.
## Test data
The data file format with three example rows using fake values.
## How to run
`k6 run` commands per scenario with the environment variables.
## Reading the results
Five bullets: which numbers decide pass or fail, what `dropped_iterations` means, and which server-side metrics to watch alongside.
## Placeholders
List of values to fill in, or "None".
</output_format>
````
