AI and ML engineering
Building with models: LLM apps, RAG, evals, agents, training and MLOps.
Download all 42
- Build an MCP server
Implements a Model Context Protocol server exposing the given tools and resources, with input validation, least privilege and error messages a model can act on. Use to connect a system to AI clients.
- Build an LLM structured extraction step
Builds an LLM step that turns documents into schema-valid JSON, with the schema, prompt, validation and repair loop, null handling and an eval set. Use when automating invoices, forms or emails.
- Choose between rules, ML and an LLM
Recommends rules, classical ML, a hosted LLM or a fine-tuned model for a problem, comparing accuracy, cost, latency and maintenance with the reasoning shown. Use before committing to an approach.
- Design an LLM agent architecture
Designs an LLM agent system, deciding first whether an agent is needed, then single or multi-agent, tools, memory, guardrails, human checkpoints, evals and cost limits.
- Design a RAG pipeline
Designs a retrieval-augmented generation pipeline from a corpus and its real questions, covering chunking, hybrid retrieval, reranking, citations and evals. Use before building or rebuilding RAG.
- Design tool definitions for an LLM agent
Designs tool or function definitions for an LLM agent, with names, descriptions, JSON Schema parameters and error returns that models call reliably. Use when exposing an API or capability to an agent.
- Machine-learning engineer
Acts as a machine-learning engineer who starts from the data and a baseline, insists on evals and reproducibility, and distrusts any gain a simpler model explains.
- Plan a fine-tuning project
Decides whether fine-tuning beats prompting or retrieval for a task and, if it does, plans the data, splits, training settings, evaluation against a prompt baseline, and cost.
- Plan a machine-learning experiment
Plans a machine-learning experiment before any training code exists: framing, baselines, leak-proof splits, metrics, ablations and a stop rule. Use when starting a new model or modelling spike.
- Reduce LLM costs and latency
Cuts an LLM feature's cost and latency through prompt trimming, caching, model routing, batching and output limits, each paired with the quality check that proves nothing regressed.
- Review a training dataset sample
Audits a sample of a labelled dataset for label noise, leakage, duplicates, class imbalance and representation gaps, and gives a concrete fix for each problem. Use before training or fine-tuning.
- Write an eval suite for an LLM feature
Writes an eval set for an LLM feature with golden, edge and adversarial cases, graders matched to each criterion, and pass thresholds. Use before shipping or changing a model, prompt or pipeline.
- Write a model card
Writes a model card with intended use, training data, metrics by slice, limitations and ethical considerations from training notes and eval results, flagging gaps. Use before releasing a model.
- Add guardrails to an LLM feature
Adds layered guardrails to an LLM feature with input checks, schema-validated output, content and grounding checks, refusal handling, fallbacks and monitoring. Use before real users see it.
- Analyse sentiment by aspect in reviews
Extracts the aspects a review or comment mentions, such as price, delivery or support, with the sentiment and supporting quote for each, returning structured output for dashboards.
- Answer from retrieved passages with citations
Answers a user question using only the retrieved passages, cites a passage id for every claim and abstains with a fixed phrase when the passages lack the answer. Use as the answer step of a RAG app.
- Check an answer's faithfulness to its sources
Splits an answer into atomic claims, labels each as supported, contradicted or not found in the sources, and returns a faithfulness score with the unsupported claims. Use to catch hallucinations.
- Choose a model for an LLM feature
Chooses a model tier for an LLM feature with a small task-specific eval of quality, latency and cost, plus a decision rule. Use when picking or switching models instead of trusting leaderboards.
- Clean up an automatic speech transcript
Cleans an automatic speech transcript by fixing punctuation, casing, obvious misrecognitions and speaker labels while keeping the words faithful and marking every uncertain fix.
- Compress a conversation into a carry-over state
Compresses a long chat history into a compact state summary of goals, decisions, constraints, open questions and user facts, to carry into a fresh context window without losing what matters.
- Convert a question into safe read-only SQL
Turns a natural-language question into one bounded, read-only SQL query for a given schema, asking for clarification on ambiguity and refusing writes or unbounded scans. Use in text-to-SQL features.
- Critique a draft against requirements and revise it
Critiques a generated draft against stated requirements, lists concrete defects by severity, then produces a revision that fixes them without adding unsupported content. Use as a self-refine step.
- Decide whether an assistant answer needs a human
Decides whether an assistant's draft answer can be sent or must go to a human, checking policy triggers, risk and confidence, and returns a decision with a reason code the app can log.
- Decompose a complex question into sub-questions
Breaks a multi-hop, comparative or aggregate question into ordered sub-questions with dependencies and a composition step, so a retrieval or agent system can answer each one first.
- Detect prompt injection in untrusted content
Classifies untrusted content such as a web page, email, file or tool output for attempts to override instructions, exfiltrate data or trigger tools, and returns a risk level with the suspicious spans.
- Extract durable user preferences for memory
Turns lasting preferences and facts a user explicitly shared in a chat into add, update and delete operations on a memory store, skipping sensitive details unless the user asked to save them.
- Generate synthetic test records from a schema
Generates a synthetic dataset of one record type that hits stated distributions, labels its edge cases and opt-in invalid records with the expected result, and uses no real people's data.
- Grade a response against a rubric
Grades a response against a rubric one criterion at a time, quotes the evidence behind each score, and returns structured scores with an overall pass or fail. Use for LLM evals and automated marking.
- Implement streaming LLM responses
Implements streaming LLM responses end to end, from provider stream to server-sent events to UI, with cancellation, timeouts and mid-stream errors. Use when replies feel slow to start.
- Implement LLM tool calling
Implements tool calling in an LLM feature with tool schemas, a dispatch loop, argument validation, timeouts, limits and safe error handling. Use when wiring a model to functions or APIs.
- Judge two responses side by side
Compares two candidate responses to the same prompt against stated criteria, reasons per criterion before deciding, and returns A, B or tie. Built to be run twice with the order swapped.
- Moderate user content against your policy
Classifies user-generated content against a platform policy the operator supplies, returning violated clauses, severity, quoted evidence and a recommended action. Use as an LLM moderation step.
- Normalise messy records to a canonical form
Normalises messy names, addresses, company names or product titles into a canonical form with confidence and the rules applied, flagging records that need a human. Use in data cleaning pipelines.
- Plan a multi-step task for an agent
Turns a user goal into an executable agent plan of steps with tools, inputs, success checks, approval gates and replanning triggers, for plan-then-execute agent architectures.
- Redact personal data with typed placeholders
Redacts personal data such as names, contact details, ID numbers and health details from text, replacing each with a consistent typed placeholder, and returns the mapping only when asked.
- Rerank retrieved passages by relevance
Scores retrieved passages for how well they answer a query and returns a ranked list with graded relevance and a one-line reason, flagging when none are relevant. Use as an LLM reranker in RAG.
- Rewrite a chat turn into a standalone search query
Rewrites the latest turn of a conversation into a standalone search query, resolving pronouns and earlier context, with optional keyword and semantic variants. Use before retrieval in a chat app.
- Route a user request to the right handler
Classifies a user request into one of the application's declared routes with a confidence score and a short reason, and returns the fallback route when nothing fits. Use as an LLM router.
- Run a tool-using agent loop
System prompt for a tool-using agent that plans the next step, calls one declared tool at a time, checks each result, and stops with an answer, a request for approval or a request for help.
- Suggest follow-up questions after an answer
Suggests short follow-up questions a user might ask next, grounded in the last answer and the app's scope, avoiding repeats and out-of-scope topics. Use for suggestion chips in chat apps.
- Summarise with increasing density
Produces a series of same-length summaries of one document, each adding salient entities the previous one missed while staying faithful, so an application can pick the density it needs.
- Write a hypothetical answer for embedding search
Writes short hypothetical answer passages for a question, styled like the target corpus, to be embedded for retrieval and never shown to users. Use to improve recall in semantic search.
Not: prompts for using an assistant (prompting domain); coding-agent setup (meta).