Point Reprompt at your pipeline's real execution traces, tell it which model you want to switch to, and it finds — and proves — a prompt that gets you the same results for less.
Every few months a cheaper or better model comes out. But the prompt your team spent weeks tuning was written for the old one — and the exact same words can give you a different answer on the new one. That's not a "your pipeline" problem. It happens to every team, on every model change.
Drop in your pipeline's real execution logs. Reprompt auto-builds a visual dependency graph of every stage, then shows it live and animated as optimization runs — generating, critiquing, refining, scoring — colored by status.
A self-evolving prompt optimizer that mutates, critiques and refines candidates across multiple rounds until they match your original output — judged by an independent AI model, never the one being tested, so it can't grade its own homework.
Deterministic checks, semantic embedding similarity and an AI judge, combined into one composite match score per candidate. A number you can defend, not a vibe.
Test several target models side by side per stage — and pick a different model for each stage. A curated model library with cost, context window and prompt-style cards; bring your own API keys.
See exactly what changed between your original prompt and the winning one, and browse every stage's inputs, outputs and costs across all test runs in a spreadsheet-style explorer.
Auto-generated, human-editable acceptance criteria per stage, with an approval gate before any migration runs — and automatic budget limits so an optimization run can never spend unbounded API costs.
Eight steps, most of them automatic. You approve two decisions — what "correct" means, and which model wins.
You upload real execution traces — actual logs from your pipeline running in production or testing, inputs, prompts, outputs, per stage. Reprompt parses these into a formal pipeline structure automatically: it detects each stage, figures out which stages depend on which others, and builds a visual dependency graph. No manual mapping required.
Your whole pipeline appears as an interactive node graph — every stage, which model it currently uses, token counts, latency. Browse the raw data behind it: every input, output and cost per stage, across every trace you imported, down to any individual run.
Before touching any model, Reprompt auto-generates a rubric per stage — hard rules (must match schema, no hallucinated IDs) plus judged criteria like tone and reasoning quality. You review and approve these first, so the system optimizes toward a real target, not "seems similar."
Choose the cheaper model(s) to migrate to — compare several candidates at once, or assign a different model to each stage if one needs more capability than another. Bring your own key for any provider: OpenAI, Anthropic, Gemini, or self-hosted.
For each stage: generate several prompt variants → cheap first-pass score to cut the weak ones → an independent judge critiques exactly why they failed → rewrite based on that critique → repeat for several rounds → final sweep across phrasing and format → pick the winner. The judge is never the model being tested, so it can't grade its own homework.
An animated version of the pipeline graph shows which stage is working right now and exactly what it's doing — generating, critiquing, refining, running the final sweep — plus a running activity log and the actual reasoning text, in real time. Not a spinner.
A before/after comparison: your original prompt next to the winning one, diff-highlighted, with score, cost and latency for each. Tested multiple models? A side-by-side comparison gives a clear recommendation — not just raw numbers with no verdict.
Budget limits are enforced automatically in the background, so a run can never spend unbounded API cost chasing a perfect score. You don't configure this — it just protects you.
We're onboarding a small group of teams running multi-stage LLM pipelines. Tell us where to reach you.