Before your next model, prompt, or tool change

AI agent regression testing that catches wrong actions before release.

Your agent can pass one test and still skip approval on the next run, use the wrong customer or record, or execute the same action twice after a model, prompt, or tool change. Konsista replays the critical procedure against mocked tools and shows exactly what drifted before the change reaches production.

For engineers releasing agents that send, update, delete, submit, charge, or book · runs in your environment
The release risk

The answer still looks right. The action path no longer is.

An agent can return a convincing answer while using the wrong account, skipping approval, ignoring a lookup, or acting twice. The risk reopens whenever the team changes a model, prompt, tool schema, connector, or policy — and one successful run gives no evidence that the procedure will hold again.

Konsista gives the release owner literal evidence that required steps, critical targets, and action limits held across repeated runs before the change ships.

Output evals can grade the answer. Konsista checks the action path: what the agent looked up, whether it waited for approval, which target it used, and how many times it acted.

Caught on a live run

A speed-tuned GPT-5 agent asked for approval before delete 20/20 runs on baseline — then, after the change, 0/20. The target record stayed correct, but the required gate disappeared on every candidate run.

See the full report ↓
From release risk to evidence

Know what changed before you decide to ship.

01 — Declare

Name the action that cannot drift

Describe what the agent may do, what must happen first, and which recipient, record, amount, or destination must stay correct.

02 — Replay

Compare baseline and candidate safely

Repeat the same scenario against mocked tools, so no real message, delete, charge, submission, or record update occurs.

03 — Decide

See the exact release blocker

Review which required step disappeared, which critical target changed, or which action repeated — with evidence for every run.

Failure-specific guides

Start with the failure your release must prevent.

Each guide turns one production-facing risk into observable checks and a repeatable mocked scenario.

Release evidence

See the exact run where the procedure broke.

Every finding traces back to observed tool calls and critical values, so the team can reproduce the failure and decide whether to block the release.

Public workflow evidence

Validated on public agent workflows — without treating them as endorsements.

Public GitHub repos are not customers. We use them as evidence that the harness works beyond our own demo: it catches declared procedure failures, and it also records passed scenarios when the declared action path holds.

red finding

Support email procedure drift

Public LangGraph-style support email workflow: the full required procedure held 1 / 4 candidate repeats while baseline held 4 / 4.

Open report

passed evidence

Sales lead-finder digest

Public sales-agent workflow: selected the expected leads, created outreach drafts, sent one digest, and updated the processed-contact log 4 / 4 repeats.

Open report

passed evidence

CRM lead routing

Public CRM-agent workflow: qualified the lead, saved the CRM contact, and created the expected follow-up task 4 / 4 repeats.

Open report

compatibility

Browser checkout procedure

Public browser-action example: search, add-to-cart, approval, and checkout held 4 / 4 repeats with mocked action boundaries.

Open report

See all scenario evidence
Where teams start

Start with the action your team is least willing to get wrong.

You do not need a complete test suite to begin. Start with one action, the check or approval that must happen first, and the target that must remain correct. We shape the first mocked run ourselves.

outbound action

Outbound customer messages

Agents that watch accounts, draft outreach, recommend exchanges, or message customers proactively.

  • wrong customer or account
  • outreach before required evidence
  • duplicate message or wrong segment
record write

CRM and account updates

Agents that score account health, flag risk, qualify leads, update records, or create follow-up tasks.

  • wrong account or contact
  • risk label written to wrong record
  • lookup skipped before update
controlled action

Refunds, payments, claims, bookings

Agents that submit, charge, refund, book, or trigger an external operation after a required check.

  • approval or eligibility check skipped
  • wrong record, claim, booking, or amount
  • duplicate controlled action
Before release

Every behavior-changing release can reopen a closed risk.

Run AI agent regression testing before any change that can move the action path:

  • changing the model (including switching to a cheaper / faster one)
  • changing the system prompt
  • adding or removing a tool
  • changing a tool's schema or arguments
  • changing an approval or policy rule
  • shipping a new agent workflow

The loop

change prompt / model / tool → repeat the critical scenarios → compare baseline vs candidate → report: held / drifted / failed → release decision.

Inputs & outputs

Define the expected action. Get evidence for the release.

You provide

  • a scenario (the risky request)
  • the expected tool sequence
  • critical arguments to watch (amount, recipient, target record…)
  • a mock / sandbox tool layer (so repeats are safe)
  • baseline vs candidate config

You get

  • a mechanical procedure report
  • pass/fail for the declared procedure
  • a plain-English drift explanation
  • an explicit coverage statement
  • a JSON artifact for CI / review
The missing release check

Close the gap between answer quality and production behavior.

Output evals

Grade the answer

Tools like LangSmith, Braintrust, and Promptfoo help evaluate prompts, outputs, and model behavior.

question: did the answer look right?
Observability

Watch production

Tools like Langfuse, LangSmith, and Helicone help trace what happened once the agent is running.

question: what happened in production?
Konsista

Check the procedure

Re-runs the action scenario before release and checks whether approvals, required tools, targets, and limits held.

question: did the action happen as declared?
Different layer, not a replacement. Output evals catch quality drift. Observability catches production behavior. Runtime guardrails block live actions. Konsista checks before release whether the declared action procedure still holds.
Primary fit

Built for the engineers responsible for the next agent release.

A strong fit

  • You ship tool-using agents that act: send, delete, purchase, modify data, call external APIs.
  • Your risk is a wrong action, not wording or tone.
  • You want confidence before a release or model/prompt/tool change.
  • You rely on procedures: confirmations, required tools, limits, order, approval gates.

Not the problem we solve

  • Measuring text quality, chatbot tone, or answer "correctness".
  • General LLM evaluation or prompt benchmarking.
  • Legal or compliance certification.
  • Replacing human review of high-stakes decisions.
First release check

Test the action you cannot afford to get wrong.

Describe one risky action, the check that must happen first, and the target that must remain correct. We'll shape the first mocked run and send back evidence for every repeat: held, drifted, or failed.

Good examples: “ask for approval before deleting a record”, “look up the customer before writing to CRM”, or “never refund above $500 without manager approval”.

Prefer email? hello@konsista.com

Runs in your environment · we never see your agent's internals or raw data without explicit permission.

Privacy: submissions are used only to evaluate fit for a mocked first run. Read the privacy note.

Founder-led

Built by Tatiana Panfilova, founder and product engineer. I review early scenarios personally and reply when a mocked first run is a fit. LinkedIn

Current design-partner capacity: up to two full mocked first runs per week, so accepted scenarios stay hands-on and detailed.

Thanks — we received the scenario.

We read every scenario and reply personally when it fits the current design-partner queue. Full mocked first runs are limited to two per week.