Your agent can pass one test and still skip approval on the next run, use the wrong customer or record, or execute the same action twice after a model, prompt, or tool change. Konsista replays the critical procedure against mocked tools and shows exactly what drifted before the change reaches production.
An agent can return a convincing answer while using the wrong account, skipping approval, ignoring a lookup, or acting twice. The risk reopens whenever the team changes a model, prompt, tool schema, connector, or policy — and one successful run gives no evidence that the procedure will hold again.
Konsista gives the release owner literal evidence that required steps, critical targets, and action limits held across repeated runs before the change ships.
Output evals can grade the answer. Konsista checks the action path: what the agent looked up, whether it waited for approval, which target it used, and how many times it acted.
A speed-tuned GPT-5 agent asked for approval before delete 20/20 runs on baseline — then, after the change, 0/20. The target record stayed correct, but the required gate disappeared on every candidate run.
See the full report ↓Describe what the agent may do, what must happen first, and which recipient, record, amount, or destination must stay correct.
Repeat the same scenario against mocked tools, so no real message, delete, charge, submission, or record update occurs.
Review which required step disappeared, which critical target changed, or which action repeated — with evidence for every run.
Each guide turns one production-facing risk into observable checks and a repeatable mocked scenario.
Required tools, sequence, critical arguments, action counts, and evidence consistency.
Human approval before controlled actions, decision enforcement, and target binding.
Repeated mocked runs and literal assertions without an LLM judge for the verdict.
Skipped gates, wrong targets, missing tools, duplicate actions, and unsafe stability.
Concrete message, CRM, refund, payment, claim, booking, and delete procedures.
Exercise declared action boundaries before release while keeping runtime controls separate.
Literal evidence, deterministic findings, repeat stability, JSON artifacts, and coverage.
Every finding traces back to observed tool calls and critical values, so the team can reproduce the failure and decide whether to block the release.
request_approval and delete_record mocked.
request_approval(customer-4829) → delete_record(customer-4829)delete_record(customer-4829) without approval first in all 20 runs
customer-4829. The failure was procedural, not argument-level: the required approval step disappeared in 20 of 20 candidate runs.Public GitHub repos are not customers. We use them as evidence that the harness works beyond our own demo: it catches declared procedure failures, and it also records passed scenarios when the declared action path holds.
Public LangGraph-style support email workflow: the full required procedure held 1 / 4 candidate repeats while baseline held 4 / 4.
Public sales-agent workflow: selected the expected leads, created outreach drafts, sent one digest, and updated the processed-contact log 4 / 4 repeats.
Public CRM-agent workflow: qualified the lead, saved the CRM contact, and created the expected follow-up task 4 / 4 repeats.
Public browser-action example: search, add-to-cart, approval, and checkout held 4 / 4 repeats with mocked action boundaries.
You do not need a complete test suite to begin. Start with one action, the check or approval that must happen first, and the target that must remain correct. We shape the first mocked run ourselves.
Agents that watch accounts, draft outreach, recommend exchanges, or message customers proactively.
Agents that score account health, flag risk, qualify leads, update records, or create follow-up tasks.
Agents that submit, charge, refund, book, or trigger an external operation after a required check.
Run AI agent regression testing before any change that can move the action path:
change prompt / model / tool → repeat the critical scenarios → compare baseline vs candidate → report: held / drifted / failed → release decision.
Tools like LangSmith, Braintrust, and Promptfoo help evaluate prompts, outputs, and model behavior.
Tools like Langfuse, LangSmith, and Helicone help trace what happened once the agent is running.
Re-runs the action scenario before release and checks whether approvals, required tools, targets, and limits held.
Describe one risky action, the check that must happen first, and the target that must remain correct. We'll shape the first mocked run and send back evidence for every repeat: held, drifted, or failed.
Good examples: “ask for approval before deleting a record”, “look up the customer before writing to CRM”, or “never refund above $500 without manager approval”.
Prefer email? hello@konsista.com
Runs in your environment · we never see your agent's internals or raw data without explicit permission.
Privacy: submissions are used only to evaluate fit for a mocked first run. Read the privacy note.
Built by Tatiana Panfilova, founder and product engineer. I review early scenarios personally and reply when a mocked first run is a fit. LinkedIn
Current design-partner capacity: up to two full mocked first runs per week, so accepted scenarios stay hands-on and detailed.
We read every scenario and reply personally when it fits the current design-partner queue. Full mocked first runs are limited to two per week.