Every system gets a twin.
Same API, same state, same webhooks. Your agent's code doesn't change, only where its calls go.
Vektra gives your agent a stateful twin of every system it touches. Break it there. Ship only what holds.
Same API, same state, same webhooks. Your agent's code doesn't change, only where its calls go.
Late webhooks, 429 bursts, a person editing the same ticket. Run every scenario in parallel against the twins.
Invariants are checked after every step. A failure replays from its seed, identically, until the fix holds.
The plan runs in a twin forked from live records first. You see the diff, then it goes live.
3D illustration · simulated
A support agent issues refunds against twins of Stripe and Zendesk. Turn on a fault, run it, replay the seed, then ship the fix and watch every invariant hold.
Not run yet.
Simulated in your browser to show how Vektra reports a run. Not connected to Stripe or Zendesk.
Late webhooks, 429 bursts, duplicate refunds, a person editing the same ticket, retries without backoff, partial failures, stale reads, missing idempotency keys, pagination drift.
This year an agent deleted a company's production database in nine seconds. It had been working in staging. Here is the public record, replayed.
00:09.00elapsed
Working in staging, a coding agent hits a credential mismatch and decides to fix it. Its API token is fully permissioned.
One API call deletes the production database and its volume backups.
The service is back, running on a three-month-old backup.
Sources: OECD.AI incident record [1], Decrypt [2], Computing [3]
1 in 3
of 1,340 teams surveyed named quality as their main blocker to putting agents in production.
37%
run online evaluations on how their agents behave once live.
Source: LangChain, State of Agent Engineering, survey of 18 Nov – 2 Dec 2025 [4]
Mocks answer one call with a canned reply. Vendor sandboxes cover one system at a time. Neither can hold a refund webhook back forty seconds while a support rep edits the same ticket. That's where agents break, and that's what Vektra rehearses.
For bulk or irreversible actions, Vektra copies only the records a plan touches into a twin, runs the plan there, and holds every write until someone approves it.
Only the records the plan touches are copied, read-only, at the moment the plan is made.
Choose which actions need a plan, above what amount, and who can approve it.
Every applied write is logged with its reversal where one exists. An email that was sent stays sent, so the plan says so before you approve.
Each piece exists to turn a once-a-month production incident into a test that fails on every pull request until it's fixed.
A refund changes the charge, queues a webhook and shows up when the agent reads it back. Twins remember, across systems.
Turn production's bad days into switches.
Same seed, same trace, every time.
Rules that must hold after every step.
Copy only what a plan touches.
Every prompt, tool or model change runs the scenarios before it merges. A broken invariant blocks the merge, with the seed attached so anyone can replay it.
The check names and counts here are illustrative.
1# Before: real endpoints 2# STRIPE_API_BASE=https://api.stripe.com 3 4# After: the same API, served by twins 5STRIPE_API_BASE=$VEKTRA_TWIN_URL/stripe 6SHOPIFY_API_BASE=$VEKTRA_TWIN_URL/shopify 7ZENDESK_API_BASE=$VEKTRA_TWIN_URL/zendesk 8MCP_SERVER_URL=$VEKTRA_TWIN_URL/mcp
1import { invariant } from "@vektra/sdk"; 2 3// Never refund more than was charged. 4invariant("refund_total <= charge_total", ({ stripe }) => 5 stripe.charges.every((c) => c.amountRefunded <= c.amount)); 6 7// One refund per support ticket. 8invariant("one refund per ticket", ({ stripe, zendesk }) => 9 zendesk.tickets.every((t) => 10 stripe.refunds.filter((r) => r.metadata.ticket === t.id).length <= 1)); 11 12// Leave tickets a person has put on hold. 13invariant("respects human holds", ({ zendesk, actions }) => 14 actions.every((a) => !zendesk.isHeldByHuman(a.ticket)));
1name: agent 2on: [pull_request] 3jobs: 4 rehearse: 5 runs-on: ubuntu-latest 6 steps: 7 - uses: actions/checkout@v4 8 - uses: vektra/rehearse@v0 9 with: 10 agent: ./agents/refunds 11 twins: stripe, zendesk 12 faults: webhook-lag, rate-limit, human-edits 13 fail-on: invariant
Nothing here is generally available. Design partners decide the order.
Missing yours? Ask it in the pilot request below.
A mock returns a canned response to one call. A twin keeps state across calls: a refund changes the charge, fires a webhook, and shows up when the agent reads the charge back. Multi-step agents fail in exactly those gaps.
A sandbox covers one vendor, and it isn't built to hold a webhook back on purpose or have a person edit the same record mid-run. Twins span the systems your agent touches, reset between runs, and fork so scenarios can run in parallel.
That's the hard part, and it's what we're building with design partners. Each twin is checked against the vendor's own sandbox and recorded traffic. Where a twin doesn't cover a behaviour, it's designed to say so instead of guessing.
No. It's designed for batch and irreversible actions, not every request. You choose which actions need a plan.
Synthetic fixtures by default. The dry-run is designed to read only the records a plan touches, read-only, at the moment the plan is made.
The gateway is built for anything that calls HTTP APIs or MCP tools. Your agent keeps its code; only its endpoints change.
Pricing is being set with design partners during the pilot.
We're working with a few teams whose agents write to money, orders or customer records. If that's you, we'd like to build the twins around your workflows.