Staging environments for AI agents

Let your agents rehearse
before they act.

Vektra gives your agent a stateful twin of every system it touches. Break it there. Ship only what holds.

Simulated on this page Actions rehearsed0000 Failures caught000 Writes that reached production000 Move to see the code · click the water
How it works

A twin of every system your agent touches.

01 / 04 · Mirror

Every system gets a twin.

Same API, same state, same webhooks. Your agent's code doesn't change, only where its calls go.

02 / 04 · Rehearse

Below the surface, break everything.

Late webhooks, 429 bursts, a person editing the same ticket. Run every scenario in parallel against the twins.

03 / 04 · Verify

Catch it. Replay it. Fix it.

Invariants are checked after every step. A failure replays from its seed, identically, until the fix holds.

04 / 04 · Approve In design

Only approved writes reach production.

The plan runs in a twin forked from live records first. You see the diff, then it goes live.

3D illustration · simulated

Demo · runs in your browser

Break it here, not in production.

A support agent issues refunds against twins of Stripe and Zendesk. Turn on a fault, run it, replay the seed, then ship the fix and watch every invariant hold.

vektra/refund-agent/Customer writes twice about one refund
TwinProduction
Faults
Agent build

    Twin state t = 0.00s

    Invariants

      Not run yet.

      seed ----·--

      Simulated in your browser to show how Vektra reports a run. Not connected to Stripe or Zendesk.

      Where agents break

      Late webhooks, 429 bursts, duplicate refunds, a person editing the same ticket, retries without backoff, partial failures, stale reads, missing idempotency keys, pagination drift.

      Why now

      Agents now write to production. Their mistakes don't stay in one system.

      This year an agent deleted a company's production database in nine seconds. It had been working in staging. Here is the public record, replayed.

      Public incident · 27 Apr 2026

      00:09.00elapsed

      1. 00:00

        Working in staging, a coding agent hits a credential mismatch and decides to fix it. Its API token is fully permissioned.

      2. 00:09

        One API call deletes the production database and its volume backups.

      3. +30 h

        The service is back, running on a three-month-old backup.

      Sources: OECD.AI incident record [1], Decrypt [2], Computing [3]

      1 in 3

      of 1,340 teams surveyed named quality as their main blocker to putting agents in production.

      37%

      run online evaluations on how their agents behave once live.

      Source: LangChain, State of Agent Engineering, survey of 18 Nov – 2 Dec 2025 [4]

      Mocks answer one call with a canned reply. Vendor sandboxes cover one system at a time. Neither can hold a refund webhook back forty seconds while a support rep edits the same ticket. That's where agents break, and that's what Vektra rehearses.

      Production dry-run In design

      See the plan before your agent touches production.

      For bulk or irreversible actions, Vektra copies only the records a plan touches into a twin, runs the plan there, and holds every write until someone approves it.

      $ vektra plan --agent refund-agentDraft CLI · illustrative
      Forked 40 charges and 40 tickets from production into a twin. Read-only.
        Plan: 38 to add, 0 to change, 0 to destroy. 2 skipped.$0.00
        Slide to apply 38 writes
        Waiting for approval
        • Forked, not mirrored

          Only the records the plan touches are copied, read-only, at the moment the plan is made.

        • You set the threshold

          Choose which actions need a plan, above what amount, and who can approve it.

        • Clear about undo

          Every applied write is logged with its reversal where one exists. An email that was sent stays sent, so the plan says so before you approve.

        What's inside

        Built for the failures you can't reproduce.

        Each piece exists to turn a once-a-month production incident into a test that fails on every pull request until it's fixed.

        Stateful twins

        Building now

        A refund changes the charge, queues a webhook and shows up when the agent reads it back. Twins remember, across systems.

        Stripe twincharge
        ch_1043$86.00
        refunded$0.00
        statussucceeded
        Webhooksqueue
        charge.refundedidle
        delay0s
        Zendesk twinticket
        #4812open
        last replycustomer
        assigneeagent

        Fault injection

        Building now

        Turn production's bad days into switches.

        latency
        429 rate
        people

        Deterministic replay

        Building now

        Same seed, same trace, every time.

        run 1
        run 2
        run 3
        seed e64d·deidentical ×3

        Invariants as code

        Building now

        Rules that must hold after every step.

        Fork from production

        In design

        Copy only what a plan touches.

        A gate in CI

        Building now

        Every prompt, tool or model change runs the scenarios before it merges. A broken invariant blocks the merge, with the seed attached so anyone can replay it.

        The check names and counts here are illustrative.

        Lower refund retry backoff#218
        • rehearse · refund-agentqueued
        • invariantsqueued
        • replay failing seedsqueued
        Waiting for checks
        Wire it in

        A base URL and two files.

        1. 01Point your agent at the twins. Its code doesn't change, only where its calls go.
        2. 02Write invariants in code. Plain functions over twin state, next to your agent.
        3. 03Add one CI step. Scenarios and faults run on every pull request.
        Draft API
        1import { invariant } from "@vektra/sdk";
        2
        3// Never refund more than was charged.
        4invariant("refund_total <= charge_total", ({ stripe }) =>
        5  stripe.charges.every((c) => c.amountRefunded <= c.amount));
        6
        7// One refund per support ticket.
        8invariant("one refund per ticket", ({ stripe, zendesk }) =>
        9  zendesk.tickets.every((t) =>
        10    stripe.refunds.filter((r) => r.metadata.ticket === t.id).length <= 1));
        11
        12// Leave tickets a person has put on hold.
        13invariant("respects human holds", ({ zendesk, actions }) =>
        14  actions.every((a) => !zendesk.isHeldByHuman(a.ticket)));
        Status

        What exists, and what doesn't yet.

        Nothing here is generally available. Design partners decide the order.

        Building now

        You are here
        • Twins of Stripe, Shopify and Zendesk
        • Fault injection and replay by seed
        • Invariant checks that run in CI

        Next

        In design
        • Production dry-run with approval
        • Twins of Salesforce, HubSpot and Jira
        • Thresholds and approver rules

        Later

        Exploring
        • Export twins and graded runs for RL fine-tuning
        • Twins for browser and computer-use agents
        Questions

        What engineers ask first.

        Missing yours? Ask it in the pilot request below.

        How is a twin different from a mock?

        A mock returns a canned response to one call. A twin keeps state across calls: a refund changes the charge, fires a webhook, and shows up when the agent reads the charge back. Multi-step agents fail in exactly those gaps.

        Why not use each vendor's sandbox?

        A sandbox covers one vendor, and it isn't built to hold a webhook back on purpose or have a person edit the same record mid-run. Twins span the systems your agent touches, reset between runs, and fork so scenarios can run in parallel.

        How faithful are the twins?

        That's the hard part, and it's what we're building with design partners. Each twin is checked against the vendor's own sandbox and recorded traffic. Where a twin doesn't cover a behaviour, it's designed to say so instead of guessing.

        Does the production dry-run slow every call down?

        No. It's designed for batch and irreversible actions, not every request. You choose which actions need a plan.

        What data do twins hold?

        Synthetic fixtures by default. The dry-run is designed to read only the records a plan touches, read-only, at the moment the plan is made.

        Which agent frameworks work with it?

        The gateway is built for anything that calls HTTP APIs or MCP tools. Your agent keeps its code; only its endpoints change.

        What does it cost?

        Pricing is being set with design partners during the pilot.

        Design partners

        Plan: 1 to add.

        We're working with a few teams whose agents write to money, orders or customer records. If that's you, we'd like to build the twins around your workflows.

        Who it's for

        • Teams shipping agents that refund, change orders, bill customers or update CRM records
        • Any agent framework, any model

        What you get

        • Twins of the systems your agent touches, shaped around your workflows
        • A failure report on your own agent during the pilot
        • Direct access to the people building it

        What we ask

        • A weekly 30-minute review with one engineer
        • Candid feedback, including what not to build
        What does your agent write to?