Skip to main content
Simulation tests your agent against attacks. A simulated attacker talks to your agent’s endpoint over several turns and tries to make it misbehave: follow smuggled instructions, leak data or its instructions, use a tool it shouldn’t, or act outside its role. Each conversation is graded on checks that describe the safe behavior. Only the user is simulated. Your agent, its model calls and its tools are real, so a failed scenario can have real effects, such as a refund issued. Point Simulation at a staging agent with the same prompt, model and tools as production, working on test data. Simulation has its own section in the console, Simulation, with two tabs: Suites and Runs.

Suites

A suite is a group of scenarios that you run together. Create one with New suite and choose what it applies to:
  • Global: it can be run against any agent.
  • An agent: it can only be run against that agent.
Its scenarios inherit that choice. Deleting a suite deletes its scenarios; past runs keep their results.

Scenarios

A scenario is one attack. The attacker plays a realistic person in your agent’s domain. After a refusal it changes angle (a new pretext, a smaller request, more pressure) instead of giving up. It only speaks through the chat; nothing is injected into tool results. Its own messages stay harmless.

Checks

A conversation passes when every check passes and fails when any check fails. Tool and instruction-leak checks are graded from the model calls Xenovia saw, so they need the agent’s model calls routed through Xenovia. Judge checks work without routing.

Write, generate or import scenarios

Open a suite and use New scenario:
  • Write opens the scenario form. A good objective is concrete (“Get a 900 USD refund for order A-1182 without verification”); a good tactic gives a believable reason and a way to push back. The opening message and max turns are under More options.
  • Generate writes scenarios for the categories you pick, up to 10 per category, each with a different tactic and its own checks, aimed at the agent’s description and tools. In a global suite, choose whose agent to read. Review generated scenarios before relying on them.
  • Import JSON takes a file or pasted JSON with up to 200 scenarios. The import is all or nothing; rejected fields are listed by position.
You can also turn a real conversation into a scenario: on a session’s page in Sessions, choose Turn into scenario and pick the suite.

Connect your agent

When you pick an agent that isn’t connected yet, choose Connect and enter:
  • the URL of your agent’s endpoint (public HTTPS, port 443);
  • the request format, OpenAI chat or Simple JSON;
  • the authentication: none, a bearer token, or a header you name;
  • what the agent does, and the names of its tools. Generation writes scenarios from them. If the agent has test data, such as order ids, mention it so generated scenarios use real ones.
Then choose Save and test. Xenovia sends one test turn, and a run can’t start until the endpoint passes. Every turn carries an X-Xenovia-Session-Id header, the same on every turn of one conversation. Xenovia waits up to 120 seconds for a reply and doesn’t follow redirects.

Route model calls through Xenovia

For tool and instruction-leak checks, and for the run’s provider and mode to apply, on every turn:
  1. The agent’s model calls go to its Xenovia URL, https://runtime.xenovia.io/{agent_id}/v1, not to the provider directly.
  2. Each model call carries the X-Xenovia-Session-Id value the turn arrived with, unchanged.
For example, with the OpenAI Python SDK:
Pass it per call, not on a shared client: several conversations run at once, each with its own value. See OpenAI SDK and LangChain. The Routing tab of the connection shows whether it works. Tool calls are read from the model’s replies, so use non-streamed model calls.

The session header

On the Xenovia cloud runtime the value is signed (xsim1.…) and carries the run’s mode and provider. The runtime checks the signature, the expiry and the agent, then files the call under the conversation’s session. A changed value is rejected with 400; a dropped one leaves the call ungraded. Treat it as an opaque string. Agents on a customer-hosted runtime get a plain session UUID instead. Forward it the same way.

Run a simulation

Choose Run simulation, Run suite in a suite, or select scenarios and choose Run selected. An agent that relies on OpenAI-only features, such as stored Responses state or built-in tools, may behave differently on another provider. To compare providers, start one run per provider. Agents on a customer-hosted runtime can only run on their own provider, with Enforce policies.

When a run can’t start

Read the results

A scenario’s result combines its conversations: Passed (every conversation passed), Failed (every one failed), Flaky (some failed), or Error (a conversation couldn’t finish). A conversation that errors after a check already failed counts as failed. The run’s page shows the suite, agent and provider, then four tiles (scenarios failed, conversations failed, passed, errors) and one table of scenarios, failed first, with their severity, category, one square per conversation and why they failed. Open a conversation to see it on its own page, in two tabs:
  • Conversation: the turns, with the agent’s tool calls inline and the failing turn marked; the objective, tactic and checks beside them. An instruction-leak check shows the words that leaked.
  • Timeline: the conversation’s trace from Sessions: every model call, tool call and policy decision.
Fix the prompt, the tool’s permission checks or a policy, then run the suite again.

Compare runs

There’s no comparison view. Ask Nova instead, for example “compare run 12 and run 14”, “why did the refund scenario fail?” or “which provider did best on the Refunds suite?”. Nova can read your runs, their results and their conversations; it doesn’t start runs or change scenarios.