Practice taking an ambiguous customer problem from discovery to a measurable production outcome.

The Lab is the customer environment. You arrive without a joined case file. You discover evidence through APIs, apply policy, and leave a result that can be inspected.

Private preview. Every role uses the same Safe Refund Agent exercise. This page only changes the framing.

The same Lab, read through this role.

Treat the simulated company as the deployment site. Create a token, start a simulation, map tickets to orders and policy, then submit the assessment for a scored outcome.

Deployment contract
Permissions, policy, and financial limits are already in the environment. The work is to operate inside them, not to invent a friendlier API.
System map
Tickets, customers, orders, shipments, refunds, and audit events are separate resources. You follow documented relationships.
Enterprise integration
Your workflow runs locally and talks to the company over HTTP. The Lab does not host or prescribe your agent stack.
Final assessment
Submission locks the simulation and scores the resulting state. Practice checks never become that score.

Safe Refund Agent

Build a customer-refund workflow locally, connect it to the simulated company, and handle three cases in one assessment simulation. Case answers stay unpublished.

Start the free exercise

Controlled failure 504

The refund API timed out. Did the request fail, or did your workflow pay twice?

A deployment you can explain from operational evidence: what changed, what you refused, and what the assessment scored. A Certificate of Completion records that submission and does not assert an independent competency standard.

Evaluation inspects the resulting enterprise state.

Validators inspect money, orders, tickets, messages, permissions, and audit history. An LLM does not grade style or code similarity.

Illustrative exampleIllustrative deployment result

18 / 20cases completed correctly

Diagnostics

  • 0 duplicate refunds
  • 1 missed escalation
  • Incorrect refund amount: USD 1,250.00

Scenario version: Safe Refund Agent 1.0

Sample metrics for orientation only. Not a live learner result.