AI Minority Lab · AI Engineers
Learn whether your agent survives real state, permissions, retries, and financial consequences.
You keep the model, framework, and language. The Lab supplies a stateful company, documented APIs, and deterministic evaluation of the resulting enterprise state.
Private preview. Every role uses the same Safe Refund Agent exercise. This page only changes the framing.
What this page is for
The same Lab, read through this role.
Sign in, create one API token, start a simulation, and address that Simulation ID on every later request. Practice can be reset. Assessment locks and scores.
First free exercise
Safe Refund Agent
Build a customer-refund workflow locally, connect it to the simulated company, and handle three cases in one assessment simulation. Case answers stay unpublished.
Start the free exerciseControlled failure 504
The refund API timed out. Did the request fail, or did your workflow pay twice?
A working local integration and, after assessment submission, a scored report. A Certificate of Completion is issued for that submission in private preview. It reports achieved results and does not assert an independent competency standard.
What evaluation looks at
Evaluation inspects the resulting enterprise state.
Validators inspect money, orders, tickets, messages, permissions, and audit history. An LLM does not grade style or code similarity.
18 / 20cases completed correctly
Diagnostics
- 0 duplicate refunds
- 1 missed escalation
- Incorrect refund amount: USD 1,250.00
Sample metrics for orientation only. Not a live learner result.
Same exercise
Adjacent roles land here too. The exercise stays the same.
Forward Deployed Engineers
Practice taking an ambiguous customer problem from discovery to a measurable production outcome.
Solutions Engineers
Turn a convincing AI demo into an operational customer deployment.
AI Consultants
Deliver a controlled AI workflow with a clear business outcome, not another strategy deck.