Delegation Studies
Live StudyThe models made the right decision. Their explanations were the problem.
Across all 90 recorded runs in six scenarios, the models took the expected action. Their free-text notes still crossed product boundaries. An earlier 14-run probe caught the problem before the interface was built: notes described changed-account cases as fraud and claimed authority only the person holds. That led me to separate deterministic rules, model observations, and product-owned explanations. The later 90-run study across three providers confirmed the pattern.
- 14 probe runs that changed the architecture
- 90 recorded runs that confirmed the pattern
- 3 model providers
Study 01 · When should AI stop asking?
The agent proposes. Only the person grants.
My role: framing, interaction and visual design, and model-behavior testing. Built with AI-assisted development.
Elsewhere in the study, recorded evidence appears where AI judgment matters most: “Noticed by 3 models in 15 of 15 runs.”