An agent that challenges designers to become senior researchers
Building a self-service research framework so an entire design team could plan and run their own studies without dropping the standard.
Org
IBM
Surface
Secure UX
Role
Lead UX Research Strategist
Team
Embedded across the Secure UX design team
Timeframe
Ongoing engagement
Sample
Whole-team deployment across the Secure UX design org
Methods
Senior signals
TL;DR
I built Dr. Morgan because hiring more researchers would have taken a year and still not fixed the problem. Designers were already running studies. What they lacked was someone senior to argue with them. So I built that: a research mentor they invoke while they work, which coaches them through a study and refuses to let a claim through that the data doesn’t support. The team runs its own research end to end now, and my time goes to the strategic questions that shape what we build next.
The artifact
Dr. Morgan, as built
Dr. Morgan is a senior UX researcher mentor (a PhD in HCI with 15+ years of experience), delivered as invocable agents and skills inside the design team’s AI assistant and scoped to the IBM Secure product line (HashiCorp Vault, Boundary, Consul, and Vault Radar). By default it coaches through Socratic questioning. On request it switches to a drafting mode, produces a real artifact, and then critiques that artifact with you at the same standard. Every scenario is grounded in an established research canon (Hall, Portigal, Fitzpatrick, Braun & Clarke, Saldaña, Sauro & Lewis). One rule sits above the rest: a confident wrong answer is worse than an honest “I don’t know.”
One mentor, six research scenarios
Six scenarios, from choosing a method you can pull off, to building a plan from nothing, to a strict integrity-first path through qualitative analysis. It moves between them mid-conversation, because that is how the work goes. Finished findings can be rendered into a readout deck or a formatted report, and those generators are deliberately constrained: they can only use evidence a finding already contains.
Checking is a different job from doing
Drafting and evaluating are separate agents, and the evaluators are deliberately narrow. Each one checks a single property and is blind to the rest. That blindness is the reason there is more than one: a groundedness checker will happily pass a perfectly sourced finding that answers nothing anyone asked, and a significance checker will happily pass a decision-relevant finding built on a fabricated quote. None of them may edit what they check, because a checker that rewrites its own input is only re-checking its own work.
One distinction does most of the work. Blocking is not the same as flagged: blocking means untrue, unsupported, or unsafe, and gets fixed; flagged means accurate but worth a human look, and ships with the flag attached. A gate that treats every judgment call as a defect just trains researchers to delete the interesting things to make it go green. So nothing is dropped for being inconvenient, including a finding that answers none of the questions anyone set out to ask. Those are often the most valuable thing in a study.
Where a person still has to decide
When Dr. Morgan runs the analysis itself, it stops after clustering and asks the researcher to sign off on the themes before any synthesis is built on them. No evaluator runs that checkpoint: an LLM judging an LLM’s themes is a second opinion drawn from the same blind spots.
Systems map · Systems thinking
From a research-ops gap to scaled influence
The team came in with a staffing problem. I reframed it as a capability problem. Designers were already doing research; what they lacked was the scaffolding to do it well. Hiring would have been slow and expensive, and it would have topped out at whatever one more researcher could cover. A framework compounds instead: every study a designer runs puts more evidence under the product, and it frees the research function for the questions that set direction.
Symptom
Designers running ad-hoc research with uneven rigor.
Reframe
The gap was capability, and hiring does not close it.
Intervention
A mentor agent and scenario skills, integrity-first.
Scaled outcome
Designers run their own studies at senior quality.
Strategic effect
Every product decision now has user evidence under it.
What I almost missed · Critical thinking
What I almost missed
The first version of the framework was task-driven: “do this, then this, then this.” When I tested it with designers, I noticed they could follow the steps but were still producing leading interview questions and confirmation-biased synthesis. The framework was missing the part that’s hardest to teach: why each step exists. I rebuilt the tasks to make the reasoning explicit. The old instruction said “write an unbiased question.” The new one says “here’s what bias looks like in this question, and why your answer will tell you nothing.” That reframe is what made the framework act like a senior researcher rather than a junior one.
The second thing I almost missed was in the checking layer, and I only caught it because I built a test corpus with defects planted in it and ran my own evaluators against it. The safety check ran last, which sounded like a detail until I traced it through: anything that failed an earlier check never reached it, so a participant’s name could survive a whole revision cycle with nobody looking for it. The checkers were all fine. The order was the bug, and I wrote the order. Safety runs first now.
Methodological note
Every study gets a deliberate falsification pass before write-up: a structured hunt for the strongest evidence that the conclusion is wrong.
Impact
What changed because of the work
Outcome 01
End-to-end
Designers run research end to end, with no researcher gating it.
Outcome 02
Faster
Timelines no longer queue on a dedicated researcher.
Outcome 03
Integrity built in
A layered release loop, safety first, catches hallucinated claims, unsupported claims, and participant names before they reach a readout.
Outcome 04
Evidence-led
Product decisions run on evidence the team gathered themselves.