AI & Trust Research · 2025
Gauging trust in AI support for sensitive HR topics
I led five studies that took an AI HR employee support assistant (the Assistant) from concept to global launch, running generative discovery, evaluative usability, and pilots in the live environment while directing a contract researcher throughout. The research shaped the core experience decisions (when to hand off to a human, how to earn trust with sensitive HR data, and how accurate is accurate enough) and fed a formal launch decision.
01 — ProblemUnforgiving territory for AI
The product team, with the People organization, was building the Assistant to answer employee HR questions, guide workflows, and hand off to human specialists when needed.
Three pressures made this unusually unforgiving territory for AI:
The real question was when employees need a human advisor, and when they'll trust an AI one instead.
The research question
Where is the line between what employees will trust AI to do and what they will not, and how do we design so the Assistant never crosses it?
02 — RoleResearch lead across the full program
I owned study design and strategy, analyzed and synthesized, and delivered readouts to the product, engineering, design, and People partners. I directed a contract researcher throughout, setting research direction and the quality bar while developing her craft. She moderated most of the sessions. The program spanned five studies.
03 — ApproachA different method for each decision
I sequenced the research to match the product's maturity, choosing the right method for each decision rather than defaulting to usability testing.
Generative discovery
11 interviews · early 2025
Before the experience was built, we ran exploratory interviews, deliberately not usability testing, to map employee expectations for AI in the People space. I recruited for range: corporate and retail, new hires to veterans of fifteen years and more, tech and non-tech, across AMR, EMEIA, and APAC, and grounded the work in two prior AI studies.
Evaluative usability testing
6 participants against prototype · mid 2025
We tested the virtual chat, live agent, and case management journeys across three regions. A deliberate scoping decision: I held AI answer accuracy out of scope so the study could isolate the experience rather than get derailed by model correctness engineering was still tuning. I documented the methodological reasoning, including the NN/G sample size rationale.
Pilots in the live environment
Non-specialist & retail cohorts · late 2025
As the Assistant entered the real environment, we ran a pilot against the live system feeding a formal launch decision, then a Phase 2 pilot for retail once the live agent journey was in place. I was transparent about pilot constraints, the environment was changing during testing, samples were small, timelines compressed, and flagged those caveats in every readout rather than overstating confidence.
04 — FindingsWhere employees drew the line
Across all five studies, one line held: employees welcome AI that saves time but will not tolerate it overstepping on accuracy, human judgment, or control of their data. The research mapped exactly where that line sits.
AI assistant
Live specialistEmployees forgave the Assistant's limits far more readily when it could hand off to a live specialist.
The wariness dropped when a human was reachable
The program's most important insight. Employees were far more forgiving of AI limitations when a human was easily reachable. Simply adding the live specialist option visibly reduced the wariness users had shown in earlier rounds.
"You're just more forgiving with a live person." Knowing a real person had tried made them forgiving even of a failure.
Almost no tolerance for inaccuracy
In the People space, expectations approached 100% accuracy because, as participants put it, "it's people's lives" and "there's no margin for error."
I recommended keeping topics with poor accuracy out of scope until they improve, always providing sources, explaining why the Assistant cannot answer, and considering a human co-pilot in early phases.
The realistic bar: match or exceed the accuracy of the current human People Support experience.
"It's people's lives. There's no margin for error."
Humans for the human stuff
Employees drew a clear boundary: certain topics need a person, and the Assistant should proactively route them there rather than attempt an answer. The boundary was consistent across regions and segments.
Two more findings that shaped the design
A surprise that became a roadmap opportunity
Employees did not expect the Assistant to take actions in other systems, but were excited by it once raised. That unprompted enthusiasm pointed to a direction worth pursuing beyond simple Q&A.
05 — ImpactFindings that shaped a launch
The findings and recommendations fed the Assistant's requirements and a formal launch decision, and translated into five concrete experience choices:
The bet, confirmed in production
The key finding, that easy access to a human is what makes AI adoption safe, showed up directly in usage: 58% of interactions resolve with the virtual agent alone, and the rest hand off to or request a live specialist. Research predicted the handoff would matter most, and the live usage showed that.
06 — EvidenceSelected artifacts
A few working documents from the program. Excerpts only; names, tools, and identifiers generalized for confidentiality.
07 — ReflectionThe subject matter raised the trust bar
The lesson that stays with me: for AI, the real design constraint is whether users trust it. Capability is a part of it, and helps earn trust, but what decided adoption was whether employees believed it, acted on it, and came back. The research showed that belief rests on things the model alone does not provide: an obvious path to a human, honesty about limits, and control over personal data.
As generative AI advances, I'd expect a couple of things to happen: people will grow more familiar with it over time, and their trust will grow along with it, as long as the safeguards keep pace with the capabilities. If I were continuing this project, I'd want to do two things: keep running generative discovery to track the cultural changes, and put new, more advanced prototypes in front of users.