AI & Trust Research · 2025

Gauging trust in AI support for sensitive HR topics

TimelineFive studies · early–late 2025 RoleResearch lead · directed 1 contract researcher MethodsGenerative discovery · evaluative usability · live pilots
Snapshot

I led the five-study research program that took an AI HR employee-support assistant (the Assistant) from concept to global launch, running generative discovery, evaluative usability, and live-environment pilots while directing a contract researcher throughout. The research shaped the core experience decisions (when to hand off to a human, how to earn trust with sensitive HR data, and how accurate is accurate enough) and fed a formal Go/No-Go.

~80k
Interactions · 7 countries
88%
Positive/neutral vs 85% goal
5
Studies, concept → launch

01 — ProblemAbout as high-stakes as AI gets

The product team, with the People organization, was building the Assistant to answer employee HR questions, guide workflows, and hand off to human specialists when needed.

Three pressures made this unusually unforgiving territory for AI:

Sensitive personal data sat across multiple HR and case-management systems.
Wrong answers carried real consequences: people's pay, leave, and livelihoods.
Employees' trust in their own employer was on the line with every interaction.

The real question was when employees need a human advisor, and when they'll trust an AI one instead.

The research question

Where is the line between what employees will trust AI to do and what they will not, and how do we design so the Assistant never crosses it?

02 — RoleResearch lead across the full program

I owned study design and strategy, moderated sessions, analyzed and synthesized, and delivered readouts to the product, engineering, design, and People partners. I directed a contract researcher throughout, setting research direction and the quality bar while developing her craft. The program spanned five studies.

03 — ApproachA program, sequenced to the product's maturity

I sequenced the research to match the product's maturity, choosing the right method for each decision rather than defaulting to usability testing.

1

Generative discovery

11 in-depth interviews · early 2025

Before the experience was built, I ran exploratory interviews, deliberately not usability testing, to map employee expectations for AI in the People space. I recruited for range: corporate and retail, new hires to 15-plus-year veterans, tech and non-tech, across AMR, EMEIA, and APAC, and grounded the work in two prior AI studies.

2

Evaluative usability testing

6 participants against prototype · mid 2025

I tested the virtual-chat, live-agent, and case-management journeys across three regions. A deliberate scoping decision: I held AI answer accuracy out of scope so the study could isolate the experience rather than get derailed by model correctness engineering was still tuning. I documented the methodological reasoning, including the NN/G sample-size rationale.

3

Live-environment pilots

Non-specialist & retail cohorts · late 2025

As the Assistant entered the real environment, I ran a pilot against the live system feeding a formal Go/No-Go, then a retail-specific Phase 2 pilot once the live-agent journey was in place. I was transparent about pilot constraints, the environment was changing during testing, samples were small, timelines compressed, and flagged those caveats in every readout rather than overstating confidence.

04 — FindingsWhat the research revealed

Across all five studies, one line held: employees welcome AI that saves time but will not tolerate it overstepping on accuracy, human judgment, or control of their data. The research mapped exactly where that line sits.

AI assistantAI assistant
Handoff
Live specialistLive specialist

Employees forgave the Assistant's limits far more readily when it could hand off to a live specialist.

The human handoff is the foundation of trust

The program's most important insight. Employees were far more forgiving of AI limitations when a human was easily reachable. Simply adding the live-specialist option visibly reduced the wariness users had shown in earlier rounds.

"You're just more forgiving with a live person." Knowing a real person had tried made them forgiving even of a failure.

Near-zero tolerance for inaccuracy

In the People space, expectations approached 100% accuracy because, as participants put it, "it's people's lives" and "there's no margin for error."

I recommended keeping low-accuracy topics out of scope until they improve, always providing sources, explaining why the Assistant cannot answer, and considering a human co-pilot in early phases.

The realistic bar: match or exceed the accuracy of the current human People Support experience.

~100%
Expected accuracy
0%100%

"It's people's lives. There's no margin for error."

Humans for the human stuff

Employees drew a clear boundary: certain topics need a person, and the Assistant should proactively route them there rather than attempt an answer. The boundary was consistent across regions and segments.

Two more findings that shaped the design

Trust is earned early and easily lost. Employees wanted transparency about what data is accessed and stored, distinguished access from storage, and wanted to opt in by data type.
Segment and regional nuance. Corporate users preferred a professional tone; retail wanted warmth, needed the chat never to time out on shared floor devices, and valued the live agent even more. One size would not fit all.

A surprise that became a roadmap opportunity

Employees did not expect the Assistant to take actions in other systems, but were excited by it once raised. That unprompted enthusiasm pointed to a high-value direction beyond simple Q&A.

05 — ImpactFindings that shaped a launch

The findings and recommendations fed the Assistant's requirements and a formal Go/No-Go decision, and translated into five concrete experience choices:

A live specialist reachable at any point in the journey.
A professional tone with guardrails on sensitive topics.
Transparency about data handling: what is accessed, and what is stored.
Sources on every answer, so employees can verify what the Assistant tells them.
Honest scoping of what the AI would and would not attempt.
~80k
Chat interactions across 7 countries
88%
Positive-or-neutral sentiment vs an 85% goal
−2.4hrs
Faster per case · days-to-close 0.8 → 0.7

The bet, confirmed in production

The key finding, that easy access to a human is what makes AI adoption safe, showed up directly in usage: a substantial share of interactions hand off to or request a live specialist rather than resolving with the virtual agent alone. Research predicted the handoff would be load-bearing, and the live usage bore that out.

06 — EvidenceSelected artifacts

A few working documents from the program. Excerpts only; names, tools, and identifiers generalized for confidentiality.

07 — ReflectionThe subject matter raised the trust bar

The lesson that stays with me: for AI, the real design constraint is whether users trust it. Capability is a part of it, and helps earn trust, but what decided adoption was whether employees believed it, acted on it, and came back. The research showed that belief rests on things the model alone does not provide: an obvious path to a human, honesty about limits, and control over personal data.

As generative AI advances, I'd expect a couple of things to happen: people will grow more familiar with it over time, and their trust will grow along with it, as long as the safeguards keep pace with the capabilities. If I were continuing this project, I'd want to do two things: keep running generative discovery to track the cultural changes, and put new, more advanced prototypes in front of users.

Want to talk through this work?
Email Résumé (PDF)

Daniel Farooqi · Staff UX Researcher · Product and team names generalized for confidentiality.

{{ modalTitle }}{{ modalSub }}
{{ modalNote }}
{{ pg.alt }}