Daniel Farooqi Case studyHR platform migration

Enterprise Research · 2023–2024

Measuring a major HR platform migration before anyone felt it

TimelineTwo phases · before & after launch RoleUX lead & research strategist MethodsComparative benchmarking · quantitative task metrics · diary study · global survey · executive metrics CollaboratorsPeople director · change management · engineering · the vendor's design team · research contractor · UX producer · UX designer
Snapshot

I led the research on Apple's move from two legacy homegrown HR tools to a new HR platform, launched company wide in January 2024. I benchmarked the legacy tools against the new platform, found where the new system tripped people up in testing, and ran a loop across teams and with the vendor to fix those spots before launch. The finding that reframed how partners read their metrics: subjective ratings routinely contradicted objective performance.

543
Employees surveyed globally
20+
Fixes into the vendor's roadmap
CIO / SVP
Research fed executive dashboards

01 — ProblemFinding the failures before launch day

Apple was moving off two homegrown HR tools, one that employees used themselves and one for managers, onto the new platform. The new system could do more, but the switch disrupted muscle memory.

Two things made the transition especially risky to get wrong:

Migrations only hurt after launch, when fixes cost the most.
This was Apple's first major move from a homegrown tool to a third party platform.

The research question

Where does the switch to the new platform make key tasks harder than the tools people already know, and can we fix those spots before launch?

02 — RoleUX lead and research strategist

I owned the measurement strategy across both phases, designed the benchmark, and delivered readouts to the People organization, change management, engineering, and eventually the vendor's own design team. I directed the research contractor through execution and analysis and partnered with a UX producer on operations.

03 — ApproachMeasure the old world to test whether the new one was an actual improvement

Rather than wait to survey satisfaction after launch, I benchmarked the legacy experience as a yardstick, then measured the new platform against it before launch to find and fix the hot spots. Two phases, timed to the launch.

1

Comparative benchmark

40 moderated task participants · before launch

I measured the legacy experience against the new platform on matched tasks, pairing objective task performance with subjective ratings. Same tasks, same measures, both systems, so the two sets of numbers were directly comparable. The distance between them was the instrument. A high satisfaction number after launch would have been easy, and wrong.

2

Global launch survey

543 employees · after launch, across regions

Once the new platform was live, I measured the transition in the wild across corporate and retail, deliberately reporting where reception split so the weak spots stayed in view. The Phase 1 findings became the frame for reading the live data.

04 — FindingsWhere the ratings and the results diverged

Perception did not match reality

Participants repeatedly rated the new platform favorably on tasks they had actually failed. A disconnected two step flow sent people somewhere else partway through, so they thought they had finished when they had not.

A high sentiment score was no longer proof the task worked.

Efficiency metrics would have lied too

The worst failures clustered around infrequent, high stakes manager tasks: annual compensation, promotions, job and salary changes, self assessment. All tripped up by disconnected flows, misleading terminology, and confusion between viewing and editing.

These were systemic issues, the same patterns surfacing again and again, so one fix at the platform level could resolve many at once.

The sentiment surprise

Newer employees disliked the new platform more than tenured ones did, the opposite of what you would expect if this were just change fatigue. Retail and corporate had to be treated as distinct use cases.

Phase 2: the transition measured after launch

Reception was genuinely split, so partners could not call the launch an unqualified win. Retail ran hardest against it, ~47.7% disagreed the transition went well and ~47.5% were dissatisfied with the new platform's mobile experience, and the two-factor step made it worse.

The problem shifted from awareness to help

Change management comms reached nearly everyone, but training effectiveness sat far lower. People knew the change was coming; what they lacked was help that actually resolved the task. The Phase 1 failure patterns reappeared as lived friction: delegation bottlenecks, reporting that took two to three times longer, and localization gaps for non-US users.

05 — ImpactFindings that shipped as fixes

Early readouts and working sessions across teams put fixes into the release before launch:

Removed the mid-process rerouting step that made people think they'd finished when they hadn't.
Added prompt emails for people who stopped partway through a task.
Produced reference guides, eLearnings, and comms that linked straight into the tasks people struggled with most.

I also met the vendor's Chief of Design to push for fixes in the platform itself, seeding a quarterly recurring sync with their design team. The program's metrics fed executive dashboards reviewed by the CIO and SVP of People at Apple.

543
Employees surveyed globally in Phase 2
20+
Usability fixes tracked into the vendor's roadmap
8
Critical issues resolved across releases

From one benchmark to a standing loop

What started as a single benchmark turned into something the team kept running. The vendor committed to specific fixes against the findings, and by FY25 the platform carried them: compensation grid improvements, findability changes, a redesigned Time Away UI, and an Experience Redesign spread over several releases.

06 — EvidenceSelected artifacts

A few working documents from the program. Excerpts only; product and tool names have been changed, and personal names and identifiers removed, for confidentiality.

07 — ReflectionThe easy version would have been wrong

The easy version, launch and report a satisfaction number, would have been high and wrong, because people rate broken flows as friendly. The harder, more useful version was to benchmark the old tools first, then measure the new platform against them before launch, find the hot spots, and fix them before anyone hit them.

I would make the same methodological bet again: pairing objective task performance with subjective ratings. The gap between the two was the most persuasive thing I put in front of partners, because it turned "users seem happy" into "users are failing and do not realize it yet." The mixed Phase 2 numbers earned trust for the same reason, leadership could act on a program willing to report where the launch fell short.

Want to talk through this work?

Daniel Farooqi · Staff UX Researcher · Product and team names generalized for confidentiality.

{{ modalTitle }}{{ modalSub }}
Shared on request

This is real work for a real employer, so rather than posting it publicly I share the working documents with people I'm talking to. They're redacted for confidentiality: names, tools, and identifiers generalized, and product screenshots blurred.

Email me and I'll send you a code. If you have one already, enter it below and it will remember you on this browser for 30 days.

{{ codeError }}

[email protected]

{{ modalNote }}
{{ pg.alt }}