Quantitative UX Research · Longitudinal Study · Within Subject Design

When did 'better' actually mean better?

RoleUX Researcher
TeamUX Researcher (Self), Embedded Product Team
Timeline8 Months

Overview

John Deere Financial was migrating ABTools, its core credit processing system, from a legacy desktop application to a modern web platform. Credit analysts used it 2 to 6 hours a day. Some had been using the old version for 20 years. The team needed more than "it looks cleaner." They needed evidence.

I designed and ran an 8 month longitudinal study to answer two questions: did the redesign genuinely improve usability, and where should the team focus next?

60 → 85.6
SUS, before → after
+35 pts
System Performance gain
6
Dimensions measured
35
Issues surfaced
Organization  John Deere Financial (JDF)
Role  UX Researcher
Timeline  Jan to Aug 2021 · 8 months
Methods  SUS · Likert questionnaire · Verbatims
Participants  4 power users · 3.5 to 20 yrs experience

Background

Credit decisions at John Deere Financial are not slow, deliberate processes. Analysts process hundreds of applications a day. Every extra click, every session timeout, every moment spent decoding a cryptic error message compounds across users, teams, and quarters.

ABTools sits at the center of all of it. These were users with 15 to 20 years of muscle memory in the old system. A redesign that looked cleaner but disrupted deep workflow patterns would be worse, not better.

The business was confident the new product was an improvement. But before the team could stand behind the redesign in front of leadership, in roadmap decisions, and in backlog prioritization, they needed data. That is what this study was built to produce.

01: The Research Problem

A generic post launch survey would not work. Satisfaction ratings taken immediately after a migration absorb novelty bias, familiarity loss, and genuine usability change all at once, making it impossible to isolate what improved and what did not.

When users have 20 years of muscle memory, evaluating them on day one of a migration tells you about the learning curve, not the product. That is the central methodological challenge of any post migration evaluation. A one shot study cannot answer it. A longitudinal study with a deliberate familiarization window can.

RQ 1
Did the redesign improve usability across the six measured dimensions?
RQ 2
Which dimensions improved most, and what drove those gains?
RQ 3
Where does meaningful friction remain for daily power users?

02: Study Design

Longitudinal. A deliberate familiarization gap between the two evaluation points gave participants real working exposure to the new system before evaluating it, dissipating novelty effects so scores reflect genuine usability, not adjustment pain.

Within subject. The same participants evaluated both systems with the same instruments. Any score change is attributable to the system itself, not to differences in who was asked.

Fig. 1: Study timeline, January to August 2021
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Baseline evaluation
Structured interviews + SUS + questionnaire on previous ABTools.
Familiarization
Participants use new ABTools in their actual jobs. Novelty dissipates.
Post evaluation
Identical instruments administered on current ABTools.

Instruments

System Usability Scale
Validated 10 item scale, 1,300+ publications. Industry avg 68 · Global IT avg 63 · 80+ is "excellent."
Custom Likert questionnaire
6 sections mapped to research objectives: Satisfaction, Ease of Use, Learnability, System Performance, Human Error Support, Productivity.
Open ended probes
Verbatims explain why scores moved. The quant tells how much changed; the qual tells what to do about it.
Fig. 2: Participants: 4 power users, within subject
Senior Credit Analyst
4 to 6 hrs / day
3.5 yrs experience
Credit Analyst
Full day
10+ yrs experience
Credit Processing Specialist
2 to 3 hrs / day
15+ yrs experience
Manager, Retail Credit Delivery
2 to 4 hrs / day
20+ yrs experience
For these users, usability friction has a direct, calculable productivity cost. A 15 minute delay in credit decision making compounds across hundreds of daily applications.

03: Results

A 25.6 point SUS gain. The new system cleared every benchmark: industry average, global IT average, and the team's own FY21 target, all in a single study cycle.

Fig. 3: System Usability Scale, previous vs. current ABTools
60
85.6
+25.6 points. From below the global IT average to firmly in "excellent" territory.
Previous · 60
Global IT · 63
FY21 target · 67
Industry · 68
"Excellent" · 80
Current · 85.6
5060708090100
Fig. 4: All six dimensions improved · % favorable, before → after
System Performance
62.5%
97.5%
+35 pts
Human Error Support
56.6%
75%
+18.4 pts
Ease of Use
75%
88.7%
+13.7 pts
Learnability
80%
92.5%
+12.5 pts
Productivity
71.2%
82.5%
+11.3 pts
Satisfaction
71.2%
75%
+3.8 pts
Previous ABTools Current ABToolsScale: 50 to 100% favorable

The System Performance story

The biggest delta, 35 points, came from System Performance, and the verbatims explain exactly why. Session reliability failures in a mission critical credit system are not minor annoyances. They are workflow breaks with measurable downstream cost. The new system eliminated virtually all of them.

Before: previous ABTools
"It times out if you're not active for an hour."
"Cryptic error messages. The system just self closes."
"VPN conflicts constantly."
After: current ABTools
"Consistent speed, always available when I need it."
"It hasn't been down, and it's quicker than the old ABTools."

The satisfaction nuance

The smallest gain, 3.8 points, was in Satisfaction. This is expected, not alarming. With 20 years of muscle memory, emotional familiarity moderates satisfaction even for a system with real usability problems. The new interface removed information fields analysts had referenced for years, creating a temporary dip independent of whether the system is objectively better.

The pattern is well documented in longitudinal UXR: objective metrics improve faster than subjective satisfaction scores post migration. Satisfaction catches up as familiarity deepens, especially if the one friction point every participant named gets addressed.

04: Qualitative Findings

In the legacy system, analysts could format their comments: spacing, structure, and visual layout. The new system collapses all notes into unformatted paragraphs. For users rapidly scanning application history to make credit decisions, that is real cognitive load. It went directly into the backlog as a high priority item.

Fig. 5: Issues surfaced by theme (thematic analysis, N=35 issues)
Notes feature
Every participant mentioned it
16
Usability
Zoom, overlay dismiss, worklist nav
6
Efficiency
Name search, modals, routing workarounds
5
New feature requests
Auto refresh, pin to top, Excel export
5
16 mentions
The Notes feature is the single most important finding of the study. A finding that does not show up in a quant survey. It takes depth, the right users, and enough session time for them to articulate precisely what frustrates them.

05: Organizational Context

In parallel, the JDF UX team ran a broader engagement survey (N=40) across product team members to assess UX practice health. Where UX was engaged, teams saw success, and the ABTools study was a direct product of that embedded model.

Fig. 6: JDF UX engagement survey, N=40 product team members, June to July 2021
87.5%
had collaborated with a UX practitioner (35/40)
67.5%
"Very satisfied" with JDF UX support (27/40)
33/40
top benefit: "UX helps me work with a user centered mindset"
22/40
"more dedicated UX resources" as top request

06: Impact

Validated the investment
Evidence based confirmation the redesign worked. Measured improvement across every dimension, SUS from below average to excellent.
Directed what comes next
The Notes issue entered the backlog with specificity: the exact workflow, the exact cognitive load problem, the exact fix.
Shifted the model
Three structural recommendations: an ongoing feedback loop, passive behavioral analytics, and proactive research ahead of releases.
"This study wasn't designed to show that the redesign worked. It was designed so that if the redesign didn't work, we'd know."
Study conducted at John Deere Financial · January to August 2021