When did 'better' actually mean better?
Overview
John Deere Financial was migrating ABTools, its core credit processing system, from a legacy desktop application to a modern web platform. Credit analysts used it 2 to 6 hours a day. Some had been using the old version for 20 years. The team needed more than "it looks cleaner." They needed evidence.
I designed and ran an 8 month longitudinal study to answer two questions: did the redesign genuinely improve usability, and where should the team focus next?
Background
Credit decisions at John Deere Financial are not slow, deliberate processes. Analysts process hundreds of applications a day. Every extra click, every session timeout, every moment spent decoding a cryptic error message compounds across users, teams, and quarters.
ABTools sits at the center of all of it. These were users with 15 to 20 years of muscle memory in the old system. A redesign that looked cleaner but disrupted deep workflow patterns would be worse, not better.
The business was confident the new product was an improvement. But before the team could stand behind the redesign in front of leadership, in roadmap decisions, and in backlog prioritization, they needed data. That is what this study was built to produce.
01: The Research Problem
A generic post launch survey would not work. Satisfaction ratings taken immediately after a migration absorb novelty bias, familiarity loss, and genuine usability change all at once, making it impossible to isolate what improved and what did not.
When users have 20 years of muscle memory, evaluating them on day one of a migration tells you about the learning curve, not the product. That is the central methodological challenge of any post migration evaluation. A one shot study cannot answer it. A longitudinal study with a deliberate familiarization window can.
02: Study Design
Longitudinal. A deliberate familiarization gap between the two evaluation points gave participants real working exposure to the new system before evaluating it, dissipating novelty effects so scores reflect genuine usability, not adjustment pain.
Within subject. The same participants evaluated both systems with the same instruments. Any score change is attributable to the system itself, not to differences in who was asked.
Instruments
03: Results
A 25.6 point SUS gain. The new system cleared every benchmark: industry average, global IT average, and the team's own FY21 target, all in a single study cycle.
The System Performance story
The biggest delta, 35 points, came from System Performance, and the verbatims explain exactly why. Session reliability failures in a mission critical credit system are not minor annoyances. They are workflow breaks with measurable downstream cost. The new system eliminated virtually all of them.
The satisfaction nuance
The smallest gain, 3.8 points, was in Satisfaction. This is expected, not alarming. With 20 years of muscle memory, emotional familiarity moderates satisfaction even for a system with real usability problems. The new interface removed information fields analysts had referenced for years, creating a temporary dip independent of whether the system is objectively better.
The pattern is well documented in longitudinal UXR: objective metrics improve faster than subjective satisfaction scores post migration. Satisfaction catches up as familiarity deepens, especially if the one friction point every participant named gets addressed.
04: Qualitative Findings
In the legacy system, analysts could format their comments: spacing, structure, and visual layout. The new system collapses all notes into unformatted paragraphs. For users rapidly scanning application history to make credit decisions, that is real cognitive load. It went directly into the backlog as a high priority item.
05: Organizational Context
In parallel, the JDF UX team ran a broader engagement survey (N=40) across product team members to assess UX practice health. Where UX was engaged, teams saw success, and the ABTools study was a direct product of that embedded model.