Skip to content
Jessica Fleur

Case study · Ultimate Sackboy

Running Player Usability Tests

Ran playtests with ten players and found the score and multiplier confused them.

The game

Ultimate Sackboy is an endless runner built by Exient around Sackboy, PlayStation’s knitted hero, on iOS and Android.

Role
UX Designer
Date
November 2020
Skills
  • Usability Testing
  • Survey Design
  • Qualitative Research
  • Observational Recording
  • Qualitative Documentation
  • Feedback Resolution
Tools
  • PlaytestCloud
  • Microsoft Excel
  • Microsoft PowerPoint
  • Confluence

The challenge

Ultimate Sackboy was heading for launch and we still hadn’t checked whether players understood how it worked. I ran a round of playtests with ten of them to find out.

01

Setting Up the Test

I recruited ten testers against the personas we’d already built for the game.

The brief the ten testers were recruited against

The sessions ran on PlaytestCloud, which records the screen and the player’s commentary, sends back a transcript, and shows where on the screen they tapped.

Each session was followed by a survey, capped at ten questions by the platform. I used them on the things watching wouldn’t tell me: what players thought the score meant, what they thought the bubbles were for, whether the UI made sense, and what they made of the art.

02

Watching It Back

I built a profile for each tester first, so their answers came with some context: age, device, country, how much they played and what they usually played.

The ten tester profiles
How familiar testers were with the franchise

Then I read through the survey results and watched the recordings back, timestamping anything the survey hadn’t picked up. Each session ran fifteen minutes.

Survey results
A session recording with my timestamped notes

I coded all of it into one sheet by category: usability, mechanics comprehension, playability, bugs, and anything they volunteered as good or bad.

Observations coded by category
03

What Players Didn’t Understand

The main finding was that the numbers on screen weren’t reaching anyone. Ten testers gave four different answers for what the score represented, four for what the multiplier did, and three for what the coloured segments on the progress bar meant. That readout is the game’s main feedback on how a run is going.

What the score meant: four answers from ten players
What the multiplier did: also four
The progress bar’s colour segments: three

The confusion went wider than the numbers. Testers split four ways on what they were trying to achieve in a run, and nobody could say why they were collecting bubbles or how bubbles affected their score.

What they thought they were trying to do
Bubbles: what players made of them

For each issue I wrote up what players said, what I’d watched them do, and what I thought would fix it: animating the intro scene to show the objective, rewording the lose screen, moving the score next to the multiplier icon, and separating the distance indicator from the score comparison on the progress bar.

Score: observations and proposed fixes
Multiplier: observations and proposed fixes
Progress bar: observations and proposed fixes
04

The Report

Not all of it was negative. The art style rated 8.6 out of 10 and testers praised it without being prompted. I also had them rate how fun the game was, whether they’d play it again and whether they’d recommend it.

Art style: 8.6 out of 10
Fun, replay and recommend

Separately from the comprehension problems, I logged the playability issues by how often they came up. The two most frequent were deaths the player couldn’t avoid and target scores they couldn’t reach. I had screenshots for one of them: the player moves left, and there isn’t time to reach the right-hand platform before the train arrives.

Blockers, grouped by where they came up
Playability issues, ranked by how often they came up
The player moves left, and there isn’t time to reach the right-hand platform before the train.

All of it went into a report with the proposed fixes, which I presented to the whole company, C-level included. Several of the fixes were built. We ran the same process at every stage after that, through the rest of development and up to launch.