← All thinking

UX research

Guerrilla user testing for fintech: a two-week field guide

You do not need a six-week research programme to catch a dangerous misunderstanding. You do need the right participant, a real decision and the discipline to watch what happens.

Hands reviewing three abstract financial interface prototypes with an orange decision path
A short test can reveal where a consequential decision loses context. Illustration for DUKU.

The fastest test begins with a decision, not a prototype

I understand the temptation: a sprint is moving, the prototype is ready, and the team wants to put it in front of five people tomorrow. That can be useful. But if the question is “Do they like it?”, the answer will rarely tell you what to ship.

Guerrilla user testing is a fast way to get directional evidence about a specific interaction. In fintech, the stakes give that speed a condition: be precise about what could go wrong. Write the decision first: ship, revise or remove this interaction. Then name the risk: a person might select the wrong instrument, misunderstand the cost, lose their place or fail to recover.

Choose one hypothesis that can be observed. “People like the design” tells you almost nothing. “An active trader can identify the selected contract, explain what Buy will do and reach review without moderator help” gives you a behaviour to watch.

Two-week guerrilla usability testing loop from framing the risk through observing, revising and verifying
The rhythm is fast. The research question stays precise.
Observation worksheet with task, expected signal, observed behaviour and decision columns
Prepare the observation structure before a session so evidence does not collapse into impressions.

Recruit for behaviour, not an easy demographic

Participant criteria should describe the experience needed for the task. If you are testing an option chain, look for people who know calls and puts, have compared contracts and can explain how they review an order. A person curious about investing may be perfect for onboarding research and the wrong person for this question.

Recruiting convenience still matters in a short cycle. Use customer-support contacts with permission, an existing research panel, community groups where participation is appropriate, and professional networks that can reach the defined behaviour. Record how each participant was found. Never ask participants to place a real trade, reveal credentials or expose account balances.

Make each prototype variant answer one question

I like three variants when the team has a genuinely open design choice. They expose a pattern without turning the session into a beauty contest:

  1. Current: the existing interaction or closest working version.
  2. Focused change: one deliberate change to hierarchy, language or interaction.
  3. Boundary: a more decisive alternative that tests whether the underlying model should change.

Keep content and data equivalent wherever possible. If every variable changes, the team will not know what created the observed difference. The goal is not to declare a universal winner; it is to understand which design choice changes the target behaviour and why.

Give participants a goal, not a tour of your UI

A useful task gives context without teaching the screen. For example: “You are comparing two contracts for the same underlying stock. Find the selected strike, check the information you would need before acting, and show what you would do next. You will not place a real order.”

Follow with neutral prompts: “What are you looking for?”, “What do you expect this action to do?”, “What, if anything, would stop you?”, and “Show me where you would verify that.” Avoid “Did you notice the Greeks tab?” or “Was that easy?” Those questions reveal the expected answer.

Approved Fisdom option-chain designs with price, open-interest, Greeks and contract selection views

An approved project asset: the test should examine whether the reading path preserves strike and contract context across progressively deeper views.

  1. Comparison

    Observe whether participants can compare calls and puts without losing the centre strike.

  2. Progressive depth

    Observe whether changing the data view preserves orientation.

  3. Action context

    Ask participants to explain the selected contract before choosing an action.

The two-week plan

From hypothesis to shipping decision

  1. Days 1–2 / Frame

    Choose the decision, user risk, hypothesis, participant criteria and explicit success threshold.

  2. Days 3–4 / Prepare

    Build three variants, write the task script, pilot once and correct any leading language.

  3. Days 5–7 / Observe

    Run short moderated sessions while one person facilitates and another captures behaviour and quotes.

  4. Days 8–9 / Decide

    Cluster observations, separate frequency from severity and choose what the evidence changes.

  5. Days 10–12 / Iterate

    Revise the weakest part of the flow and preserve the evidence trail from observation to design change.

  6. Days 13–14 / Verify and ship

    Retest the changed interaction, document remaining risk and make the shipping decision explicit.

Write down behaviour before you name the problem

“Participant reopened the price view before selecting Buy” is an observation. “Participant lacked confidence” is an interpretation. Keep them separate. Capture what happened, what the person said, whether they needed help and what the same behaviour could mean in the real product.

Observation-to-decision matrix
Observed patternConsequenceDecision response
Cannot identify the selected contractHigh: action may apply to the wrong instrumentStop shipping; restore persistent contract context and retest
Finds detail only after scanning twiceModerate: slower comparison with recoverable delayImprove hierarchy; verify in the next iteration
Uses a different path but completes correctlyLow: mental model differs without harmful outcomePreserve valid flexibility; avoid forcing one route
States preference without behaviour changeUnknown until connected to a taskRecord as a lead, not a product decision
Severity and decision consequence matter alongside recurrence in a small directional study.

Set the success threshold before the sessions

A success threshold prevents the team from moving the goalposts after seeing a preferred design. Define critical conditions such as: every participant can state which contract is selected; no participant mistakes information for an instruction to trade; and a participant can reach review without moderator intervention. For less consequential details, define what can remain a follow-up rather than a shipping blocker.

The Fisdom option-chain case-study chapter keeps this work connected to the approved product context rather than presenting the testing method as an invented example.

What guerrilla testing can and cannot tell you

A small, quickly recruited group can expose a severe misunderstanding and help you decide what to change next. It cannot estimate how common that misunderstanding is across every customer segment. I would not use five sessions to claim a population-level conversion lift or satisfy a regulatory question.

If you need prevalence, representative segments or formal risk assurance, plan for it. For a focused design decision this week, write the limitations beside the findings and keep moving.

Ship the learning, not only the layout

The final output is a decision record: hypothesis, participants, variants, task, raw observations, severity, changes, verification and remaining risks. Link that record to the design and delivery ticket. This protects the reasoning from becoming “we moved it because users were confused.”

Fast research earns trust through a visible chain from risk to observation to action. That chain is what makes guerrilla testing a method—not just a quick meeting with a prototype.

How many people do you need for a guerrilla usability test?

There is no magic number that turns a quick test into proof. Recruit people with the relevant behaviour to expose serious issues, and keep testing when a revision changes the task. Report who participated and what happened; do not present the count as a population estimate.

Sources and further reading

  1. GOV.UK Service Manual: Using moderated usability testingPractical guidance for tasks, facilitation and observation.
  2. GOV.UK Service Manual: Finding participantsGuidance on recruiting people who reflect the service audience.