What Does The FCA Expect From Consumer Duty Communications Testing?
The Consumer Understanding outcome asks firms for something specific: an approach to testing communications that provides assurance that customers can identify and understand the information they need to make effective decisions. Not an assumption. Not a legal sign-off confirming the wording is technically accurate. Assurance, based on evidence, that real customers get it.
That distinction matters more than it first appears. A communication can pass every compliance review, use approved wording, and still leave customers unable to say what their product costs, what happens if they miss a payment, or what action they need to take next. Testing exists to catch that gap before customers do.
The FCA is explicit that testing should be proportionate. Proportionate to the purpose of the communication, its importance to the customer’s decision, the audience it reaches, the vulnerability characteristics likely to be present in that audience, and the potential harm if the information is misunderstood. A low-stakes marketing email and a default notice warrant different levels of scrutiny, and firms are expected to make that judgement rather than apply one standard everywhere.
It’s also worth being clear about what testing isn’t. It isn’t a proofreading exercise to confirm a communication is factually correct or legally cleared. Those checks matter, but they answer a different question: is the wording right, not did the customer understand it. And testing doesn’t end when a communication goes live. Monitoring how customers actually respond after launch, through contact volumes, journey behaviour, complaints and other signals, is part of the same ongoing obligation, not a separate afterthought.
One point is worth stating plainly, because it shapes everything that follows in this guide: the FCA has not mandated a fixed testing methodology. It takes an outcomes-focused, proportionate approach. There is no single approved test, no universal sample size, and no checklist that automatically satisfies the requirement regardless of context. Firms are expected to design an approach that fits their communication, their customers, and their risk, and to explain and evidence that approach. Anyone telling you otherwise, including any agency, is overstating how certain the rules are.
Which Consumer Communications Should You Test
Not every communication needs the same depth of testing, and testing everything to the same standard is neither proportionate nor a good use of budget. A risk-based prioritisation approach lets firms direct the most rigorous testing at the communications where misunderstanding causes the most harm.
Useful criteria for prioritisation include:
- The importance of the customer decision the communication supports
- The potential consumer harm if the information is misunderstood
- The complexity of the product, message or underlying terms
- The size of the customer population affected
- The presence of vulnerability characteristics within the target audience
- Whether the communication is new or an established, previously tested piece
- Existing evidence of misunderstanding, gathered from complaints, contact-centre queries or journey abandonment data
- Recent changes to the product, pricing, terms or customer journey that the communication describes
In practice, this tends to push certain communication types toward the top of the list: product information, onboarding journeys, financial promotions, renewal communications, price and fee disclosures, risk warnings, arrears and default communications, cancellation information, communications about significant changes to terms, and digital journeys and calculators where customers self-serve decisions with real financial consequences. These are the moments where a customer is deciding something, committing to something, or losing money if they misread the message. Lower-risk, purely informational content can usually be tested more lightly, or monitored rather than formally tested each time it’s updated.
Consumer Duty Communications Testing Methods
There is no single best method for testing customer understanding. The right method depends on the question you need answered, and different questions call for different tools.
Qualitative Depth Interviews
Depth interviews are the right tool when you need to understand why customers misunderstand something, not just whether they do. Think-aloud exercises, where customers talk through a document or journey as they work through it, expose the exact point where confusion sets in. Asking customers to paraphrase a key message in their own words, rather than simply confirming they read it, reveals whether the message actually landed. Probing specific terminology lets you find the words and phrases that don’t mean what the firm assumes they mean, and open discussion often surfaces interpretations the firm never anticipated.
Quantitative Comprehension Testing
Where qualitative work diagnoses the problem, quantitative testing measures its scale. This means testing whether customers understand specific information: correct comprehension of key facts, unprompted recall (can they state the information without being shown it again), recognition (can they identify the correct information when shown options), understanding of the consequences of a decision, and the ability to identify the appropriate next action.
This is a different exercise from asking customers whether they thought a communication was clear. Perceived clarity and actual comprehension frequently diverge. Customers regularly report that something was “clear” while getting the underlying facts wrong. If a test only asks about clarity, it measures confidence, not understanding.
Surveys And Customer Callbacks
For lighter-touch or ongoing testing, where the risk profile doesn’t warrant a full research programme every time, surveys and customer callbacks offer a proportionate alternative. The FCA’s 2026 review specifically cites surveys, comprehension checks and customer callbacks among the approaches it has seen firms use proportionately. These methods work well for monitoring understanding on an ongoing basis, rather than as a one-off deep dive, and can be built into business-as-usual processes more easily than a full qualitative or quantitative study.
A/B And Comparative Testing
When a firm is redesigning an existing communication, comparative testing pits the current version against the redesign and measures which one performs better on comprehension, error rates, task completion, drop-off and correct next-action selection. This is a useful discipline because it resists a common assumption: that a prettier or shorter communication automatically communicates better. Sometimes it does. Sometimes a shorter version cuts information customers need. The outcome has to improve on the measures that matter, not just on visual appeal.
Behavioural And Journey Data
Drop-off rates, repeated page visits, help-centre contacts, calls following a communication, error rates, abandonment and the actions customers actually take all provide signals about where a communication might be failing. This data is genuinely useful for identifying potential problems worth investigating further.
It shouldn’t, however, be treated as proof of understanding on its own. The FCA has specifically criticised firms that collect management information without showing how it connects to customer understanding, and has noted that sales performance or an absence of complaints doesn’t, by itself, provide reliable assurance that customers understood what they were told. A customer can buy a product, never complain, and still have misunderstood a material fact about it. Behavioural data points you toward where to look. It doesn’t answer the question by itself.
Consumer Duty Testing Sample Sizes: How Many Customers Do You Need?
This is the question that generates the most anxiety, and the honest answer is: there is no universal FCA-mandated sample size for communications testing. Any framework, template or agency that hands you a fixed number and calls it the compliant sample size is giving you false certainty.
What actually determines whether a sample size is defensible is a combination of factors: the objective of the test, whether the method is qualitative or quantitative, the size and diversity of the customer population, how many distinct customer segments matter to the analysis, the level of risk associated with misunderstanding, the level of statistical confidence the firm wants in its conclusions, the expected variation in how different customers respond, and whether the firm intends to generalise the findings to its entire target market or treat them as indicative.
Sample Sizes For Qualitative Testing
Qualitative research doesn’t work on a magic number. The relevant concept is saturation: the point at which additional interviews stop surfacing new problems or new patterns of misunderstanding. A relatively small number of in-depth interviews can genuinely uncover usability and comprehension problems that a much larger, shallower survey would miss. That said, a small qualitative sample demonstrates that problems exist and gives insight into why; it does not, by itself, demonstrate the scale of a problem across the whole customer base, and it shouldn’t be presented as though it does.
Sample Sizes For Quantitative Testing
Quantitative sample sizes are a statistical question, driven by the size of the customer population, the confidence level required, the acceptable margin of error, the expected distribution of responses, and whether the firm needs reliable results for subgroups within the sample, not just the sample as a whole.
The table below is illustrative only. It shows, in general statistical terms, how sample size requirements tend to shift as precision demands increase; it is not an FCA rule or benchmark.
| Requirement | Illustrative Effect On Sample Size |
| Wider margin of error (e.g. +/-10%) | Smaller sample sufficient |
| Narrower margin of error (e.g. +/-3%) | Larger sample required |
| Higher confidence level (e.g. 99% vs 90%) | Larger sample required |
| Analysing the total sample only | Smaller sample sufficient |
| Reliable analysis of an important subgroup (e.g. a vulnerable cohort) | Each subgroup needs its own adequately sized sample |
The last row is where most testing programmes fall short in practice, and it deserves its own section.
When Your Overall Sample Size Can Be Misleading
A headline sample size can look reassuring while hiding a weak evidence base. Consider a test with 300 respondents overall, which sounds like a solid sample. If only 15 of those 300 represent an important vulnerable cohort, whatever conclusions the firm draws about that cohort’s understanding rest on 15 responses, not 300. The overall number tells you almost nothing about the reliability of the subgroup finding.
This is why sample composition matters as much as headline sample size. A defensible testing programme looks at who is actually represented within the sample, not just how large the sample is, and treats subgroup findings with the caution their smaller base warrants.
Testing Communications With Vulnerable Customers
Vulnerable customers deserve their own section because aggregate results routinely obscure exactly the problems Consumer Duty is designed to catch. A communication that performs well on average can still fail badly for a specific cohort, and averaging that failure away defeats the purpose of testing in the first place.
Building this into a testing programme starts with identifying the vulnerability characteristics genuinely relevant to the target market, rather than a generic list. That might mean lower financial capability, lower digital confidence, sensory or cognitive barriers, language needs, or other factors specific to the customer base a particular communication reaches. Recruitment then needs to actively include participants who reflect those characteristics, and accessibility requirements need to be built into the research design itself, not bolted on afterwards.
The FCA’s March 2026 review is directly relevant here. It identified insufficient testing across accessibility needs, language requirements and lower financial capability as an area firms need to improve, and it gives examples of firms that test specifically with vulnerable cohorts and measure their comprehension rather than relying on how the wider customer base responds. The clear message: if a firm’s testing sample doesn’t include the customers most likely to struggle, its conclusions about understanding are incomplete, however strong the headline numbers look.
What Should You Measure In A Consumer Understanding Test?
Understanding isn’t one thing. It breaks into four distinct areas, and a testing programme that only measures one of them will miss real problems.
- Comprehension: did customers understand the important information itself? Measured through comprehension accuracy, key-message recall and misunderstanding rates.
- Findability: could customers actually locate the information they needed, when they needed it? Measured through task completion rates and time taken to locate important information within a document or journey.
- Decision support: could customers use the information to make an informed decision, not just recite it back? Measured through appropriate next-action selection and the quality of the decisions customers describe making based on what they read.
- Behaviour and outcomes: what actually happened after customers received the communication? Measured through the downstream behavioural signals discussed earlier, contact volumes, drop-off, errors, alongside the caveat that these signals support rather than replace direct comprehension evidence.
A testing programme built around all four areas gives a far more complete picture than one focused solely on comprehension scores, because it connects understanding to what customers actually do with that understanding.
Setting Pass/Fail Criteria And Comprehension Thresholds
Success criteria need to be defined before results come in, not chosen afterwards to fit whatever the data shows. That sequencing matters for credibility as much as for rigour.
Not all information carries equal weight. Critical information, the facts that materially affect a customer’s decision or financial outcome, should carry more weight in a pass/fail assessment than minor or supporting details. And averages are a poor substitute for genuine assurance: an overall comprehension score of 85% can still mean a critical subgroup understood almost none of the key message, which is exactly why subgroup results need to be examined alongside the headline figure, not instead of it.
Whatever threshold a firm sets, it should be able to explain why that threshold was chosen, not simply state the number. This is worth underlining because of how the FCA’s 2026 review has sometimes been (mis)quoted in the market: the review describes one firm using an internal target of at least 80% correct recall as its own internal benchmark. That is one firm’s own practice, cited as an example, not a figure the FCA requires or endorses as a universal pass mark. Treating 80%, or any other number, as a regulatory threshold misrepresents the guidance and could leave a firm relying on a defence that doesn’t exist.
From Testing To Improvement: The Test-Learn-Retest Cycle
Testing that stops at “we found a problem” doesn’t deliver assurance. Assurance comes from closing the loop: identifying the risk, establishing a baseline understanding of current performance, testing, diagnosing why customers struggled, redesigning the communication, retesting the redesign, launching it, monitoring real-world outcomes, and reviewing the whole cycle periodically.
The step firms most often skip is the retest. Conducting research and making changes based on it feels like the job is done, but without retesting, there’s no evidence the changes actually fixed the problem rather than simply looking like an improvement. The FCA has specifically flagged poor practice here: firms that made changes to communications but never subsequently established whether those changes worked. Good practice, by contrast, means recording what changed, why it changed, and what impact that change actually had on customer understanding. That record is what turns an assumption into evidence.
What Evidence Should Firms Keep?
This is the section worth treating as a working template, because the FCA’s core concern in this area is straightforward: firms saying testing occurred without being able to produce meaningful evidence of it.
A defensible evidence trail for each tested communication should include:
- The specific communication and version tested
- The reason it was prioritised for testing
- The testing objectives
- The methodology used and why it was chosen
- Recruitment criteria for participants or customers
- The sample size and the rationale behind it
- Participant or customer characteristics, including relevant vulnerability characteristics
- How vulnerable cohorts were represented within the sample
- The questions or tasks used in the research
- The results, in full
- Subgroup findings, not just aggregate results
- Problems identified through the testing
- The decisions made in response
- Changes implemented as a result
- Retest results following those changes
- Approval and sign-off records
- Post-launch monitoring data
- Any follow-up actions taken
Held together, this record does two things at once: it demonstrates the testing happened, and it demonstrates the firm can explain and defend why it happened the way it did. Both matter under scrutiny.
Common Consumer Duty Communications Testing Mistakes
Several of these map directly onto weaknesses the FCA’s March 2026 review identified as needing improvement.
- Testing readability rather than understanding
- Asking “was this clear?” instead of testing actual comprehension
- Relying on internal compliance review as a substitute for customer testing
- Treating an absence of complaints as proof customers understood
- Using unrepresentative convenience samples
- Ignoring vulnerable customer cohorts within the sample
- Reporting only aggregate scores, with no subgroup analysis
- Testing once and never retesting after changes
- Making cosmetic changes without addressing the underlying cause of misunderstanding
- Collecting management information without connecting it to customer decisions or understanding
- Failing to document why changes were made, or what effect they had
Building A Defensible Consumer Duty Testing Framework
- Prioritise communications according to consumer risk
- Define precisely what customers need to understand
- Select the research method that fits the question being asked
- Recruit a sample that is relevant and large enough to be defensible, including relevant vulnerable cohorts
- Measure comprehension and decision-making, not perceived clarity alone
- Analyse overall and relevant subgroup outcomes
- Improve the communication based on the evidence gathered
- Retest important changes before treating them as resolved
- Monitor real-world outcomes after implementation
- Document every decision and maintain a full evidence trail
FAQs About Consumer Duty Communications Testing
Does The FCA Specify A Minimum Sample Size For Consumer Duty Testing?
No. There is no universal FCA-mandated sample size. Defensible sample sizes depend on the testing objective, the method, the customer population, the number of segments that matter, and the level of confidence required. What the FCA expects is a sample size the firm can justify, not a fixed number it can point to.
Do All Customer Communications Need To Be Tested?
No, and testing everything identically isn’t what proportionality requires. Firms should prioritise communications using a risk-based approach: importance of the decision, potential harm from misunderstanding, complexity, population size, vulnerability, and whether the communication is new or has changed. Lower-risk communications can be monitored more lightly.
Can Readability Scores Demonstrate Consumer Understanding?
Not on their own. Readability metrics assess how easy text is to read at a surface level; they don’t measure whether customers actually understood the substance of the message, retained it, or could act on it correctly. They can be a useful early screen, but they aren’t a substitute for testing comprehension directly.
Should Consumer Duty Communications Be Tested Before Or After Launch?
Both. Testing before launch helps catch comprehension problems before customers are affected. Monitoring after launch, through behavioural data, contact-centre signals and ongoing checks, is also part of the obligation, since real-world use can reveal issues that pre-launch testing didn’t surface.
How Often Should Customer Communications Be Retested?
There’s no fixed schedule the FCA mandates. Retesting is expected after material changes to a communication, and periodic review is good practice for higher-risk communications even without a specific change, since customer populations, products and context can shift over time. The frequency should track the risk level of the communication.
What Evidence Should Firms Retain From Communications Testing?
A full record covering what was tested and why, methodology, sample size and rationale, participant characteristics and vulnerability representation, results including subgroup findings, decisions made, changes implemented, retest results, sign-off, and post-launch monitoring. The goal is evidence that shows both what was done and why it was defensible.
Need Independent Evidence That Your Communications Work?
Testing your own communications internally has an obvious limitation: the team that wrote the communication is rarely the right team to judge objectively whether customers understood it. Independent testing removes that conflict, and it gives compliance, marketing and the board a shared evidence base rather than competing assumptions about how customers actually respond.
Hub Agency’s Specialist Compliance Support practice runs Consumer Duty communications testing for financial services firms who need more than a tick-box exercise. That means qualitative and quantitative research designed around the specific decision a communication supports, sampling that properly represents vulnerable and higher-risk cohorts rather than glossing over them, and a documented evidence trail built to withstand scrutiny, not assembled after the fact. With 15 years of specialist financial services experience, we understand both sides of the brief: communications need to be bold enough to convert and compliant enough to approve.
If you need to know, with evidence rather than assumption, whether your customers actually understand what you’re telling them, get in touch with Hub Agency’s Specialist Compliance Support team.