candid responseConsultation Edition ← Back to the home page

How we keep ourselves honest

Every claim links to the words underneath it.

A report is only worth what its weakest claim is worth. So each figure carries the number of people it is drawn from, each quote is checked against what was actually said, and anything we cannot stand behind is left out rather than softened. Your submission counts as one response however many members are behind it, so its influence rests entirely on the quality of the evidence inside it.

What we will never say

The refusals are the more useful half. Every line in this column is something a competitor could tell you, and we will not.

We will never sayBecauseWe say instead
"68% of members think X" It confuses how often something was raised with how many people believe it. Silence on a theme is not disagreement. "Raised in an estimated 68% of responses (95% CI 64 to 72%, n=88)"
"Representative of the membership" Participation is self-selecting. Government's own consultation analyses state this plainly about their own data. The achieved sample set against your membership profile, and a statement of who is missing
A margin of error There is no valid inferential basis from a self-selected sample to a whole membership. Confidence intervals on the sample proportion only, described as such
"Significantly more likely" between cohorts Themes are model-assigned labels carrying measurement error a significance test does not model. On a self-selected sample a p-value is theatre. Non-overlapping 95% intervals, described as a gap worth examining, never as a tested result
A sentiment score, such as "7.2" Spurious precision on a constructed scale. Sentiment direction per theme, evidenced with quotes
NPS, satisfaction scales, derived importance These need complete structured data on scaled items. Conversational data is missing by design. Add a short structured module, or decline the work
"Worth £Xm to the sector" No valid basis for projecting to a population. A clearly labelled illustrative scenario with its assumptions on the face of it

The governing line. We quantify how often something was raised within the sample we achieved, and never infer to your whole membership. That is the same position UK government takes when it counts themes in free-text consultation responses.

How a member's words reach a page

Six gates. A quote that fails any one of them does not appear, and the default at every stage is silence.

01

The member says it

A conversation of roughly five minutes, stored as their actual words and not only as a paraphrase, so the source of every later claim still exists to be checked.

02

The machine proposes a quote, and the code checks it

Every candidate quote is matched against the transcript and discarded if it is not verifiably there. Paraphrases, inventions, tidied grammar and sentences stitched together from different points in the conversation are all rejected.

03

Contact details are removed, and the removal is logged

Emails and phone numbers are stripped, and we record that we did it. Deliberately narrow, because this is not anonymisation. The member decides that themselves at the next step.

04

The member sees exactly what would be published

Their own words and the description that would sit beside them, because context identifies a firm as surely as a name does. In a 400-member body, "a regional contractor in the North West with about 50 staff" names someone.

05

They choose how far it goes

Not quoted, quoted anonymously, quoted with a description such as firm size and region, or quoted by name. You set the ceiling, the member chooses within it, and the ceiling is enforced on our server rather than in their browser.

06

They can take it back, including after publication

Withdrawal is first class. A withdrawn quote disappears from the next generation of the report. Consent given is not consent locked.

The line is in the database, not the policy. A response carries no name, email or user reference, ever. Identity lives in a separate table holding no reference to any response or consultation. The only path from an answer to a person runs through an explicit permission grant, and a database constraint physically refuses to attach an identity to an anonymous one.

What we have measured, and what we have not

Three independent passes over a real 88-response consultation.

99.2%Mean pairwise agreement. Running the same responses through theme mapping three times produces almost identical answers.
86 of 88Responses mapped identically on every run. The two that moved are the genuinely borderline ones.
0 to 2.3Points of variation in how prevalent a theme was. Four of five themes were identical across all runs.

What these numbers are not. They measure whether the analysis is stable and what it refuses. They do not measure whether the themes are right, because that needs a human-coded comparison on the same data and we do not have one yet. Until we do, we report consistency and refusals, and we do not claim accuracy. That is the honest answer, and it is the reason a person reviews the theme set before anything is published.