Questionnaires in User Research
[Work-in-progress] This page is intended to perhaps provide some increased awareness on the range of uses of questionnaires in user research. The primary point is that, anecdotally, surveys are sometimes assumed to be suitable only for answering extremely simplistic questions - primarily at the level of counting some kind of atheoretical, discrete category. In contrast, as used in, e.g., psychological science, questionnaires - more precisely, developed scales involving combinations of questionnaire items - are at heart about making relationships and patterns measurable.Rich versus impoverished questionnaires
This isn't to say that simple counting and calculating averages doesn't play a role - it's what I'd call "basic descriptives", step 1 of the analysis or results section of a report. But I'd want to emphasize what you could be counting and averaging - these can be psychological, theoretical, interesting concepts. E.g., components of a model of trust in government or a given organisation; or a range of beliefs and attitudes about AI; or a granular scoring of specific educational needs. The set of variables you choose to measure can be impverished or rich. Using only an impoverished survey will obviously give you boring results - that's not the fault of the method.
Of particular interest here are developed scales - that is, sets of items of which you combine the responses into a single value. The "developed" bit means - the scale isn't something you pulled out of your neck, it's something from the scientific literature. People with significant expertise will have invested significant effort to generate, test, select, and validate the scale. Using developed scales means your far less likely to be building your house on sand than if you just make up Likert scale that sounds good to you - although, pragmatically, that might also happen. But what you do is at least test the reliability and validity of the scales you used, developed or ad-hoc. (Reliability basically means, if you did your measurement twice, would you find the same scores? There are methods and scores that can tell you that. Validity means, do those scores measure what you think they measure. Assessing validity is a bit more of a creative endeavour, as you have to look at whether patterns of findings support your interpretation.)
Comparisons
One of the points where you start getting at the real value of a (rich) questionnaire comes when you start looking at theoretically meaningful differences. There are two types of comparison I tend to use.
First, comparison between subscales of a developed scale. For instance, the Trusting Beliefs scales has three subscales, for benevolence, integrity and capability. It can be interesting to test whether one of those does conspicuously worse than the others. For this kind of comparison, you need some kind of multidimensional concept that has meanignful, relevant comparisons.
Second is comparison of the average scores of one particular variable over different sub-populations, for instance, different job levels in a workforce. This could tell you, for instance, where attitudes towards some topic of interest are especially low. To do this, you need one or more sets of categories to create sub-populations, or "cells", to compare. This can be called "mapping". With two sets of categories, heatmaps give a nice visualisation of the distribution of scores over the two sets of categories (e.g., job level and gender).
Associative relationships
The other main kind of analysis to do on survey data is to test associations, or connections. This is where correlations or regressions come in. In this case, we're looking at relationships between "measurement-on-a-scale", or scale, variables - that is, variables that give you a measurement like you'd get using a ruler or a thermometer. These are approximately continuous values, where the numerical value is meaningul as a quantity, not just a label (like when you might arbitrarily code, say, gender using male=0 and female=1).
Associations tell you: People who score higher on variable X tend to score higher/lower on variable Y. With rich data, you might have complex patterns of such associations that can, combined with and informed by theory and subject matter expertise, give you interesting interpretations. For instance, you might want to know what predicts higher versus lower job satisfaction within a given team, out of a set of possible predictors derived from a work psychology model.
Note that the well-known rule of "correlation is not causation" is both improtant to keep in mind, but also important not to let block reasonable interpretations. It may be valuable just to select down a limited set of data-driven hypothesis on causal relationships that can then be assessed conceptually and selected for experimental testing.
Quantitative personas
Briefly: With rich data, basic dimension reduction and clustering techniques can be combined to create quantitative personas. These are a kind of "skeletal" personas that are derived from data. A persona in this sense is defined as having a particular profile of high and low scores over a range of variables.
The underlying methods give you an "efficient" set of personas, i.e., ones that cover a wide range of the variations in your data with only a restricted combination of scales.
These personas can subsequently of course be further developed and embellished, e.g., with nicknames and imaginary photographs - of course, at the risk of introducing the well-known concerns around biases, stereotyping, and lack of justification of traditional personas.
Beyond
There are many more analysis methods, which can be more complex, more principled, or more informative; but also likely more demanding of the amount and quality of data.
On sample sizes
It's important to understand the critical factor of sample size in quantitative research. Using too-small samples will means trying to interpret noise - a waste of time at best, deceptive at worst. There are statistical procedures for getting ballpark figures that will be needed for specific analyses, details of methodological choices, and specific expected effects; these procedures are called statistical power analyses. Such analyses are concerned with the ability to justifably draw inferences from the data from the observed sample to the population of interest. Note that the "population" can be as abstract as you like - e.g., all analogous people at other times and places. There's an element of interpretation here.
Conclusion
The use of questionnaires offers more than is perhaps sufficiently widely recognized in user research, but necessary conditions include (1) rich rather than impoverished questionnaire design, (2) sufficient analytical know-how, and (3) adequate sample size. Practically, there is also a need for sufficient stakeholder support, as there may be barriers due to ignorance or, to name an elephant in the room, perverse incentives. The latter can include stakeholder fear of results seen as threatening, e.g., if the results involve an assessment, or researcher resistance to a competing methodology. Over time, however, I hope awareness and acceptance of such methods to shift as their value becomes more widely known - in terms of information and trustworthiness.