top of page
Search

On Literature and Statistics

Writer: Yulia Kuzmina
Yulia Kuzmina
10 minutes ago
3 min read

I’ve finished working on the report for my research grant. It was difficult, but also very interesting, and now I have a whole bunch of ideas for future research. But that’s not what I want to write about today. Or, rather, not exactly.

Reading Lolita in Tehran

Recently, I read a wonderful book by Azar Nafisi, Reading Lolita in Tehran. In a way, it rescued me from the gloom and discouragement that are probably familiar to anyone trying to find their place in life after emigrating, losing their familiar surroundings and work, and perhaps even a little bit of their identity.

The book reminded me how much we need good literature. In any truly good book, we can find something of ourselves and look at our own lives and surroundings through the lens of completely different situations and characters, even if they lived in another country and another time.

It also reminded me how important it is to have something of your own to do—something that can sustain you, even if not materially, then at least emotionally.

Although my work has nothing to do with literature, while reading this book I kept thinking about who I am, what I do, and what I want to do. What is it that I genuinely love doing?

For me, it is working with data.

Modelling the world

At first glance, what could be further from people and their problems, from art and literature, than data and numbers? Sometimes I feel this distance myself. I pull myself away from some fascinating dataset and suddenly wonder: how is any of this actually connected to reality?

When we analyze data and test different statistical models, in a sense we are working with abstractions, with things that do not really exist in the world. Statistical models often ask us to imagine an idealized reality. We talk about averages, expected values, and hypothetical individuals defined by particular combinations of variables, even though no actual person in our dataset may correspond to any of them. What can a regression coefficient tell us about a particular person? Almost nothing.

But perhaps literature does something surprisingly similar. Literary characters are not real people either. They are constructions: simplified, exaggerated, selective, shaped by an author to capture something about human experience. They may resemble people we know, but they do not have to correspond exactly to anyone who has ever existed. In this sense, they too are models, literary models rather than statistical ones.

And this is perhaps why both can tell us something about reality precisely by not reproducing it exactly.

Approximate answer to the right question

The world of statistical models is a world created by humans, just as the world of literature is. And it can be beautiful in its own way. Both can reveal something new about ourselves and the world around us. Statistics allows me to ask questions and get answers.

And perhaps the questions matter even more than the answers. John Tukey, one of the most influential statisticians of the twentieth century and often described as one of the fathers of data science, put it beautifully:

“Far better an approximate answer to the right question, which is often vague, than an exact answer to the wrong question, which can always be made precise.”

So statistics, to me, is not only about finding answers. It is also about learning how to ask the right questions.

For a long time, I tried to find one particular field that I wanted to devote myself to. I have done research in education, cognitive psychology, and psychometrics. I have attended lectures on biology and genetics, studied in a neuroscience master’s program, taken part in archaeological expeditions, and read books on evolution and anthropology.

And I find all of these things fascinating.

But there is something that connects all these very different fields: much of what we know about them would have been impossible without statistics and data analysis.

So how far is data analysis from real-world problems? I don’t think it is far at all.

There is another quote by Tukey that I have always liked:

“The best thing about being a statistician is that you get to play in everyone’s backyard.”

I think this is exactly what attracts me to statistics. I don’t have to choose what I love more. I get to play in everyone’s backyard.

So I’m going to try to bring this blog back to life and write more about interesting datasets, unexpected findings, and research that has taught us something new thanks to statistical models.

And, of course, about the methods themselves too.


 
 
 

Comments


bottom of page