Two decades ago, a psychology experiment with millions of participants was nearly impossible to imagine. Gathering data was expensive and time-consuming, requiring dozens or maybe hundreds of people (often college undergraduates) to troop into a physical lab to take part.
Today, researchers can create an online survey and watch it gather hundreds of thousands of responses from diverse participants around the world. They can access millions of tweets with a few lines of computer code. And they can use newly powerful computer-analysis techniques to glean insight into human behavior from these and other large data sets.
Propelled by these tools, big data research is taking off in fields as diverse as cognitive, personality, social and industrial/organizational psychology.
"Five years ago, many people were talking about big data psychology, but I wondered whether they would actually do it," says Samuel Gosling, PhD, a personality researcher at the University of Texas, Austin, who has been gathering data online since the late 1990s. "But to my happy surprise, they did. At this point it’s well under way."
Fertile research ground
Big data has become a buzz phrase, but what does it mean? How big is big? There’s no one answer to that question, researchers say. When computer scientists talk about big data, they are usually talking terabytes, petabytes or larger—amounts that require distributed computing systems to analyze, says psychologist Sean Wojcik, PhD, a senior data scientist at the digital media company Upworthy and co-author (with Eric Chen, PhD) of "A Practical Guide to Big Data Research in Psychology" (
Psychological Methods
, Vol. 21, No. 4, 2016).
Psychologists, however, rarely work with data sets that large. "We often just use it to mean ‘data that’s much bigger than we’re used to,’" Wojcik says. "There’s no threshold."
That’s because even large data sets that might not impress a computer scientist can provide fertile ground for psychological research. What those data sets consist of varies by subfield in psychology.
Often, they involve online surveys or social media postings. For example, in their research on personality, Gosling and his colleagues combined meteorological data from every zip code in the United States with data from more than 1.6 million participants who took an online personality test. They found that people who grew up in more temperate climates were more likely to be agreeable, open and emotionally stable than people who grew up in colder areas (Nature Human Behaviour, Vol. 1, 2017).
At the University of Pennsylvania, meanwhile, a consortium of psychologists and computer scientists is working on the World Well-Being Project—founded by positive psychology pioneer Martin E.P. Seligman, PhD—an attempt to measure worldwide well-being by analyzing social media postings. In one recent study, they found that by analyzing the language in millions of tweets they could predict which U.S. counties consume more alcohol than others (PLOS ONE, online publication, April 2018).
In a different field, cognitive psychologist Brendan Johns, PhD, is exploring texts, not tweets. Johns, an assistant professor in the communicative disorders and sciences and computational linguistics departments at the University at Buffalo, uses big data analysis methods to analyze Wikipedia, publicly available ebooks and other digital troves of written language. His goal is to understand how people learn the meanings of words from the structure of language, and how that learning affects memory and other forms of cognition.
"We can train our models on a corpus of 2 billion words, and that’s been a big change," he says.
Promise and challenges
So, what do these far-flung applications of big data have in common? Broadly speaking, large data sets change the kinds of questions that psychology can seek to answer, researchers say.
For Lyle Ungar, PhD, a professor of computer science and psychology at the University of Pennsylvania who co-leads the World Well-Being Project, that change is embodied in a shift from "hypothesis testing" to "hypothesis generation." Most of psychology research, he points out, involves coming up with an experiment to test one hypothesis. "But that’s only half of science," he says. "And that’s not big data—which is primarily data driven, not hypothesis driven." In a typical study, for example, Ungar might collect millions of tweets from people with attention-deficit hyperactivity disorder (ADHD), then explore those data to find ways that the tweets of people with the disorder differ from the tweets of people without it, all while having no particular hypothesis in mind ( Journal of Attention Disorders opens in new window, 2017). That kind of insight into the daily experiences of people with ADHD could help lead to better treatments.
Kevin Grimm, PhD, a research methods psychologist at Arizona State University, agrees that searching for the unexpected is an important aspect of big data research, but adds that the key is that big data analysis methods provide a systematic way to do that. "It’s important that we test our hypotheses with confirmatory methods, but also [that we] test for other trends that wouldn’t be found unless you explored."
Another advantage of big data, says Gosling, is that the high-powered studies allow researchers to come closer to understanding the complexity of real-world human behavior.
"Most of human behavior is extremely complicated and you simply cannot examine the interactions of so many things unless you have sufficient power to do it," he says. In traditional psychology research, researchers could examine only a few factors or do studies in extremely controlled environments that may not approximate the real world. "Until the advent of the big data age, we didn’t have tools that were well matched to the complexity of the phenomena we wanted to study."
The potential rewards of exploring large data sets are great, but for many psychologists the barriers to entry can be high. "The question of where to start is a very big hurdle to people," Wojcik says. "A lot of psychologists’ training is in SPSS, and that’s not an ideal tool for analyzing very large data sets."
Learning a new programming language that’s a better fit, like R or Python, can seem daunting. But Wojcik suggests thinking of the time investment as similar to the data-gathering stage of traditional psychology research. Once you learn R or Python, "data collection can be incredibly fast," he says. "In a traditional lab, you might devote several months to collecting data. You can devote that time to learning R instead."
On a broader scale, big data collection—particularly from social media postings—brings up a host of ethical and privacy challenges for psychology and other fields. "I do think about it a lot—what it means to consent," says Ungar. "For example, when participants give me access to their [Facebook] posts, I don’t take what their friends post on their page because those people haven’t consented."
Overall, says Gosling, psychologists can contribute to the debate over data privacy by being good stewards of data themselves and by using their expertise to help understand the factors that play into people’s decisions about privacy and data security.
"How do people decide when they feel safe and don’t feel safe sharing data? That’s a question psychology can help answer."
Interested in learning more? APA conducts an annual Advanced Training Institute on big data. The next one will be held in 2019.

