← Back to writing list

Jul 2024

What Search Data Reveals That People Do Not Say

A reading reflection on new forms of data, the difference between correlation and causality, and why data still requires skepticism.

Book reflection
Everybody Lies: Big Data, New Data, and What the Internet Can Tell Us About Who We Really Are

Everybody Lies was written by former Google data scientist Seth Stephens-Davidowitz. Rather than teaching the mechanics of data analysis, the book explores what Google Search data can and cannot reveal—including truths people may not disclose in surveys or conversations.

Reading the book alongside Professor Hung-Yi Lee’s GenAI course and Google’s Data Analytics course raised several questions for me: How do big data and small data differ? How does generative AI relate to data analysis? And how does big data connect to UX research? My current understanding is that generative AI creates new outputs from existing patterns, big-data analysis is often used to identify trends and correlations, and small-scale research can help explore causality and solve specific problems.

Thinking in Practice

Innovation can come from finding a new type of data

The book describes how combinations of Google search terms can help track indicators such as unemployment more quickly. I was reminded of Google’s Android earthquake-alert system, which uses sensors in phones as “mini seismometers.” I have not researched that example in depth, but it illustrates the same idea: an existing device can become a new source of data.

A useful correlation may appear before its cause is understood

Some examples in the book show that a relationship can be useful even when its cause is not yet clear. Tartar sauce may sell more during hurricanes, and the size of a horse’s left ventricle may correlate with racing performance. In these cases, the relationship can support prediction before it provides an explanation.

Large software companies may run thousands of A/B tests. What surprised me was how a small interface or copy change could affect click-through rates. In the example shown here, adding a right-facing arrow to an ad significantly increased clicks, even though the result is difficult to explain from the interface alone.

Figure 1 Illustration from the book showing the addition of a button to the right of an ad.
Figure 1 Illustration from the book showing the addition of a button to the right of an ad.

3. Data does not remove the need for skepticism

People may present different versions of themselves to friends, surveys, and search engines. Search data can reveal biases that respondents may not state openly, while published numbers and analyses can also mislead.

For analysts, this makes research design important: respondents need conditions that help them answer honestly, and researchers still need to consider confirmation and observer bias. For individuals, the lesson I took from the book is simple:

A personal experiment

The book offered me a different perspective through practical examples of data analysis. I could not fully follow the author’s interpretation of some correlations—particularly in the chapters about sex and pornography—and I remain uncertain about those cases.

I have also begun collecting my own spending data. Beyond understanding where my money goes and adjusting categories, I hope the record will help me identify a financial allocation that suits me.