Jul 2024
What Search Data Reveals That People Do Not Say
A reading reflection on new forms of data, the difference between correlation and causality, and why data still requires skepticism.

Everybody Lies was written by former Google data scientist Seth Stephens-Davidowitz. Rather than teaching the mechanics of data analysis, the book explores what Google Search data can and cannot reveal—including truths people may not disclose in surveys or conversations.
Reading the book alongside Professor Hung-Yi Lee’s GenAI course and Google’s Data Analytics course raised several questions for me: How do big data and small data differ? How does generative AI relate to data analysis? And how does big data connect to UX research? My current understanding is that generative AI creates new outputs from existing patterns, big-data analysis is often used to identify trends and correlations, and small-scale research can help explore causality and solve specific problems.
Thinking in Practice
Innovation can come from finding a new type of data
The value of big data is not only collecting more information faster. It can also come from finding a new kind of data for an existing problem.
The book describes how combinations of Google search terms can help track indicators such as unemployment more quickly. I was reminded of Google’s Android earthquake-alert system, which uses sensors in phones as “mini seismometers.” I have not researched that example in depth, but it illustrates the same idea: an existing device can become a new source of data.
A useful correlation may appear before its cause is understood
Some examples in the book show that a relationship can be useful even when its cause is not yet clear. Tartar sauce may sell more during hurricanes, and the size of a horse’s left ventricle may correlate with racing performance. In these cases, the relationship can support prediction before it provides an explanation.
Large software companies may run thousands of A/B tests. What surprised me was how a small interface or copy change could affect click-through rates. In the example shown here, adding a right-facing arrow to an ad significantly increased clicks, even though the result is difficult to explain from the interface alone.

3. Data does not remove the need for skepticism
People may present different versions of themselves to friends, surveys, and search engines. Search data can reveal biases that respondents may not state openly, while published numbers and analyses can also mislead.
For analysts, this makes research design important: respondents need conditions that help them answer honestly, and researchers still need to consider confirmation and observer bias. For individuals, the lesson I took from the book is simple:
Do not compare your private searches with other people’s public social-media lives, and remain skeptical of conclusions presented as objective data.
A personal experiment
The book offered me a different perspective through practical examples of data analysis. I could not fully follow the author’s interpretation of some correlations—particularly in the chapters about sex and pornography—and I remain uncertain about those cases.
I have also begun collecting my own spending data. Beyond understanding where my money goes and adjusting categories, I hope the record will help me identify a financial allocation that suits me.