Volunteers outdoors wearing matching shirts, standing together

Data Critique

What information is included in our dataset?

The dataset used for this study comes from the World Happiness Report 2020 (WHR2020) and consists of 153 countries. Each record in the dataset represents an individual country. The main characteristic of interest is the Ladder Score which reflects the average response to the Cantril Ladder. Additionally there are 6 other variables that are associated with a country’s overall happiness; these are GDP per capita, social support, healthy life expectancy, freedom to make decisions about their lives, generosity and corruption perception. The database also displays the confidence intervals around the Ladder Scores. Along with regional data on where the countries are located and the dystopia residual, which is the portion of a country’s score attributable to the 6 explanatory variables. A couple other variables from other sources (i.e. World Bank, WHO) and associated with Gini Index and health are also included.

What information, events, or phenomena does our dataset illuminate?

Because all countries answered the same questions using the same scale, the database allows for discovering relationships about global trends associated with happiness (e.g. whether there is an association between countries’ levels of happiness with each other or by region; whether there is an association between happiness and income, health, freedom or whether countries are happier or less happy than predicted by their GDP). Additionally, the dataset also shows how much each of the 6 variables (log GDP per capita, social support, healthy life expectancy, freedom to make decisions about their lives, and generosity and corruption perception) explains the overall ladder/happiness score of a country, allowing the dataset to show how different countries might have happiness influenced in different ways.

What can this dataset not reveal?

One main thing the dataset cannot show is variation in happiness inside a country. Each country has one average number with some residual, so we cannot see differences between differing regions, rich and poor people, or young and old people. Additionally, it only shows correlation and not causation, and it only shows how people rate their life on one question, not what happiness really means to them.

How was this data generated?

The data in our dataset was generated from the World Happiness Report. In order to collect the data, they used surveys from the Gallup World Poll. People in each country were asked a life-evaluation question called the Cantril Ladder. They imagine a ladder from 0 to 10, where 0 is the worst possible life and 10 is the best possible life. Then, they rate where they feel their life is right now. Those answers are averaged for each country to create the main “Ladder score” or happiness score. The rankings usually use three year averages, so one unusual year does not affect a country’s happiness ranking too much.

What are the original sources?

The original sources include the World Happiness Report 2020, the Gallup World Poll, the World Happiness Report 2019 online data, the World Bank Global Database of Shared Prosperity, and the World Health Organization. The World Happiness Report 2020 provides the main happiness ranking, Ladder Score, and calculated “explained by” variables. The Gallup World Poll provides survey-based variables and survey metadata, such as sample size, interview mode, languages, excluded population, social support, affect, confidence in government, and related measures. The World Happiness Report 2019 online data provides historical variables from 2005–2018, including positive and negative affect, standard deviation of ladder score, Gini-related variables, and confidence in government. The World Bank provides income and inequality measures such as median income/consumption and Gini Index. Finally, the WHO provides health-related variables, including suicide rates, mortality rates, pregnancy and birth indicators, public health measures, sanitation, and physical health indicators.

Who or what organization funded the creation of the dataset?

Most of the research costs were covered by a series of research grants from the Ernesto Illy Foundation and illycafe; other institutional sponsors included the Sustainable Development Solutions Network, Center for Sustainable Development at Columbia, Center for Economic Performance at the LSE, UBC Vancouver’s Economic school, and Oxford’s Wellbeing Research Center. The main data collector for the dataset was Gallup Research Poll. Additional financial partners mentioned include the Davines Group, Blue Chip Foundation, The William, Jeff and Jennifer Gross Family Foundation, and Unilever.

What information is left out of this spreadsheet?

One factor left out of the dataset is that it does not include each person’s individual responses or general information such as their age, gender, status, level of education, or which part of the country they live in. Another important factor that is not included in the dataset is the explanation or circumstance behind people’s ratings. Having all the responses combined together for each country limits our ability to see the differences between people within a country. If the spreadsheet was the only source of information given, we would inaccurately assume everyone in a country has the same amount of happiness. A final missing factor is people’s backgrounds such as culture, experiences, relationships, or dynamics. Therefore, the spreadsheet gives a good general idea of happiness around the world, but it doesn’t include details needed to deeply understand each country’s happiness.

Data Ontology

With some of the top contenders coming from Nordic countries, this can develop a perception that other places should be socially constructed in a certain way in order to achieve a similar World Happiness Level. This neglects how different countries may naturally have more variation depending on factors such as gender, geography, and socioeconomic status—making an overgeneralization to why these countries may be ranked as such. Additionally, focusing on factors such as GDP to rank a country’s happiness may come from Western standards of associating money and economic success with lifestyle. This can be challenged by how other lower ranking countries may still have typically “happy” people despite not having the same luxuries of the countries at the top of the list. Considering that many of the non-Western countries differ significantly from Western ones in terms of ideology, determining countries’ subjective happiness based on objective values such as GDP can create misleading ideological effects coming from Western norms.