Monday, 23 September 2024

RCE: How Good is your data?

Most of you have now completed the process of gathering data for your assigned country. The dataset you gathered consists of 104 countries coded over 1863 country-years by the AI and over 2414 country-years by you (based on the AI’s responses – several of you extended the coding period significantly). A few of you are still working on your data-gathering; this post will be updated as new data comes in. In the meantime, the existing scores and justifications can be explored interactively here. This post is the first of several, focusing on the question of whether the data you gathered is any good: does it measure the degree of democracy and the type of regime in particular countries over time?

Sources of Error

There are a few of potential sources of error in the data you gathered. First, there is error in the measurement instrument – the questionnaire. It is possible that the questions we asked were not good at determining the degree of democracy in particular states (perhaps they were not good enough at “operationalizing” an unobservable concept like democracy), or that the scales we used to assign points to countries were inappropriate (not clearly separating categories that should be separated). But overall the questions used seem reasonable and similar to questions used by professional organizations trying to measure democracy and non-democracy.

Second, there’s potential error from poor responses to the questions – perhaps the AI’s responses hallucinated facts, or gave biased or otherwise unreliable answers. All of you checked these answers, however, and most of you were generally satisfied with their quality, though perhaps some people checked more thoroughly than others.

Finally, there is simple data-entry error. Some of our questions were on an 8-point scale, some on a 6-point scale, and some on a 4-point scale, which introduced some confusion - I can see that some people occasionally assigned 8 points to a question which should have been marked out of 6, or 6 points to a question which should have been marked out of 4. The questions in section 7 were in any case confusing regarding whether higher values were more peaceful or less peaceful, and I could see that in at least a few cases this affected these responses (both by the AI and by you). The prompt we used had some errors as well; though these were corrected early on, they may nevertheless have affected some responses. The AIs were also bad at adding up scores (some of you as well!) – and in some cases the final score produced does not match the score you would get by adding up the actual sub-scores for each question. I have tried to correct for these problems by generating calculated columns that simply add the scores entered by you to each question in the questionnaire, and caps the maximum score for any sub-question to 8, 6, or 4, depending on the section.

After these corrections, it turns out that both the data you generated and the data the AIs generated is generally pretty decent (Table 1 and Table 2), correlating well with the data generated by both Freedom House and the Varieties of Democracy Institute (both widely used and well regarded), at the 0.8-0.9 level (which is similar to the correlations between democracy scores obtained by different professional measurement projects, which are around the 0.83 level).

This is a good sign! It suggests that our regime data is not bad, at least by the conventional standards of existing political science measurement efforts. Moreover, though the AI scores were well correlated with both V-Dem and Freedom House (at the 0.89-0.86 level – see Table 1), your own scores, which correct the AI scores, were even better correlated (at the 0.91-0.88 level – see Table 2), suggesting that your attempts to “correct” the AIs improved the data, at least if we take V-Dem as the “gold standard” for this sort of measurement. (Freedom House and V-Dem are themselves correlated at the 0.93 level, partly because they measure slightly different things, partly through the inherent uncertainties of this sort of exercise).

Table 1: Correlations between AI-generated democracy scores and scores by V-Dem and Freedom House
term Final AI score (as entered) Adjusted AI score (as entered) Calculated AI score (generated, w. scale corrections) V-Dem Polyarchy score Freedom House score
Final AI score (as entered) NA 0.97 0.97 0.86 0.84
Adjusted AI score (as entered) 0.97 NA 0.95 0.82 0.80
Calculated AI score (generated, w. scale corrections) 0.97 0.95 NA 0.89 0.86
V-Dem Polyarchy score 0.86 0.82 0.89 NA 0.93
Freedom House score 0.84 0.80 0.86 0.93 NA

Table 2: Correlations between student-corrected democracy scores and scores by V-Dem and Freedom House
term Final student score (as entered) Adjusted student score (as entered) Calculated student score (generated, w. scale corrections) V-Dem Polyarchy score Freedom House score
Final student score (as entered) NA 0.97 0.99 0.90 0.88
Adjusted student score (as entered) 0.97 NA 0.96 0.88 0.86
Calculated student score (generated, w. scale corrections) 0.99 0.96 NA 0.91 0.88
V-Dem Polyarchy score 0.90 0.88 0.91 NA 0.93
Freedom House score 0.88 0.86 0.88 0.93 NA

High correlations with professionally-produced scores are nice, but they represent only an average; specific countries may be scored more or less accurately. For example, the AI scored Laos and Hungary most differently from V-Dem (the difference calculated as the root of the mean squared error from the V-Dem score – Table 3); Hong Kong and Jordan were also problematic. But error values relative to V-Dem were generally small.

Table 3: Top differences between corrected AI scores and V-Dem polyarchy score. Higher values indicate bigger differences relative to V-Dem.
Country Root Mean Squared Error with V-Dem score
Laos 0.43
Hungary 0.41
Jordan 0.29
Hong Kong 0.27
Switzerland 0.22
Lebanon 0.22
Oman 0.22
Guinea 0.21
El Salvador 0.20
Kenya 0.20

Similarly, the hardest country after adjustment for you (not the AI) was Sao Tome (relative to V-Dem), with Sri Lanka and Jordan following close behind. But again, divergences from V-Dem were generally small.

Table 4: Top differences between corrected student scores and V-Dem polyarchy score. Higher values indicate bigger differences relative to V-Dem.
Country Root Mean Squared Error with V-Dem score
Sao Tome and Principe 0.42
Sri Lanka (Ceylon) 0.30
Jordan 0.29
Kuwait 0.26
Panama 0.24
Switzerland 0.22
Hong Kong 0.22
Afghanistan 0.21
Guinea 0.21
Oman 0.20

The particular tool you used does not seem to produce much of a difference in terms of accuracy, but it seems that Claude was the least accurate overall, while Gemini and Meta were a bit more accurate. People who used multiple tools selected the best scores; it helped to do some comparison shopping.



Figure 1: Overall error (RMSE) with respect to V-Dem score, by AI tool

Finally, it looks as if the further back in time, the more likely there are some errors (Figure 2 and Figure 3). For example, the AI scores tend to have a greater divergence from V-Dem for years further in the past, though the end of the cold war seems to have been tricky too. Some modeling suggests, however, that it’s not just how far back in time the country is coded, but how large it is: smaller countries have larger errors relative to V-Dem, perhaps because there are fewer sources to consult, or fewer people studying them.



Figure 2: Root mean square error per year relative to V-Dem score, corrected AI answers.


Figure 3: Root mean square error per year relative to V-Dem score, corrected student answers.

Personalism and Accountability

Six questions in the questionnaire measured the degree of accountability of the leadership, and function as the inverse of an index of personalism: the more accountability, the less personalism. Since a leading theory of the causes of repression and war is that personalism at least increases the probability that these things would happen (see Weeks 2014, in the recommended readings; Peceny and Butler 2004), it is worth looking at the reliability of this particular measure. There is little consensus on how to measure personalism in political science, but I calculated an index of personalism for my book, and the accountability score you gathered is well correlated with it (negatively, as we should expect - see Table 5).

Table 5: Correlation between the POLS209 index of accountability and the Márquez index of personalism.
term POLS209 index of accountability Márquez index of personalism
POLS209 index of accountability NA -0.85
Márquez index of personalism -0.85 NA

Country Scores

It is possible to graph the POLS209 democracy scores per country. As we can see, most of the time both the AI scores and your own are pretty close to the V-Dem scores (sometimes within the measurement error interval), but there are some exceptions – the countries noted above in particular.



Figure 4: All AI country scores. V-dem polyarchy score in dotted line for reference.


Figure 5: All student country scores. V-dem polyarchy score in dotted line for reference.

References

Peceny, Mark, and Christopher K Butler. 2004. “The Conflict Behavior of Authoritarian Regimes.” International Politics 41 (4): 565–81. https://doi.org/10.1057/palgrave.ip.8800093.
Weeks, Jessica L. 2014. Dictators at War and Peace. Cornell University Press.

1 comment:

  1. This is so interesting! On the topic of further dated back countries being harder to research, my two regimes included one in the late 80s early 90s, which was indeed harder to find sources, and the other was in 2021, which really is not studied yet outside of freedom house. I strangely experience both sides of the coin. Other wise this was a very eye opening assignment into what AI really understanding, it rarely can get into the weeds of an issue and do real academic analysis from what I saw with mine.

    ReplyDelete