Yang, Eddie and Margaret E. Roberts. "The Authoritarian Data Problem." Journal of Democracy, vol. 34 no. 4, 2023, p. 141-150. Project MUSE, https://dx.doi.org/10.1353/jod.2023.a907695.
This article by Eddie Yang and Margaret E. Roberts discusses the effect of increased use of AI technology on non-democratic regimes. AI works off of a global exchange of data that allows the AI to be knowledgeable about a range of diverse opinions, cultural norms and context based on the data set the technology has been trained on. Because of the globalised nature, AI is likely to become a platform for geopolitical conflict. The focus of this article is the way repressive regimes affect AI, and vice versa. AI largely relies on the data provided by authoritarian regimes for important information about these regimes. The tendency of these regimes to restrict and erase important information coming out of their communities, and instead inserting manipulated data and propaganda can lead to a skew or inaccuracy in AI’s output, the example they used is where AI may parrot government slogans or justification for anti-democratic behaviour. The data AI is exposed to has the potential to lead to replication of political biases, creating a potential political tool for autocrats. On the other side of the same coin, democratic countries openness with data provides repressive regimes with examples of subversive content to feed their AI, in order to increase knowledge and output, that is limited within their own environments because as a result the censorship practices. The authors even point out the irony of this stating “This is the irony of the free world—data from democracies can be used to boost the performance of repressive AI.” Although this information provided by democracies is not enough to prevent the issue of the ‘dictators dilemma’ from cropping up.
Polluting Data
This article explains, in line with what has been taught in lectures, that non-democratic repressive states, in order to further their narrative and ensure the health of the regime, will control and shape media, both traditional and social media. Censorship and propaganda are a means of manipulating public opinion, controlling the flow of information and ensuring there is no publicised regime dissent. Countries such as China, Russia and Iran who had sophisticated censorship and propaganda systems are able to tightly monitor and shape the online content in order to serve regime interests by deleting posts, manipulating search results and ban accounts as well as influence self-censorship by civilians who fear legal or extra-legal consequence for posting political content. This article then asks the question, how do these practices affect AI? The data, or lack thereof, is then collected by AI systems to be used in the output of both predictive and generative AI. Because of the massive manipulation of the data, the results generated by AI can be used as a tool for these regimes through polluting the data with propaganda and misinformation that supports their views will alter how AI in democratic countries will respond to political questions.
The authors have evidenced these claims by comparing the results of AI that has been trained on China’s censored platform Baidu Baike which if their version of Wikipedia, although the submission monitored and censored by the government against AI from Chinese Wikipedia, an uncensored platform. And noticed that the AI would create different meanings depending on the platform they were trained on. For example output trained from Baidu Baike were more closely associated with negative adjectives than the Wikipedia trained AI whereas words associated with the Chinese Communist Party, and historical events and figures would be more positively associated under Baidu Baike than Wikipedia. Responses to queries about Mao Zedong in Wikipedia trained AI would include both achievements and failures, whereas Baidu Baike’s response would not include any failures in its summary.
Language plays a particularly important role as the majority of the data and texts coming in the languages of these states (e.g. Chinese, Russian, Farsi) are coming from within the censored regime and therefore, they have found that AI responses in those languages are more likely to reflect the authoritarian information and ideals to people speaking that language, even outside of the regime. English output on the other hand is found to largely reflect the opinion of democratic countries.
The Dilemma
The flow of information from democratic states creates a problem for autocracies given that it may diversify the output of AI in these autocratic regimes. This is a danger as its output may begin to reproduce content that is considered subversive within the state and risk corrupting the controlled information within the regime, the authors refer to this as ‘reverse pollution’. Thus, autocracies will seek to control and censor the output of AI in the same manner that they control the media. China for example has banned the use of ChatGPT, and their regulation for other AI models stipulates that any output must be within the parameters of their ‘socialist views’ implying requirements for any output that is a departure from the regime’s worldview should be censored before reaching the user.
The repression that is rampant in autocratic regimes is what gives rise to what the authors coin ‘the Dictators Dilemma’. This is the tension between the corruption of data generation in these regimes, (censorship, propaganda, misinformation etc.) and the reliance of AI on accurate and quality data in order to perform. By limiting the information available they stunt the abilities of their AI thus losing the ‘AI race’ but this restriction of information also prevents the generation of valuable political data that enable it to be used as a repressive tool e.g. by auto-censoring or predicting regime dissent. It is not only government censorship, but the self-censorship of citizens as a result of surveillance practices in these regimes that prevent AI from collecting relevant data. People are not willing to openly share their opinions or express political disobedience, because AI cannot compute information that is not provided in the data, this creates bias in the data that is not necessarily reflective of genuine public opinion thus an ineffective tool for predicting dissent or monitoring public opinion.
Purpose of the Article
The wider context that this text exist within is the discussion of the expansion of AI, a new technological tool rapidly increasing in popularity, that is inherently political because of the social nature of the data is collects. This has led to questions about the data they are trained with, and whether this can create a bias that affects the output when asked political questions. Because this is a technological tool, the automatic jump is that the solution will come from a technological fix, through tweaking the algorithm. Yang and Roberts take and interdisciplinary approach, given the political nature of the data, preferring to ask questions through the lens of political science. The overarching purpose of this text is not necessarily providing answers - but putting forwards areas for further investigation. They raise questions around the state influence on AI data and the causal relationships between that training data and the output and particularly because of the globalised exchange of ideas and data that AI is a platform for, how to safely govern this technology in line with democratic principles.
Critique
While this article puts forward a good argument for the effect of AI creating this Dilemma for autocratic regimes, the examples they use mainly focus on advances countries with sophisticated censorship methods and advances technological capabilities. It would be interesting to know what the effect the rise of AI has had on autocracies that do not have the technological capabilities to censor in effectively or control the output of AI in the same way.
No comments:
Post a Comment