Who Gets to Teach the Machines? Julian Posada on Platform Extractivism, Data Labor, and the Politics of Artificial Intelligence

INTERVIEW SERIES: The Hidden Architecture of Technology INTERVIEWER: Gokhan Colak INTERVIEWEE: Julian Posada

Assistant Professor of American Studies & Co-Director of the Certificate in Computing, Culture, and Society at Yale University | Julian Posada

From Platform Capitalism to Platform Extractivism

Your research examines the extractive dynamics embedded within digital platforms. How would you define platform extractivism, and in what ways does it differ from earlier forms of capitalist extraction? More specifically, what kinds of resources are being extracted by contemporary platforms—data, attention, labor, knowledge, social relations, or something else—and how does this extraction remain largely invisible to the users who participate in these systems?

Platform extractivism describes a deliberate process through which digital platforms organize the extraction and commodification of value from human labor, knowledge, and the social and territorial conditions that make them possible. This process remains part of capitalism, and it is rooted in the longer legacies of coloniality. What is new is the growing importance of data as an object of extraction and, above all, the platform as an infrastructure that actively organizes this process.

Extractivism has always depended on infrastructure. Extracting minerals requires more than a mine: it requires roads, railways, ports, ships, energy systems, and the labor needed to operate them. Platform extractivism similarly connects workers, users, data, and technology companies. As Chapter 3 of the book argues, platforms are not passive marketplaces; they actively recruit, rank, discipline, and make workers dependent on particular flows of work. And digital infrastructure never floats free of physical infrastructure: it rests on data centers, cables, electrical grids, semiconductor supply chains, devices, and human bodies.

Data is central, but it should not be treated as a natural raw material, following Geoffrey Bowker’s formulation in Memory Practices in the Sciences that data must be “cooked with care.” Data is an abstraction from reality, produced through choices and processes including collection, classification, and interpretation. Drawing on Karl Polanyi, I therefore describe data and labor as fictitious commodities (the latter being part of Polanyi’s original list): the market treats them as tradable inputs while obscuring their inseparability from human beings, collective knowledge, and social relations.

The extraction remains invisible because much of it occurs behind the interface. Amazon Mechanical Turk, one of the earliest major microtask platforms, offers the perfect analogy. It was named after the eighteenth-century chess automaton that appeared autonomous while concealing a human operator inside. AI presents a similar facade. We see a facial-recognition system or a chatbot, but not the servers, cables, communities, and data workers whose annotations and feedback made its performance possible.

The Invisible Workers Behind Artificial Intelligence

Artificial intelligence is frequently presented as an autonomous technological achievement, yet contemporary AI systems depend on extensive forms of human labor, including data annotation, content moderation, classification, transcription, evaluation, and other forms of computational labor. How should we rethink the concept of AI labor in light of these often invisible workers? To what extent does the language of “automation” conceal rather than eliminate human labor?

The history of automation is not simply a history of machines replacing workers. It is also a history of people operating machines, adapting their work to them, repairing them, and supplying the judgments that allow them to function. AI continues this pattern. The language of automation can conceal human labor precisely because the human contribution is moved away from the moment when a user encounters the system.

In the book, I define data work as the labor involved in the collection, curation, classification, and verification of data, which also includes the evaluation of model outputs. With large language models, workers write demonstrations, rank responses, test safety boundaries, and provide feedback that shapes how systems behave. The InstructGPT research, for example, makes explicit that human demonstrations and rankings were used to fine-tune a model through reinforcement learning from human feedback. The model may generate an answer automatically, but that apparent autonomy rests partly on prior human judgment.

We should therefore think about data work in different contexts and as performed by different social actors. It can be a full-time occupation in a business-process outsourcing center, fragmented platform labor completed from home, or an additional task absorbed into another profession. The term itself emerged partly from research on doctors, nurses, and other medical professionals who input, correct, and interpret data while doing clinical work. In Venezuela, the population I studied was varied, and many workers had bachelor’s degrees. Economic crisis and pandemic restrictions made platform data work one of the few available income sources, but the broader point is that almost anyone can become a data worker when AI is inserted into a job or everyday workflow.

The action of these systems may be considered “automated” but it’s always important to look at how the system was produced. The important questions are where the labor has been relocated, how visible it is, who is recognized as a worker, and under what conditions that work is performed.

Data Labor and the Question of Value

Data is often described as the fundamental resource of the contemporary digital economy. Yet data does not simply exist—it is produced through human activity, interaction, communication, and labor. How should we understand the relationship between data production and value creation? Do users, platform workers, and communities effectively perform forms of uncompensated or undercompensated labor when they generate the data from which technology companies derive economic value?

Recent developments on machine learning research have contributed to the current relationship between data and value. Earlier symbolic approaches to AI tried to encode rules explicitly. This second wave of AI instead emphasized learning from statistical patterns in large datasets. As internet use expanded and more human activity became “datafied,” companies gained access to the data and computer resources needed to train increasingly complex models. Data did not simply become valuable on its own; it became valuable because institutions built systems capable of collecting, organizing, and converting human activity into a machine-readable form.

That conversion is a production process. Blog posts, books, images, search queries, and online interactions have histories and authors. Annotated datasets add another layer of labor: workers decide what an image contains, whether a text is hateful, or which model response is better. These users, workers, and communities can therefore contribute value without receiving a proportional share of it. Hamid Ekbia and Bonnie Nardi call this dynamic “heteromation”: people create economic value through clicks, profiles, posts, and other interactions while their contribution is hidden, unpaid, or treated merely as use of a service. Platform data workers are more explicitly recognizable as workers, yet they are frequently underpaid and denied the protections associated with employment.

I would not reduce every online action to a wage-labor relation, but neither should we ignore the asymmetry. Companies can appropriate collectively produced knowledge and turn it into privately owned systems. The value of AI is created not only by engineers or by the model itself, but through an extended social process involving users, data workers, families, communities, and the infrastructures that sustain their participation.

The Geopolitics of AI Production

The development of artificial intelligence is frequently represented as a global technological project, but the infrastructures, capital, expertise, labor, and data required to build AI systems are distributed very unevenly across the world. How does the contemporary AI economy reproduce existing geopolitical inequalities between the Global North and the Global South? In your view, does AI represent a new form of technological dependency, and if so, what forms does this dependency take?

This question reaches the central argument of the book: we must situate AI within longer histories of capitalism, coloniality, and extractivism while still identifying what is distinctive about it. AI directly feeds on data, or abstractions of reality and human knowledge, while requiring extraordinary concentrations of capital, computing power, and infrastructure.

The resulting production system is global. It includes mineral extraction, chip fabrication, device assembly, data centers, electrical grids, undersea cables, datasets, and dispersed human labor. Kate Crawford’s Atlas of AI is especially useful for seeing these material, labor, and environmental layers together. The label “cloud” obscures the fact that AI depends on land, water, energy, and bodies located in particular places. Much of the lower-paid data work, resource extraction, and manufacturing is located outside the centers that own the most profitable firms and models.

At the same time, the capacity to develop leading models and capture their returns is highly concentrated. The 2026 Stanford AI Index shows the United States and China trading the lead in model performance, while the United States retains enormous advantages in private investment and the production of top-tier models. Meanwhile, UN Trade and Development warns that many countries—primarily in the Global South—remain excluded from the infrastructures and capabilities needed to shape AI’s future.

This is a form of technological dependency, but it is not historically unprecedented. Countries and communities may supply minerals, energy, labor, languages, and data while depending on foreign cloud providers, chips, models, standards, and platforms. Value moves from peripheries to operational centers, and dominant taxonomies can flow back in the other direction. AI therefore extends older dependency relations through a new data-intensive infrastructure: the inputs are global, but ownership, agenda-setting power, and profit remain concentrated.

Epistemic Power and the Production of Knowledge

AI systems increasingly participate in the production, classification, ranking, and circulation of knowledge. This raises questions not only about technological power but also about epistemic power: who determines what becomes visible, searchable, classifiable, credible, or knowable. How do the datasets, taxonomies, models, and infrastructures of AI influence what kinds of knowledge are recognized or marginalized? What happens to forms of knowledge that remain poorly represented within dominant datasets?

AI systems do not encounter knowledge neutrally. Long before a model ranks or generates an answer, people and institutions have decided what to collect, how to classify it, which categories count, and whose interpretation will be treated as ground truth. As Geoffrey Bowker and Susan Leigh Star show in Sorting Things Out, classification systems shape both worldviews and social relations; they are part of the infrastructure through which reality becomes administratively and computationally legible.

In Platform Extractivism, I relate this concept to a double helix of extraction and imposition. Knowledge and labor flow toward powerful firms, but meaning also flows back: clients instruct workers how to interpret the world and encode their categories into data. Platform extractivism is therefore not only an appropriation of knowledge; it is also an imposition of particular ways of knowing.

I identify this epistemic dominance at three connected levels. First, dominant industry narratives make technological expansion, national competitiveness, and promises of future human benefit appear more urgent than present labor or environmental harms. Second, managerial algorithms and platform interfaces measure speed and conformity while limiting workers’ ability to question instructions, appeal decisions, or communicate contextual knowledge. Third, task documents define taxonomies from the top down. Our data-production dispositif research shows how these institutional, technical, and discursive mechanisms work together.

Gender and hate-speech classification make the stakes concrete. In one task, “other” did not recognize trans or nonbinary people; it was effectively reserved for images with no discernible face. Research such as Joy Buolamwini and Timnit Gebru’s Gender Shades demonstrates the consequences of building gender classification systems around reductive categories and unrepresentative data. In a hate-speech task I studied, a Venezuelan worker classified a post calling for Latinos to be expelled as hateful, but the client’s hidden answer apparently did not; the worker was removed from the task. A contextual judgment was converted into a binary answer, and the client’s worldview prevailed.

Knowledge that is poorly represented does not simply remain absent. It can be treated as error, noise, or an exceptional “other,” while systems trained on dominant categories reproduce that marginalization at scale. Epistemic diversity therefore requires more than adding data. It requires redistributing the authority to define categories, purposes, and meanings, including to workers and the communities affected by the systems.

Whose Culture Becomes Training Data?

Large-scale AI systems are trained on enormous quantities of cultural and linguistic material produced by societies around the world. Yet the communities that generate this material do not necessarily control how it is collected, transformed, or commercially used. How should we think about the relationship between cultural production, data extraction, intellectual property, and collective ownership in the age of generative AI? Could the extraction of cultural and linguistic resources by AI companies be understood as a contemporary form of digital colonialism?

Generative AI does not merely ingest data; it transforms cultural production. In Platform Extractivism, I describe knowledge as a cumulative, collective resource. Its value emerges from generations of individual and social activity, yet platforms can appropriate and commodify it as though it were freely available raw material.

Copyright is an essential part of this debate, but it does not resolve every issue. Discussions of generative AI are sometimes reduced to a misleading binary: either models store complete copies of their training materials, or they store nothing and simply “learn” as humans do. We need to distinguish among the copying involved in constructing training datasets, the statistical relationships encoded during training, and the possibility that parts of protected works can subsequently be reproduced. A. Feder Cooper and James Grimmelmann offer a useful vocabulary here: memorization, extraction, regurgitation, and reconstruction are related but distinct processes. Not all model learning constitutes memorization, but some models can reproduce substantial portions of particular works under certain conditions.

At the same time, copyright tends to organize cultural production around identifiable works and individual or corporate rights holders. It is less equipped to address languages, styles, traditions, and bodies of knowledge created collectively across generations. A community may have a strong claim to govern the use of its linguistic or cultural resources even when no single person owns them under conventional intellectual-property law. The political question is therefore broader than infringement: who has the power to collect culture, define acceptable use, transform it into private infrastructure, and capture the resulting value?

My own approach is not to name an entirely new colonialism. I situate platform extractivism within older and continuing patterns of capitalism and coloniality. Coloniality is not a completed historical event; it is a lingering structure of accumulation, domination, and epistemic hierarchy. Silvia Rivera Cusicanqui’s image of ideas flowing like rivers from South to North captures this complexity: knowledge is extracted, absorbed into larger intellectual currents, and often returned without recognition or control for the communities from which it came. Generative AI extends that historical problem through contemporary platforms and models.

Automation, Precarity, and the Transformation of Work

AI is often justified through promises of increased productivity and the liberation of human beings from repetitive work. At the same time, automation can intensify workplace surveillance, fragment employment, deskill occupations, and create new forms of precarious digital labor. How do you understand this contradiction? Is AI actually transforming the nature of work, or is it primarily transforming who bears the economic and social costs of technological transformation?

The apparent contradiction begins to dissolve once we stop treating automation as a binary choice between human work and machine replacement. My focus is on a division of labor in which humans and machines continue to work together, even as human contributions become hidden, poorly compensated, or displaced toward the margins of organizations. Automation rarely eliminates the need for labor altogether; more often, it reorganizes who performs which tasks, under what conditions, and in which parts of the world.

David Autor has shown why technological change has historically transformed the composition of employment rather than simply causing work to disappear. Automation substitutes for workers in some tasks, complements them in others, and can generate new forms of labor demand. But the persistence of employment at the aggregate level should not be confused with the absence of displacement or harm. Daron Acemoglu and Pascual Restrepo, for example, document employment and wage losses in American labor markets heavily exposed to industrial robots.

The consequences therefore depend not only on whether jobs continue to exist, but on where work goes, how it is divided, how workers are paid and managed, and which institutions absorb the disruption. Platforms can create access to income for populations excluded from formal labor markets, including migrants and people who need to work from home. At the same time, they can fragment employment, intensify surveillance, weaken bargaining power, and transfer the costs of equipment, training, downtime, and social protection to workers and their families. Venezuelan data workers experienced both sides: platforms offered income during a severe crisis while leaving households and communities to absorb the risks.

The book describes two recurring promises: freedom from work through automation, and freedom of work through flexible platforms. Neither promise is distributed equally. AI is transforming the content and organization of work, but it is also transforming who bears the economic and social costs of that change. And it’s also important to point that productivity is not the same as emancipation. The political question is who controls the technology, who captures its gains, and whether workers have the power and protections needed to shape the transition.

AI Ethics Beyond Individual Responsibility

Much of the contemporary debate surrounding AI ethics focuses on issues such as bias, transparency, fairness, privacy, and responsible use. These are important concerns, but do you think this framework is sufficient? What changes when we move from an ethics of AI systems toward a political-economic analysis of the infrastructures, corporations, labor relations, ownership structures, and extractive processes through which those systems are produced? In other words, can AI truly become “ethical” without addressing the material conditions of its production?

Bias, transparency, fairness, privacy, and responsible use are indispensable concerns, but they are not sufficient. In my 2020 commentary “The Future of Work Is Here”, I examined early AI ethics frameworks and found that labor was rarely mentioned. When it appeared, it was usually as a future problem of technological unemployment. Yet AI was already reshaping work, and workers around the world were already producing the data on which the industry depended. The “future” of work was present, but it remained invisible in frameworks largely written from the perspective of privileged economies.

The problem is not that principles are meaningless. It is that principles without institutions, enforcement, or political power can become a substitute for change. Brent Mittelstadt’s argument that principles alone cannot guarantee ethical AI remains important: apparent agreement on words such as fairness can conceal disagreement about whose interests count and who is accountable when those principles are violated—a reflection that is also important for current discussions on AI “alignment.”

Instead, I prefer to focus on political-economic analyses that move upstream. This approach asks who owns the infrastructure; who can afford the computing resources; where the minerals, energy, and data originate; how labor is contracted; and who captures the value. It also reveals how responsibility is displaced along the supply chain. For example, in the book, I describe how some AI companies claim that working conditions are the responsibility of data vendors, while some vendors argue that clients force prices downward. Here, the worker disappears between two institutions pointing at one another.

AI cannot become ethical without addressing these material conditions. An adequate framework must include fair pay and contracts, safe conditions, collective representation, supply-chain responsibility, and meaningful participation by workers and affected communities. The Fairwork AI principles are useful precisely because they connect AI governance to labor rights and worker voice. Ethics should not begin only after a system is deployed and produces a biased output. It must begin with the labor, resources, ownership relations, and decisions through which the system comes into being.

Can AI Infrastructure Become Socially Accountable?

If contemporary AI infrastructures are shaped by concentrated corporate ownership, enormous computational resources, proprietary datasets, and global networks of precarious labor, what would meaningful social accountability actually look like? Should AI infrastructure be treated primarily as a private technological asset, or could certain forms of AI infrastructure be understood as a public or collective resource? What institutional, regulatory, cooperative, or community-based models might make AI development more accountable to the people whose labor and data sustain it?

Meaningful social accountability begins when the people whose labor, data, knowledge, and communities sustain AI possess more than visibility. They need decision-making power. The central principle of the final chapter of Platform Extractivism is “no governance without workers.” Workers’ lived experience should not be treated as supplementary evidence gathered after experts have designed a system; it should be foundational to how problems, standards, and remedies are defined.

That requires institutions for representation: unions, worker associations, collective bargaining, participatory documentation, and co-design. It also requires enforceable responsibility across the AI supply chain. Companies that purchase data work should not be able to outsource accountability along with the labor, while vendors should not be able to blame poor conditions entirely on client prices. Transparency is useful only when it is connected to rights, remedies, and the power to change a practice.

AI infrastructure should not be understood exclusively as a private asset. The United States’ National AI Research Resource, which expands researcher access to computing, datasets, models, software, and expertise, demonstrates that governments can treat parts of AI capacity as shared research infrastructure. Public provision, however, is not automatically democratic. Accountability also depends on who sets priorities, whose data are included, what public obligations attach to access, and whether affected communities can refuse or redirect a project.

Regulation and independent scrutiny remain essential. Supply-chain due diligence, labor-centered dataset documentation, licensing, and third-party audits can make obligations traceable and enforceable. Research on AI accountability infrastructure is valuable because it shows that evaluation tools alone are insufficient; auditing also needs resources for harm discovery, communication, advocacy, and consequences. Fairwork’s cloudwork research offers a complementary model by evaluating platforms according to pay, conditions, contracts, management, and representation.

Finally, ownership itself can change. Platform cooperatives place workers or users in shared control of the technology, data, production process, and surplus. No single institutional model will dismantle platform extractivism. Meaningful accountability will require a plurality of public, regulatory, cooperative, and community-led arrangements—but all of them should be judged by whether those who sustain AI can genuinely influence its development and share in its benefits.

Beyond Extraction: Imagining Alternative AI Futures

Finally, if the dominant model of AI development is characterized by the extraction of data, labor, knowledge, and cultural resources, what would an alternative model of AI production look like? Could we imagine AI systems built around principles such as data sovereignty, worker ownership, cooperative infrastructures, community governance, fair compensation, and epistemic diversity? More fundamentally, what would need to change in our understanding of technology, labor, and value for AI to become not merely more efficient, but genuinely more democratic and socially transformative?

An alternative AI future does not have to reproduce the scale, universality, or priorities of the dominant industry. The assumption that progress always means larger models, more data, and greater computing power is itself part of the extractive model. We should imagine not one universal alternative, but many situated systems developed around the needs, knowledge, and values of particular communities. Smaller scale can be a feature rather than a deficiency.

A powerful example is Te Hiku Media, the nonprofit Māori radio organization profiled in Karen Hao’s “A New Vision of Artificial Intelligence for the People”. Te Hiku and its partners developed speech- recognition and natural-language tools to support the revitalization of te reo Māori. Their Papa Reo project begins with a community-defined problem and treats data sovereignty as foundational. Under its Kaitiakitanga license, data is cared for through guardianship rather than treated simply as property, and benefits are meant to flow back to the communities from which the data originates. The community governs the purpose, terms, and appropriate scale of the system.

Cooperative models offer another direction. Platform cooperatives show that workers and users can jointly own and govern digital infrastructure. James Muldoon’s Platform Socialism pushes this argument toward a broader horizon of social ownership and democratic control. Applied to AI, this could include worker-owned data platforms, community-governed datasets and models, public computing resources with enforceable social obligations, and institutions in which affected people can shape or refuse development.

Three concepts must change. Technology should be understood not as an autonomous engine of progress, but as a collectively governed means whose purposes are politically chosen. Labor should be recognized not as an invisible cost, but as knowledge and co-creation, with fair compensation, representation, ownership, and decision-making power. And value should not be measured only through profit, productivity, scale, or benchmark performance. It must also include dignity, cultural survival, care, ecological sustainability, and community capacity.

Whatever capabilities AI may acquire, the measure of success should be collective human flourishing rather than the concentration of benefits in a few corporations and countries. A democratic AI future begins by asking not simply what a system can do, but who defined the problem, who made the system possible, who governs it, and who is able to benefit.

Leave a Reply

Discover more from PR CARNET WORLD

Subscribe now to keep reading and get access to the full archive.

Continue reading