Cover image by Philipp Katzenberger
The Model Trained on You
They ask us to believe the model learns from the sum of human knowledge—from libraries and archives and the open web—but that is a story for investors, not for us. Every large language model begins with public data: Common Crawl, Wikipedia, open-source code repositories, the scattered detritus of collective expression rendered into training tokens. But the public internet is noisy, inconsistent, and widely accessible, which means it confers no competitive advantage; it is the floor, not the ceiling. The real differentiator comes from a source much closer to home: from us, in the act of using the product. Chat logs, moderation decisions, support tickets, code snippets, file names, the precise shape of our cursor movements, the corrections we make to what the machine generated—every interaction becomes a training datum, and every datum becomes a commodity. We produce the value; they own the output. That is the structure, and structures do not negotiate. This is not the first time we have seen technology erode our agency—our digital dependency has been reshaping power, privacy, and class dynamics for years—but the scale here is different.
Consider the mechanism. In April 2026, GitHub announced that interaction data from Copilot’s Free, Pro, and Pro+ users—inputs, outputs, code snippets, associated context—would be used to train AI models by default, unless users actively opted out. Data may be shared with Microsoft for AI development purposes; the fine print points toward the same concentration of ownership that defines the entire stack. Business and Enterprise customers are exempt under contract, but the individual developer—the student, the freelancer, the worker—has no such protection, only a toggle buried in account settings. Anthropic made a similar shift in August 2025: Claude conversations now feed into model training unless opted out, with retention periods extended from thirty days to five years. The consent flow is not neutral—the data-sharing toggle comes pre-checked, and the acceptance button locks that choice in, making agreement the path of least cognitive resistance. A Stanford study of six leading U.S. companies found that all of them use user inputs for training by default and that their privacy policies, written in the convoluted legal language that has defined internet-era consent, are functionally illegible to the people they claim to inform. Obfuscation here is not a failure of communication but a material condition of the relation itself, a structural feature that serves accumulation by appearing to serve transparency. The engineering mindset that treats ethics as someone else’s problem is what allows these defaults to be framed as neutral technical decisions rather than value judgments.
This arrangement is a relation of production applied to the most intimate residue of our cognitive lives. In capitalist social formations, labour is alienated because the worker is separated from the product of their work, from the process of production, from their own species-being, and from other workers—alienation understood not as a feeling but as a structural fact, the thing we make belonging to someone else, the activity of making it controlled by someone else, until we become a thing among things, an object to be used. When we type a prompt into a chat interface, we are performing labour—cognitive, communicative, generative labour—but the output of that labour is not ours; it is scraped, stored, aggregated, and fed back into a model that will eventually be sold back to us, often at a higher price and under terms we had no part in setting. The user is simultaneously the worker and the raw material, the producer and the produced-upon, and this dual role is what makes platform capitalism’s extraction so seamless: we cannot withdraw our labour without withdrawing from the tools that increasingly mediate our capacity to labour at all. To understand how we got here, we need to look at the surplus value that is extracted from every interaction and how it is rationalised by a system that treats our biographies as business expenses.
The concept of liberatory alienation—the notion that automation will free us from drudgery so that we may ascend to creativity and self-fulfilment—presumes that a technical capacity can, by its mere existence, upend the relations of production that determine how that capacity is deployed. It cannot. AI within a capitalist framework remains AI controlled by corporations, designed to maximise the rate of return on invested capital, not to redistribute the gains of productivity to those whose activity generated them. The supposed liberation is a thin ideological cover for deeper data-driven control, and we see this most clearly in the structural asymmetry of consent: the ability to opt out is presented as a consumer choice, but the burden of action falls entirely on the individual while the default serves the platform, and the platform knows—because it has the data to prove it—that most people will not navigate the menus, will not read the policy documents, will not remember to uncheck the box. That foreknowledge is not incidental, but the condition of its profitability. A system that depends on the inaction of those it extracts from has engineered that inaction into its architecture, and to call the result a choice is to misname coercion as freedom. The promise of a liberatory future is further undermined by capitalist realism, the pervasive ideology that makes it nearly impossible to imagine a genuine alternative to the current order, let alone build one.
Some argue that privacy-preserving techniques—differential privacy, federated learning, on-device redaction—can solve the problem at the technical level. But these methods do not eliminate privacy harm; they redistribute it across actors, institutions, and stages of the machine learning lifecycle, shifting the site of extraction without altering its logic. OpenAI’s Privacy Filter model, released in April 2026, promises to redact personal information from unstructured text, but foundation models generate privacy violations far beyond what personally identifiable information filtering can detect—individuals can be reidentified from small amounts of unstructured text, and large language models excel at those inferential connections that make redaction insufficient. The problem is not that the technology lacks sophistication but that sophistication is deployed on behalf of capital rather than on behalf of us, and a technical solution to a structural problem is always a palliative administered to forestall the demand for a cure. This dynamic is reminiscent of the perverse joy of checking boxes, where bureaucratic systems replace real human outcomes with the satisfaction of completing processes and following procedures.
What, then, would a cure require? We must begin by refusing the frame of individual consent altogether. Opting out is not liberation; it is a permission slip to be excluded from exploitation, granted at the discretion of the exploiter and revocable at the next terms-of-service update. The real question is not whether we agree to have our data used but why anyone should own the data of our collective lives in the first place—data that is not extracted from a void but produced by collective activity, mediated by platforms we did not design, governed by terms we did not write, and appropriated by entities we did not elect. The training of a large language model on user interactions is not a service improvement; it is primitive accumulation, the enclosure of the digital commons, the conversion of shared social activity into private property. The platform does not create value; it captures value that we produce in common, and the capture is legitimised by a legal apparatus that treats the act of clicking ‘Accept’ as meaningful consent despite the absence of any genuine alternative. To imagine a different kind of relationship with technology, we might look to models of open source that prioritise community and transparency over corporate control, or to the principles of sousveillance that seek to monitor power from below rather than simply being monitored by it.
This is not a problem of bad actors or insufficient regulation. The GDPR, the EU AI Act, the patchwork of state-level privacy laws in the United States—these are reactive structures that chase a moving target, defining harm after the fact and carving out exceptions for ‘legitimate interests’ that reliably align with the interests of capital because capital writes the lobbying briefs and funds the consultations and staffs the revolving doors between industry and oversight. The material reality is that the AI industry depends on the continuous extraction of user-generated data because the public internet is a finite and increasingly polluted resource, because synthetic data degrades model quality over generations, and because product-specific interaction data is the only source of genuine competitive differentiation. If a firm wants a model that outperforms the others, it needs data that no one else has, and the cheapest, most abundant, most continuously replenished source of that data is the user, whose every click, correction, and conversation is harvested by default under the fiction of informed consent. The incentive structure is not hidden; it is openly stated, and incentives embedded in ownership relations drive outcomes far more reliably than regulatory frameworks designed to manage rather than transform those relations. The myth of meritocracy serves a similar function in the political sphere, justifying inequality as the natural outcome of fair competition while obscuring the structural advantages that make the game unfair from the start.
The alienation we experience in our interactions with AI systems is the intended operation of a mode of production that treats human cognitive activity as an input to be optimised, a cost to be minimised, a resource to be exhausted until something cheaper can be synthesised to replace it. The model is trained on us because we are the only source of novelty that cannot yet be fully simulated, and our continued participation is secured not by coercion in its crude form but by the restructuring of work and communication around tools that make non-participation increasingly costly. The modern development burnout machine is a perfect example of this dynamic, where our tools are better than ever yet the cognitive strain of work feels harder, because the underlying pressures to produce and conform remain unchanged. The toggle in the settings menu will not change that relation, only obscure it; the task is not to petition for better privacy policies but to recognise the structure, name it as what it is—exploitation dressed in the language of convenience—and to organise collectively around the principle that the products of our shared cognitive labour belong to us, not to the shareholders of the platforms that merely mediate its expression. The model is trained on us. Will we continue to produce value for free, or demand that the relations of production be transformed to match the social character of the labour that sustains them?