7 October 2026

Why is AI so contentious?

Brookings Institution | Martin Beraja, Noam Yuchtman

The Brookings Papers on Economic Activity fall 2026 conference featured research by Martin Beraja and Noam Yuchtman examining why artificial intelligence remains highly contentious despite its economic potential. A Pew Research Center survey reveals that 50 percent of United States adults feel more concerned than excited about daily artificial intelligence integration.

This widespread public ambivalence stems from a deep-seated moral grievance. The authors term this the technology's "original sin." Developers built these models using the words, creations, and expertise of human creators without obtaining their consent. Consequently, conventional policy remedies like universal basic income, retraining programs, or development pauses fail to address this core ethical violation. Resolving this friction requires novel public or private frameworks that restore ownership and attribution to the original creators. Without these mechanisms, public resistance to technological adoption will likely persist, complicating future commercial deployment across the global economy.

Comment

The reliance of developers like OpenAI on the Common Crawl dataset to train large language models exposes a fundamental vulnerability in the emerging technology ecosystem. By scraping intellectual property without explicit consent, the Common Crawl ingestion pipeline generates a persistent moral and legal friction that traditional regulatory frameworks cannot easily resolve. This tension complicates the commercial integration of OpenAI's enterprise models, as public backlash threatens the social license required for widespread deployment.

Consequently, the long-term viability of OpenAI's GPT-4 architecture depends on establishing verifiable attribution mechanisms rather than relying on post-hoc licensing agreements. This shift will likely force a restructuring of the Common Crawl data acquisition pipeline, driving up the operational costs of training future frontier models. Ultimately, the financial burden of securing legitimate training data will consolidate market power among a few well-capitalised firms like OpenAI.

Strategic Question for Discussion
If the Common Crawl dataset becomes legally or socially unviable for training, how will frontier developers like OpenAI maintain model performance without triggering unsustainable data acquisition costs?
The transition away from uncompensated data harvesting points to a bifurcated development landscape. My assessment is that OpenAI will increasingly rely on synthetic data generation and exclusive, high-value licensing partnerships to bypass the Common Crawl bottleneck. This trajectory indicates that while model capabilities may continue to advance, the entry barriers for new competitors will rise significantly.
Share your assessment in the comments below.
💬