This development creates a paradox for the platform. Wikipedia remains a critical foundation for the knowledge base of large language models, serving as a vast trove of human-created, verified information. However, the very technology that relies on this data is intercepting the traffic that historically kept the site afloat. For Wikimedia, traffic is not merely a metric of popularity; it is the engine of its financial survival. Approximately 80% of the organization’s operating budget comes from small donations, many of which are prompted by banners displayed to visitors on the site. Meehan explained that fewer readers visiting the site directly translates to fewer potential future editors, disrupting the pipeline that has allowed Wikipedia to grow for two decades.
The Economic Tension of Free Data
At the heart of the issue is a complex economic dynamic. Wikipedia makes its content freely available to the public and has no intention of changing this open-access model. This stance leaves the organization with limited leverage over AI companies that consume its data at massive scales but do not compensate it. Wikimedia operates a program that charges large-scale commercial users for reliable, high-volume access to its data. Publicly disclosed enterprise customers, whose payments constitute 10% of the organization’s operating expenses, include Amazon, Google, Microsoft, Meta, and Perplexity.

Notably, OpenAI and Anthropic, two of the largest developers of generative AI models, are not among the publicly disclosed enterprise customers. Meehan stated that while Wikimedia has other agreements with companies that are not publicly identified, the reception from the broader tech industry has been mixed. Some companies have embraced the principle that those benefiting from Wikipedia should help sustain it, while others have acknowledged excessive data usage but declined to pay for enterprise access. “We’re not asking for charity,” Meehan said, emphasizing the professional nature of the transaction.
“Wikipedia has arguably never been more valuable to the broader information ecosystem. We’re the backbone,” Meehan said when asked about Wikipedia’s importance to AI knowledge engines.
The distinction between using Wikipedia data for AI models and using AI to create Wikipedia content is also a point of contention. While AI tools are already being utilized within the Wikimedia community for specific tasks such as translation, identifying broken links, and helping new editors spot errors, there is a strict boundary regarding content generation. Meehan was clear that AI should assist human editors rather than replace them. “I don’t think you will ever see AI writing or generating Wikipedia articles,” she stated. When asked if articles written five years from now would still be created by humans for humans, she answered affirmatively, reflecting the volunteer community’s preference for human-authored content.

Beyond the immediate financial and operational challenges, the shift toward AI-mediated information access raises broader questions about global equity in knowledge. Wikipedia exists in roughly 300 languages, covering many that are not commercially attractive markets for AI companies. Meehan highlighted the risk of a new knowledge divide, where access to high-quality information may increasingly depend on paid AI subscriptions rather than free, open access. If commercially developed AI becomes the primary gateway to knowledge, the languages and topics that do not yield high commercial returns may be neglected.
Wikipedia spent the initial phase of the generative AI boom establishing itself as an indispensable resource for machine learning. Its current challenge is ensuring that the AI ecosystem does not erode the human and financial foundations that made it indispensable in the first place. As page views continue to shift behind the screens of chatbots and search assistants, the Wikimedia Foundation is navigating a delicate balance between maintaining its open ethos and securing the resources needed to keep the platform operational for the next generation of users and editors.



