
Codeberg Moves to Block AI Models from Scraping Its Code Repositories
💡 - AI startups may face higher data acquisition costs, squeezing margins and favoring companies with proprietary data. - Data licensing brokers and platforms that facilitate legal code access could see increased demand. - Developers and open-source project maintainers might gain negotiating power to charge for training rights. - Investors in AI infrastructure should monitor similar policy moves at GitHub, GitLab, and Bitbucket for sector-wide risk. - Side hustlers using public code for model fine-tuning must verify terms of use or face potential legal exposure.
Codeberg, the open-source code hosting platform, is updating its Terms of Use to explicitly prohibit the use of its repositories for training large language models. The change, proposed via a pull request on the organization’s own repository, could force AI companies to seek alternative data sources and raise compliance costs.
Codeberg, a popular alternative to GitHub for open-source projects, is extending its Terms of Use to forbid what it calls 'LLM-extrusions' — the extraction of code for training large language models. The proposal was submitted as a pull request on Codeberg’s organization repository, signaling that the platform intends to codify the restriction into enforceable policy. The move reflects growing tensions between open-source communities and AI developers who scrape public code for model training.
The change targets a practice that has become common among AI companies, which often use publicly available code repositories to train models like GPT and Codex. By explicitly banning such extraction, Codeberg aims to protect contributors’ work from being used without consent or compensation. The platform’s community discussions, as captured in the pull request comments, highlight concerns over attribution and the potential devaluation of open-source labor.
For businesses and investors watching the AI supply chain, the policy shift reinforces a broader trend: the cost of obtaining high-quality training data is rising. If other hosting platforms follow Codeberg’s lead, AI firms may need to license code directly from owners or build proprietary datasets from scratch. This could increase operational expenses for AI startups and create new opportunities for data licensing marketplaces.
Real estate and crypto investors may see less direct impact, but the move underscores how regulatory and policy changes in tech can ripple across sectors. For side hustlers running AI tools, relying on scraped open-source code could become legally risky. Developers hosting code on Codeberg might gain additional leverage to monetize their work if data-licensing models emerge.
The pull request is still under review, but its momentum suggests Codeberg’s user base supports the restriction. The final language and enforcement mechanisms will be crucial for determining how strictly the ban is applied. If adopted, Codeberg will join a small but growing list of platforms erecting barriers against unlicensed AI training.
Read the full story
Original reporting and related coverage — attribution links only, not paid recommendations.
Broker buttons use invite / refer-a-friend links (rewards may be capped). Other partner links may pay OppHub a commission at no extra cost to you.
Tools & books on Amazon
Shop Amazon →Relevant gear and reads when you want to go deeper — OppHub may earn from qualifying purchases.
Build My Playbook
Turn this headline into a clear plan: what to watch, how to express it (stocks, ETFs, or options education), and how you’d know you’re wrong — for beginners and active traders. Not personalized advice.
You’ll get theme → ETFs → stocks → options education → side income → kill switches.