Scaling Law Uncertainty
The assumption that AI capability improves predictably with more compute and data — the “scaling law” — is increasingly uncertain, with evidence of diminishing returns and alternative pathways to capability improvement that disrupt investment theses built on raw scale.
What It Is
The “scaling hypothesis” — that simply training larger models on more data reliably produces more capable AI — dominated AI research and investment from roughly 2018–2024. This hypothesis justified the enormous infrastructure buildout: if capability scales with compute, then the company that can spend the most on compute will have the most capable models.
Sara Hooker (SSRN, 2026) argues this era is ending: “The relationship between training compute and performance is highly uncertain and rapidly changing.” Alternative levers — architectural innovations, data efficiency, reasoning methods, specialised models — are gaining relative importance. This doesn’t mean scaling is irrelevant, but it means raw compute is no longer the dominant determinant of frontier capability.
The CSET (January 2026) report on AI R&D automation adds a further uncertainty: even within AI R&D itself, there is no consensus on whether increasing automation of the research process will accelerate progress or lead to plateau. Workshop participants with different assumptions about how AI R&D works disagreed on trajectories even when shown the same empirical evidence.
Why It Matters (for Investors)
If scaling laws are weakening, several investment implications follow. First, NVIDIA’s pricing power depends in part on the assumption that frontier labs will continue scaling compute indefinitely — commoditisation risk increases if scaling is less productive. Second, smaller labs and academic groups may be able to close capability gaps with frontier labs through architectural innovation rather than raw compute spend — reducing the competitive advantage of massive capital. Third, the timeline for AI reaching various capability thresholds becomes more uncertain, making investments dependent on specific capability milestones higher-risk.
The VUB “Benchmarks Saturate” study (January 2026) adds a measurement problem: as models approach human-level performance on existing benchmarks, the benchmarks themselves become unreliable measures of true capability improvement — making it harder to track progress independently.
Evidence & Examples
- Hooker: “Academia has been marginalized from meaningfully participating in AI progress and industry labs have stopped publishing” — creating opacity precisely when measurement is most needed (
ssrn-5877662 (1).pdf) - Hooker’s thesis: scaling has produced a “massive windfall in capital for industry labs” and “fundamentally reshaped the culture of conducting science” — but the scaling formula is changing, and “key disruptions lie ahead” (
ssrn-5877662 (1).pdf) - CSET workshop (July 2025): experts disagree on whether AI R&D becoming more automated will accelerate or plateau AI progress; “new data on how AI R&D automation is progressing in practice may be insufficient to resolve conflicting perspectives” — the different camps make different assumptions that lead to different interpretations of the same evidence (
CSET-When-AI-Builds-AI.pdf) - CSET: existing benchmark evaluations are “insufficient for measuring, understanding, and forecasting the trajectory of automated AI R&D” — the measurement infrastructure doesn’t exist to reliably track progress (
CSET-When-AI-Builds-AI.pdf) - VUB “Benchmarks Saturate” (Jan 2026): when models surpass human-level performance, human-judged benchmarks lose discriminative power — the judge can no longer distinguish between models better than themselves (
2601.19532v1.pdf) - Epoch AI estimates effective compute for training AI models is rising 10x annually, but Jones (2026) notes the capability per unit of compute has also been rising at a similar rate — making raw compute spending a less clear signal of frontier advantage (
AIandEconomicFuture.pdf)
Tensions & Open Questions
- The “intelligence explosion” uncertainty: CSET documents that “intelligence explosion” scenarios — where AI rapidly self-improves — are neither confirmed nor ruled out by current evidence. They are low-probability but non-negligible, and difficult to detect in advance. This is the tail risk that most investment analyses do not adequately price.
- Measurement opacity: If scaling is uncertain AND benchmarks are saturating AND industry labs have stopped publishing, investors are pricing AI capability trajectories with very limited public information. This creates both risk (prices could be very wrong) and opportunity (information advantages are possible for those closest to frontier labs).
- Non-scaling alternatives: Hooker argues that “more interesting levers of progress” beyond scaling are emerging — but doesn’t specify which. CSET notes that architectural innovation, data efficiency, and improved training methods are candidates. Investments that benefit from architectural improvement rather than raw scale may be better positioned if Hooker is right.
- Competitive implications for open-source: If scaling advantages diminish, open-source models (which benefit from architectural innovation shared publicly) may close the gap with proprietary frontier models faster. This would accelerate LLM commoditisation.
Related Concepts
LLM Commoditization · AI Investment Thesis Capex and Returns · Recursive Self-Improvement and AI R&D · AI Capability Measurement · The Scaling Myth Is Finally Cracking
Onzekerheid Rond Schaalwetten
De aanname dat AI-capaciteit voorspelbaar verbetert met meer rekenkracht en data — de “schaalwet” — is steeds onzekerder, met bewijs van afnemende meeropbrengsten en alternatieve routes naar capaciteitsverbetering die investeringsthesen gebaseerd op ruwe schaal verstoren.
Wat Is Het
De “schaalhypothese” — dat simpelweg grotere modellen trainen op meer data betrouwbaar meer capabele AI oplevert — domineerde AI-onderzoek en -investering van ruwweg 2018 tot 2024. Deze hypothese rechtvaardigde de enorme infrastructuuropbouw: als capaciteit schaalt met rekenkracht, dan heeft het bedrijf dat het meeste aan rekenkracht kan besteden de meest capabele modellen.
Sara Hooker (SSRN, 2026) stelt dat dit tijdperk ten einde loopt: “De relatie tussen trainingsrekenkracht en prestaties is zeer onzeker en verandert snel.” Alternatieve hefbomen — architectuurinnovaties, data-efficiëntie, redeneermethoden, gespecialiseerde modellen — winnen relatief aan belang. Dit betekent niet dat opschaling irrelevant is, maar het betekent dat ruwe rekenkracht niet langer de dominante bepalende factor van frontier-capaciteit is.
Het CSET-rapport (januari 2026) over automatisering van AI-onderzoek voegt een verdere onzekerheid toe: zelfs binnen AI-onderzoek zelf bestaat er geen consensus over of toenemende automatisering van het onderzoeksproces de vooruitgang zal versnellen of tot een plateau zal leiden. Workshopdeelnemers met verschillende aannames over hoe AI-onderzoek werkt, verschilden van mening over trajecten zelfs wanneer ze hetzelfde empirische bewijs kregen gepresenteerd.
Waarom Het Belangrijk Is (voor Investeerders)
Als schaalwetten verzwakken, volgen er meerdere investeringsimplicaties. Ten eerste is NVIDIA’s prijsmacht deels afhankelijk van de aanname dat frontier-labs rekenkracht onbeperkt zullen blijven opschalen — het commoditiseringsrisico neemt toe als opschaling minder productief is. Ten tweede kunnen kleinere labs en academische groepen de capaciteitsachterstand op frontier-labs mogelijk wegwerken via architectuurinnovatie in plaats van ruwe rekenkrachtuitgaven — waardoor het concurrentievoordeel van massief kapitaal vermindert. Ten derde wordt de tijdlijn voor AI die verschillende capaciteitsdrempels bereikt onzekerder, waardoor investeringen die afhankelijk zijn van specifieke capaciteitsmijlpalen een hoger risico dragen.
De VUB-studie “Benchmarks Saturate” (januari 2026) voegt een meetprobleem toe: naarmate modellen menselijk prestatieniveau benaderen op bestaande benchmarks, worden de benchmarks zelf onbetrouwbare maatstaven voor echte capaciteitsverbetering — waardoor het moeilijker wordt om vooruitgang onafhankelijk bij te houden.
Bewijs & Voorbeelden
- Hooker: “De academische wereld is gemarginaliseerd van betekenisvolle deelname aan AI-vooruitgang en industrielabs zijn gestopt met publiceren” — wat ondoorzichtigheid creëert precies op het moment dat meting het meest nodig is (
ssrn-5877662 (1).pdf) - Hookers these: opschaling heeft geleid tot een “massale kapitaalstroom naar industrielaboratoria” en heeft “de cultuur van wetenschapsbeoefening fundamenteel hertekend” — maar de schalingformule verandert, en “sleutelverstoring staat te wachten” (
ssrn-5877662 (1).pdf) - CSET-workshop (juli 2025): experts zijn het oneens over of toenemende automatisering van AI-onderzoek de AI-vooruitgang zal versnellen of zal doen stagneren; “nieuwe gegevens over hoe automatisering van AI-onderzoek in de praktijk vordert, kunnen onvoldoende zijn om conflicterende perspectieven te beslechten” — de verschillende kampen maken verschillende aannames die leiden tot verschillende interpretaties van hetzelfde bewijs (
CSET-When-AI-Builds-AI.pdf) - CSET: bestaande benchmark-evaluaties zijn “onvoldoende voor het meten, begrijpen en voorspellen van het traject van geautomatiseerd AI-onderzoek” — de meetinfrastructuur bestaat niet om de voortgang betrouwbaar bij te houden (
CSET-When-AI-Builds-AI.pdf) - VUB “Benchmarks Saturate” (jan. 2026): wanneer modellen menselijk prestatieniveau overtreffen, verliezen door mensen beoordeelde benchmarks discriminerend vermogen — de beoordelaar kan niet langer onderscheid maken tussen modellen die beter zijn dan zijzelf (
2601.19532v1.pdf) - Epoch AI schat dat effectieve rekenkracht voor het trainen van AI-modellen jaarlijks 10x stijgt, maar Jones (2026) merkt op dat de capaciteit per rekeneenheid ook met een vergelijkbaar tempo stijgt — waardoor ruwe rekenkrachtuitgaven een minder duidelijk signaal zijn voor frontier-voordeel (
AIandEconomicFuture.pdf)
Spanningen & Openstaande Vragen
- De onzekerheid over “intelligentie-explosie”: CSET documenteert dat scenario’s van “intelligentie-explosie” — waarbij AI zichzelf snel verbetert — noch bevestigd noch uitgesloten worden door huidig bewijs. Ze zijn laagwaarschijnlijk maar niet verwaarloosbaar, en moeilijk vooraf te detecteren. Dit is het staartrisico dat de meeste investeringsanalyses niet adequaat inprijzen.
- Meetondoorzichtigheid: Als opschaling onzeker is EN benchmarks verzadigen EN industrielabs gestopt zijn met publiceren, dan prijzen investeerders AI-capaciteitstrajecten met zeer beperkte openbare informatie. Dit creëert zowel risico (prijzen kunnen sterk afwijken) als kans (informatievoordelen zijn mogelijk voor degenen die het dichtst bij frontier-labs staan).
- Niet-schaling alternatieven: Hooker stelt dat “interessantere hefbomen van vooruitgang” buiten opschaling opkomen — maar specificeert niet welke. CSET merkt op dat architectuurinnovatie, data-efficiëntie en verbeterde trainingsmethoden kandidaten zijn. Investeringen die profiteren van architectuurverbetering in plaats van ruwe schaal zijn mogelijk beter gepositioneerd als Hooker gelijk heeft.
- Concurrentie-implicaties voor open-source: Als schalingvoordelen afnemen, kunnen open-source modellen (die profiteren van openbaar gedeelde architectuurinnovatie) de kloof met proprietary frontier-modellen sneller dichten. Dit zou LLM-commoditisering versnellen.
Gerelateerde Concepten
LLM Commoditization · AI Investment Thesis Capex and Returns · Recursive Self-Improvement and AI R&D · AI Capability Measurement · The Scaling Myth Is Finally Cracking