Tokenmaxxing At The End Of The World
This is probably fine.
Reports show that 90% of software developers use AI in their day-to-day work. The vast majority use large, closed models from frontier providers. However, an increasing amount of total tokens produced are coming from smaller, open-weight models like DeepSeek and Qwen. While it takes ungodly industrial hardware to run the latest ChatGPT model, a home user with an AMD RX 6950 XT designed to play Red Dead Redemption can also run Qwen3.8:27b at forty tokens per second. To do so will require about 500W of constant electrical supply for the duration of the generation process. That means that 24 hours of constantly cranking out tokens would cost 12kWh, about four times what it takes to power a typical England flat. All that to produce ~3.5 million tokens. With this context, we can only imagine the damage of billions of tokens per day of the best models that the U.S. domestic AI arms race has produced to date. All for the price of a nice lunch. It is both a deal any entrepreneur would be silly to pass up, and a perverse incentive structure indicative of ironically the same kind of misalignment that "AGI" prophets warn us about.
Relation To Production
The main question with any mass technological deployment is, "Is what you are doing with it worth the cost?" A rational society would ask this about automobile production and agricultural production as much as data center production. This is not a whataboutism; it is my contention that we would likely see a >90% reduction in the data center deployment rate if we were serious about our so-called "human values."
Are there "economically valid" uses for AI? No, just as there are not "economically valid" uses for anything. Validity is not derived by descriptive economic calculations. It is derived first and foremost by values. Ah, but now another question arises just as quickly as the first. Whose values? Yours? Mine? The president's? Well, comrade, you'll be glad to hear that we spent hundreds of years asking this question. Three of the main answers we came up with were:
- proprietors and the price levels they agree on (markets)
- all impacted parties through conversation and collective commitment to the decisions we make together (democracy)
- the people who seem to really know the most about the thing and its first- and second-order effects (technocracy)-
These answers are largely in tension with each other, and the regime over production generally fuses them together according to how they wish to enact their will through society. Markets produce a stochastically self-organized informational-logistical system that can be managed by the owning class via their employed technocratic elite. This elite class also maintains a heavy influence in the halls of democracy, fiddling the knobs until consent is sufficiently manufactured and the whole thing can chug along in the direction they choose without boiling over into a revolution. Right now, AI capital expenditure is decided mostly by the capital owners themselves via the mobilization of massive amounts of debt. This debt is rooted in institutions and networks that were built over decades, a kind of primitive accumulation of the financial ecosystem by technocratic cartels.
It's important to note that this debt and its mobilization represent an actual opportunity cost for any other purpose that it could have been mobilized for. Hospitals. Medicine. Schools. Solar panels. Above all, anything that requires a chip. We've seen RAM, GPUs, and tech generally skyrocket in price. This has had direct impacts on the affordability of cars, phones, televisions, medical equipment, you name it. We are living through a crisis that is partly the result of ecological breakdown, and partly intensified by a bourgeois scramble to accumulate as much of the crumbles as they can on the way down.
I think it's safe to say that we have given the technocracy and the capitalist class enough time at the wheel of the the forces of production. It's time that we seized it and made more rational choices.
Let's Get Dialectical
Thesis (Altman Hypothesis)
LLMs are the vanguard of intelligence. Soon there will just be token taps for LLMs and nothing else. LLMs will cure cancer and solve climate change in a matter of years. The only real question is if they will need us, and if they don't, what will they do to us?
Antithesis (LLM Critic)
The reintroduction of classification ML such as Jev as groundbreaking are proof of LLM-induced amnesia. We are treading the very water that we are wasting on these technologies. Classical machine learning, as well as emerging architectures like JEPA and Jev, are able to perform real business functions using orders of magnitude smaller compute resources. LLMs are useless when it comes to basically anything we could realistically need AI for. They require too many resources and produce too many errors. They're a fad that is on the way out as soon as this bubble bursts and people have to pay the real price per token.
Synthesis
LLMs are part of a historical process of refining how society wields technology toward its goals, and in that process there has never been a permanent winner because innovation is endemic to the process. We are years into an era of over-valuation, a problem that has grown exponentially due to LLM prominence. JEPA and Jev prove that LLMs aren't the best tool for everything, and small, open-weight models prove that you don't need massive data center roll-outs to reap the benefits of LLMs any more than you need 10 acres in order to enjoy fresh produce. Even as the billionaire industry leaders try to steer the entire economy at their whims, innovation and dissent at the fringes keeps moving toward the center and taking huge swaths of the territory that the billionaires believe to be their birthright. A rational society would be penalizing the misuse of AI as it does any other tool, while also legislating data provenance laws and economic policies that fundamentally change the incentive structure these companies currently operate on. Even if the bubble were to burst tomorrow, small open-weight models would still have the kind of unit economics that make sense for tasks like coding. If we give up this fight for autonomy in this arena, then the billionaires will successfully ban foreign models, re-establish their monopoly in the space, and continue to drive a larger portion of the economy into their capital expenditure on data centers. The footprint of their inevitable downfall will be so enormous that it makes 2008 look like a disappointing Christmas bonus.
The Takeaway
The dichotomy isn't world model vs. language model - the language model has a version of a world model - with different gaps then our own because the dataset is fundamentally different. Humanity runs on a linguistic-representational engine layered on top of a prediction-action-observation dataset. The former is an emergent property of the latter. The world that can be circumscribed by cognition may in absolute terms be approachable via language alone. But it is not the most efficient way to do so, and probably by orders of magnitude, with likely very long term implications for the embodied competence that drives actual industrial innovation. This is what JEPA demonstrates at the level of automation. Better to calculate loss of a prediction against its result in a simpler, more proximate latent space then to defer to a higher representational space simply because it produces a chain-of-thought that we can gawk at. As Jev shows us, it is also better to produce a probability spread than a meandering autoregression of next-token predictions. When we are obliterating decades of renewable energy gains generating a constant stream of blather just to get a bot to turn a screw, we are pushing the limits of, "Worse is better." Still, it is important to note that just as linguistically derived insights propagate downward (i.e. a new tool progressing from present-at-hand to ready-to-hand in the body-mind of the apprentice), we will likely see more and more surface area aising among a heterogenous ecosystem of speaking and non-speaking artificial intelligences. The task at hand is to actually center those human values we claim to be so worried about losing to "AGI."