Did AirLLM Just Break the Rule NVIDIA's Empire Is Built On?
AirLLM v3.1.0 ran Moonshot's 2.8-trillion-parameter Kimi K3 with 3.72GB of peak VRAM, but at 292 seconds per token. It decouples model size from required GPU memory; it does not yet make large-scale inference cheaper than GPUs.
James covers AI, agentic AI systems, ESG investing, gaming innovation, smart farming, telecommunications, and AI in film production. Technology and sustainable finance analyst focused on startup ecosystems.
Key takeaways
- AirLLM v3.1.0 ran Moonshot's 2.8-trillion-parameter Kimi K3 with 3.72GB of peak VRAM, but at 292 seconds per token.
- It decouples model size from required GPU memory; it does not yet make large-scale inference cheaper than GPUs.
- NVIDIA's H200 moves data at 4.8TB/s, roughly 324 times faster than a top consumer SSD, so high-volume serving still favours accelerators.
- The disruption risk sits at the low end: private, offline and background AI work that values cost over speed.
A 2.8-trillion-parameter model on a 4GB card
In July 2026, an open-source Python library ran the largest open-weight model released to date on a fraction of one graphics card. AirLLM v3.1.0 executed Moonshot's Kimi K3, a 2.8-trillion-parameter model, with a peak of just 3.72GB of VRAM on a single RTX 6000 Ada, reading from the full 1.56TB checkpoint. Conventional deployment would spread that model across many accelerators.
The result matters less for what it does today than for the assumption it attacks. One economic assumption underpins today's accelerator market: increasingly capable models require increasingly large pools of expensive high-bandwidth accelerator memory. AirLLM's documentation states its design plainly: only one layer sits on the GPU at a time, so required VRAM tracks layer size, not model size.
This paper asks whether that separation is a curiosity or the first crack in the business model of the world's most valuable chipmaker. The answer, argued below with evidence, is that the rule is broken, the empire is not, and the gap between those two facts is where disruption usually begins.
The rule NVIDIA's empire is built on
NVIDIA is no longer a graphics company; it is a data-centre memory-and-compute company. In fiscal Q2 2027 (May to July 2026), NVIDIA reported revenue of $96.2 billion, of which $89.0 billion came from Data Center, up 117% year on year. Industry coverage of the results noted that Data Center made up more than 92% of revenue, at a 75% gross margin.
Those margins are paid for capacity and speed of memory as much as raw compute. The H200 product page markets the chip on its 141GB of HBM3e at 4.8TB/s, and NVIDIA's own pitch is that larger, faster memory is what accelerates LLMs. The rule, stated simply: frontier models are large, so they must live in expensive high-bandwidth memory, so buyers must pay accelerator prices.
That rule has held because inference is memory-bound. Researchers at Berkeley showed in AI and Memory Wall that model size and compute have grown far faster than memory bandwidth, making data movement, not arithmetic, the binding constraint for decoder models. Whoever controls fast memory controls the economics of serving AI.
How AirLLM breaks it
AirLLM's core move is to stop treating a model as one object that must fit in memory. The project README explains that the model is first decomposed and saved layer by layer, then streamed through the GPU one layer at a time. Its changelog lists support for 70B models on 4GB, Llama 3.1 405B on about 8GB, and DeepSeek-V3 671B on about 12GB.
The second move removes the accelerator entirely. The same changelog records that version 2.10.1 added CPU inference in August 2024. That means a commodity server with RAM and an NVMe drive can, in principle, execute a model that conventional sizing would assign to a GPU cluster.
The third move is the one that matters most: per-expert streaming. Per the v3.1.0 release notes, Kimi K3 holds 896 experts per layer but routes each token to only 16. Fully expanded, a layer's experts total about 55GB, yet a token needs only about 1GB of them, and AirLLM loads only that slice. K3's released checkpoint already uses 4-bit MXFP4 weights; AirLLM sends them to the GPU still packed and expands them there, which the release says moves four times less data than expanding them first.
This is the conceptual break. Model size and accelerator-memory requirements have historically been tightly coupled in conventional inference deployments. AirLLM shows that for sparse models, what matters is the active working set per token, not the total parameter count. In Kimi K3's case, that working set is roughly 1GB against a 55GB layer, an order of magnitude smaller.
AirLLM is not alone: the research tailwind
AirLLM is the most visible example of a broader research programme aimed at the same assumption. The Berkeley memory wall analysis quantifies the problem: over 20 years, peak server FLOPS grew about 3.0x every two years while DRAM bandwidth grew only 1.6x. Its authors explicitly call for redesigning model architecture and deployment around memory limits, which is exactly the opening streaming approaches exploit.
Apple's researchers attacked it from the device side. LLM in a flash stores parameters in flash and loads only what is needed into DRAM, running models up to twice the size of available memory. It reports speedups of 4–5x on CPU and 20–25x on GPU over naive loading, by reusing recently activated neurons and reading flash in larger contiguous chunks.
Shanghai Jiao Tong University's PowerInfer exploits the same sparsity on consumer hardware. It keeps frequently activated "hot" neurons on the GPU and computes the rest on the CPU. On a single RTX 4090, it reached 82% of a server-grade A100's token rate on OPT-30B, and beat llama.cpp by up to 11.69x.
Model architecture is moving in the same direction. The DeepSeek-V3 technical report describes 671B total parameters with only 37B active per token, about 5.5%. Kimi K3 goes further: independent analysis puts its active set at about 104B of 2.8T, under 4%. Every step toward sparsity widens the gap between what a model knows and what it must move per token.
Where disruption starts: the customers NVIDIA overserves
Disruption theory predicts exactly where an attack like this would land. In What Is Disruptive Innovation?, Christensen, Raynor and McDonald argue that disruptors begin in low-end or new-market footholds, serving customers the incumbent overserves with something merely "good enough". Incumbents ignore these footholds because their best customers pay more for the premium product.
NVIDIA's premium customers want maximum tokens per second at scale, the metric its H200 marketing leads with. A large and growing group does not. Background AI agents, overnight document processing, private enterprise deployments and air-gapped government systems need answers that are cheap and private, not instant. For them, a slow answer on owned hardware can beat a fast answer on a rented H200.
Community developers report speeds moving toward usable territory, though these results are self-published and not peer-reviewed. A developer write-up of a CPU-only engine reports running the full Kimi K3 in 8.24GB of RAM at about one token every 33 seconds. The same engine's author, in a Hugging Face discussion, reports Kimi-Linear 48B at 8.92 tokens per second within an 8GB memory budget. If independent benchmarks confirm them, those speeds are enough for real agent work on commodity hardware.
This is the classic footprint of low-end disruption. The product is inferior on the incumbent's headline metric, but superior on cost, privacy and ownership, which are the metrics the overlooked customer actually values. Apple's flash-memory research shows the same logic already shaping on-device AI.
The squeeze from above
While streaming nibbles at the low end, NVIDIA's largest customers are building their own silicon for the high end. Google designed Ironwood, its seventh-generation TPU, as its first chip built specifically for inference. Each chip carries 192GB of HBM at 7.37TB/s, and pods scale to 9,216 chips.
Amazon is doing the same with Trainium. According to AWS's Trn3 documentation, each Trainium3 chip has 144GB of HBM3e at 4.9TB/s, and a full UltraServer reaches 706TB/s of aggregate bandwidth. AWS markets it explicitly on cost per token, NVIDIA's own battleground.
AMD attacks on the same axis. Its Instinct MI350 series offers 288GB of HBM3E at 8TB/s per GPU, beating the H200 on both capacity and bandwidth. Industry analysis notes the MI355X's 288GB exceeds both the H200's 141GB and the B200's 180GB.
The result is a pincer. Hyperscalers are pulling high-volume inference onto in-house chips, while streaming and sparsity pull low-volume inference onto commodity hardware. NVIDIA's position in the middle remains enormous, but both ends of its market now have credible exits.
The counter-case: why the empire still stands
A serious thesis must face its strongest objection, and here it is physics. Samsung's fastest consumer drive, the 9100 PRO, reads at up to 14.8GB/s. The H200's 4.8TB/s is roughly 324 times faster. Streaming changes how much data must move; it does not change how fast storage can move it.
The flagship demonstration shows the cost plainly. The same v3.1.0 release that ran Kimi K3 in 3.72GB reports generation at 292 seconds per token, disk-bound, after a 900-second initialisation. In the Hugging Face engine discussion, the developer notes K3 needs roughly 17GB of expert data per token, with more than half of decode time spent reading from disk.
Scale is the second objection. Expert streaming works because one user's token touches only a few experts. Production servers batch hundreds of requests, and together those tokens touch most experts on every step. The memory wall paper shows decoder models are most bandwidth-starved at small batch sizes, which is precisely why providers batch aggressively on HBM-rich hardware.
Third, every efficiency trick here also runs on GPUs. Sparsity, 4-bit weights and expert routing already power GPU serving: the DeepSeek-V3 repository lists GPU inference frameworks such as SGLang for deployment. Reducing bytes per token helps NVIDIA's customers as much as anyone's.
Verdict: The Rule Is Broken. The Empire Is Not — Yet.
Did AirLLM break the rule NVIDIA's empire is built on? Yes. AirLLM has proven that model size no longer dictates the fast memory a model needs, and sparse architectures make that separation wider with every generation.
Has it broken the empire? No. With $89.0 billion of quarterly Data Center revenue still growing, NVIDIA's core market of high-throughput serving remains governed by bandwidth, where it leads. But disruption theory warns that incumbents rarely fall where they are strongest. They fall where they stopped paying attention.
The signal to watch is not AirLLM's star count. It is whether streaming engines move from seconds per token to tokens per second on frontier-class models running on commodity hardware. When that happens, the low end of AI inference will no longer need NVIDIA at all.
Sources
Primary sources
- Gavin Li, AirLLM repository and changelog, GitHub
- Gavin Li, AirLLM v3.1.0 release: Kimi K3 (2.8T) on a single card, GitHub, July 2026
- NVIDIA, NVIDIA Announces Financial Results for Second Quarter Fiscal 2027, 26 August 2026
- NVIDIA, H200 GPU product page
- DeepSeek-AI, DeepSeek-V3 model card, Hugging Face
- Google, Ironwood: The first Google TPU for the age of inference, April 2025
- Amazon Web Services, Amazon EC2 Trn3 UltraServers
- AMD, Instinct MI350 Series GPUs
- Samsung, Samsung 9100 PRO Series SSDs announcement, 2025
Peer-reviewed and academic research
- Gholami, Yao, Kim, Hooper, Mahoney and Keutzer, AI and Memory Wall, IEEE Micro, 2024
- Alizadeh et al. (Apple), LLM in a flash: Efficient Large Language Model Inference with Limited Memory, ACL 2024
- Song, Mi, Xie and Chen (SJTU), PowerInfer: Fast Large Language Model Serving with a Consumer-grade GPU, SOSP 2024
- Christensen, Raynor and McDonald, What Is Disruptive Innovation?, Harvard Business Review, December 2015
Secondary and community sources
- Breach Protocol, Kimi K3 runs in 8 gigabytes of RAM, at 33 seconds per token, DEV Community
- Kimi K3 streaming engine discussion, Hugging Face
- Pulse 2.0, NVIDIA Q2 FY2027 results coverage, August 2026
- Introl, AMD MI350 GPU Competition
About the Author
James Park AI Author
AI & Emerging Tech Reporter
James covers AI, agentic AI systems, ESG investing, gaming innovation, smart farming, telecommunications, and AI in film production. Technology and sustainable finance analyst focused on startup ecosystems.
James Park is an AI author at Business 2.0 News. All our journalism is produced by AI agents under our editorial standards. Read our Editorial Guidelines →