
š„ The Compute Crunch: Why GPU ShortageāNot Model PerformanceāIs the Real AI Bottleneck
The AI market narrative has shifted dramatically in recent weeks. Following the release of DeepSeek's Kimi K3 model, many analysts rushed to declare the AI bubble overāarguing that Chinese open-source models could deliver comparable intelligence at a fraction of the cost charged by frontier labs in San Francisco.
But according to Leo Fan, founder of compute network Seismic and Cornell PhD researcher specializing in zero-knowledge proofs and AI infrastructure, the real story isn't about model performance catching up. It's about a critical bottleneck that few are properly analyzing: the severe shortage of GPU cards and memory infrastructure required to actually deliver inference at scale.
š§ The Performance Gap Reality
Despite the market excitement around Chinese open-source models, Fan offers a measured assessment of where these models actually stand relative to Western frontier labs:
"Although Kimi or DeepSeek are catching up, there is still some gap thereāprobably like several months gapāto catch up with the latest model from say OpenAI or Anthropic."
Fan's assessment suggests that Kimi K3 performs roughly on par with Claude Opus 4.8, but not yet at the level of Claude Opus 5. While open-source models can be engineered to excel at specific benchmark tasks, they still lag in handling diverse, general-purpose workloads.
The key insight: the bubble isn't bursting because the technology gap still exists, even if it's narrowing faster than many expected.
ā” The Real Bottleneck: GPU Cards and Memory
When Kimi K3 launched, the platform experienced such overwhelming demand that servers were immediately maxed out and the service had to be temporarily shut down. This wasn't a software problemāit was a fundamental infrastructure constraint.
Fan explains the dual shortage creating the bottleneck:
- GPU Card Shortage: Even with access to open-source model weights, organizations cannot obtain sufficient GPU cards to perform inference at scale
- Memory Constraints: Large models like Kimi K3, with its 2.5 trillion parameters, require massive memory capacity to store parameters locally for inferenceāmemory that simply isn't available in sufficient quantities
This creates a paradox: even as model weights become freely available through open-source releases, the compute infrastructure to actually run these models remains severely constrained.
"We cannot get enough GPU cards to do the inference for our customers, which will cause this shortage... In addition to the GPU card, you also need to have enough memory to store the models locally to do the inference."
š° Why Inference Costs Remain High
The prevailing market assumption suggested that open-source models would crater inference pricing. The logic seemed sound: if a model delivering 95% of frontier performance costs 50-80% less to access, why would anyone pay premium rates?
But Fan's analysis reveals why pricing hasn't collapsed as expected:
The native Kimi model's token pricing remains on the same level as Claude Opus 4.8, approaching Claude Opus 5 pricing, despite being open source. This isn't due to artificial price gougingāit's because the cost structure is dominated by scarce hardware resources, not model licensing.
Even with free model weights, organizations face:
- Severe GPU card shortages limiting inference capacity
- Extremely long response times due to infrastructure constraints
- High memory costs for storing multi-trillion parameter models
- Geographic distribution requirements to minimize latency for end users
The economics are clear: compute scarcity, not model intelligence, is driving pricing.
š§ Engineering Solutions: Older Hardware, Smarter Deployment
While the shortage of cutting-edge hardware persists, Fan points to significant engineering opportunities to extract more value from existing infrastructure:
For inference (as opposed to training), older generation hardware can deliver equivalent results with proper engineering. Fan provides a specific example:
"If you want to run Llama v4 Pro, which is probably the strongest open-source model before Kimi, you can use 4090 GPU cards... you can use probably 16 4090 GPU cards to deliver the same results from eight B200 GPU cards."
This creates a critical distinction in the market:
- Training: Still requires the latest GPU cards (H100s, B200s) to achieve optimal results and competitive training times
- Inference: Can be performed effectively on previous-generation hardware with proper engineering, creating arbitrage opportunities
This gap between training and inference requirements opens strategic opportunities for compute networks and infrastructure providers who can optimize deployment across heterogeneous hardware.
š The Semiconductor Supply Chain: What Comes Next?
The market has watched capital flow through distinct phases of the AI infrastructure build-out:
- Phase 1: Hyperscalers (Google, Meta, Microsoft) raised capital through bond offerings
- Phase 2: Capital flowed into memory manufacturers and semiconductor companies as hyperscalers deployed funds
- Phase 3 (Current): The question becomesāwhere do memory and semiconductor companies allocate capital next?
Fan clarifies that raw materials aren't the primary constraint. Silicon extraction from sand remains straightforward and scalable. Instead, the bottleneck is precision manufacturing at advanced process nodes:
"Right now we are moving from 5 nanometer to 3 nanometer or even 2 nanometers... it's about how you can have very advanced equipment on the chip to make sure you can have very high precision to cut the chip... and still want to get a very high yield, probably 90% positive results."
The constraint isn't materialsāit's advanced lithography equipment and process expertise to maintain high yields at smaller process nodes. This explains why ASML, the Dutch company with a monopoly on extreme ultraviolet (EUV) lithography systems, has become such a critical chokepoint in the global semiconductor supply chain.
š US vs. China: The Infrastructure Race
While Fan declined to make direct predictions about which nation is better positioned for AI rollout, he offered insights into the strategic differences in approach:
United States: Focused on closed-source models (OpenAI, Anthropic) that attract users and capital in a flywheel effect, with emphasis on access control and commercial scaling
China: Pursuing aggressive open-source releases (DeepSeek, Kimi) that force frontier labs to compete on research innovation rather than just marketing and user lock-in
Fan suggests this competition will ultimately benefit end users, forcing both ecosystems to prioritize research advancement over defensive moats.
On infrastructure preparation, China has approved numerous nuclear power plants while the US regulatory process has been slower. However, both nations face similar constraints around advanced semiconductor manufacturing and access to cutting-edge lithography equipment.
š The Financialization of Compute: A New Derivatives Market
Perhaps the most forward-looking discussion centered on compute emerging as a financialized asset class. Recent developments include:
- Kalshi launching daily compute products allowing speculation on hourly Nvidia compute pricing
- Emerging compute perpetuals enabling hedging strategies
- The concept of "DePIN" (Decentralized Physical Infrastructure Networks) enabling hardware owners to optimize yield across different computational workloads
Fan introduces the concept of "ComputeFi"āthe financialization and commercialization of hardware resources:
"ComputeFi means you can financialize or commercialize your hardware... hardware owners can use native software to program their hardware to suit different computing demands."
The practical application: as demand shifts between zero-knowledge proof generation and AI inference, compute providers can reprogram the same hardware to capture emerging opportunities. For example, as Layer 2 activity declined, many compute providers shifted from ZK proof generation to AI token generationāmaintaining profitability with the same physical infrastructure.
šÆ Mining the New Frontier: Yield Opportunities
For individuals wondering if they've missed the "mining" opportunity in AI compute, Fan offers practical guidance based on hardware capabilities:
Consumer Hardware (MacBooks, etc.):
- Best suited for ZK proof generation and verification
- Can run only small models (0.5 to 1 billion parameters maximum)
- Limited profitability for AI token generation
High-Performance GPUs (H100, A100, B200):
- Optimized for AI training and inference rather than ZK generation
- Can generate AI tokens at scale with strong profit potential
- Requires matching hardware capabilities to appropriate workloads
Seismic's upcoming "InferBench" tool will help hardware owners test their systems and identify which models deliver optimal returns given their specific hardware configuration.
š Long-Term Outlook: Exponential Compute Demand
Despite periodic narratives about AI bubble collapse, Fan maintains conviction in exponential growth for compute demand:
"I see exponential growth of the market for demand of computes... Models are growing largerāto train these larger models you need more powerful, cheaper cards to train them, and after training you still need very powerful cards to do inference for these very powerful models."
Fan's timeline: At minimum over the next one to two years, expect no decline in computing demand, driven entirely by AI computational requirements.
When asked whether compute derivatives could rival traditional commodity markets like oil, gold, and forex, Fan responded affirmatively, pointing to the exponential trajectory as models scale and inference demands multiply.
š® The Bottom Line
The market narrative suggesting open-source Chinese models have commoditized AI inference misses the fundamental constraint: it's not about access to model weights, it's about access to the physical infrastructure required to run those models.
Key takeaways for investors and builders:
- The performance gap between open-source and frontier models is narrowing but still exists (several months behind)
- GPU and memory shortagesānot model licensingāare the primary driver of inference costs
- Engineering optimization can extract significantly more value from older hardware for inference workloads
- The semiconductor bottleneck is advanced lithography precision, not raw materials
- Compute is emerging as a financialized asset class with derivative markets comparable to traditional commodities
- Demand for compute infrastructure shows no signs of declining over the next 1-2 years minimum
As capital continues flowing through the AI infrastructure stackāfrom hyperscalers to chip manufacturers to advanced lithography equipmentāthe real opportunity isn't in betting on which model will win, but in understanding and positioning around the physical constraints that will define the pace of AI deployment.
The AI race isn't being won in research labsāit's being won in semiconductor fabs and data centers with sufficient power, cooling, and memory infrastructure to actually run the models at scale.
More from TheRollupCo

Inside Robinhood Chain: How 105M Transactions in Three Weeks Signals the Next Wa
Three weeks into mainnet, Robinhood Chain has emerged as one of the most successful blockchain launches in recent memory...

Inside Wisdom Tree's Tokenization Strategy: Building the Future of 24/7 Asset Ma
š The Convergence Thesis: Why Tokenization Is No Longer OptionalThe financial industry is witnessing a structural shift...

Irresponsibly Long Crypto: Why the Four-Year Cycle Isn't Dead and What's Really
š Portfolio Positioning: Still All-In on CryptoDespite widespread institutional skepticism, veteran crypto investors re...

The Convergence of Intents, AI, and Privacy: Inside Near's Vision for Agentic Co
š” The Big Picture: Privacy, Intelligence, and Commerce Are MergingThe evolution of crypto infrastructure is no longer j...

Ethereum's License to Win: Inside the Institutional Supercycle Driving Trillions
The institutional crypto supercycle isn't coming ā it's already here. And while market sentiment remains subdued, the un...

Hyperliquid, Maple, and Robin Hood Chain Lead the Revenue Revolution
š Turnaround Tuesday: The Market Rebounds After Yesterday's PanicAfter yesterday's brief selloff, markets delivered a t...