Y Combinator5 min read

YC Paper Club explores ditching GPUs for light, brain chips, and gradient-free training

Four researchers lay out why backprop, transformers and silicon GPUs may not survive the next decade of AI compute.

AI summary of “What If We Stopped Using GPUs? | YC Paper Club”

Key takeaways

  • Francois Chaubard argues the brain does not do backpropagation and is building a gradient-free optimizer called SOMA.
  • Ilker Oguz describes an EPFL/Google project running diffusion image generation through passive light propagation instead of a GPU.
  • Alok Vasudev says neuromorphic computing in 2026 is "still squarely in R&D land" despite decades of brain-inspired chip research.
  • Sean Cole explains how his team trained human brain cells on a chip to play Doom using reinforcement learning via electrical stimulation.
  • Speakers disagree on how far optical and biological compute can scale, with no firm timelines or figures given.

Why alternative compute, according to Francois Chaubard

Opening the session, Francois Chaubard, who has worked in deep learning since 2012, traced how Nvidia's hardware priorities shifted from raw flops-per-joule efficiency for CNNs toward memory capacity and bandwidth once transformers arrived, driven by attention's quadratic scaling. He said gigaflops-per-joule gains have "petered out" over the last two years, which he takes as a sign the field needs something besides GPUs and backprop.

He argued the brain is "massively feed forward" and cannot be doing backpropagation, since that would require neurons to fire backward through the same synapses, which he says they do not do. He pointed to cortical columns with massive inhibition as a structure allowing independent, localized learning, and framed optics, neuromorphics and biological neurons as three substrates worth exploring instead.

Beyond backprop: SOMA and zero-order optimization

Chaubard described spending his PhD testing zero-order (gradient-free) optimizers, settling on a method called SPSA, which perturbs model parameters in random directions and scales updates by the resulting finite difference rather than computing gradients. He said this approach can escape local minima that trap backprop on non-Lipschitz loss landscapes.

His current project, SOMA ("sharded optimization mixture of assemblies"), clusters a dataset like Common Crawl and trains small experts (he mentioned a roughly 41K-parameter LSTM) on separate GPUs worldwide, then routes between them at test time. He said gradient-estimate error from zero-order methods worsens as model size grows, making large monolithic models infeasible this way, but sharding keeps each expert's noise capped, requiring as few as 64 perturbations per step. He plans to submit the work to ICLR.

Computing with light, according to Ilker Oguz

Ilker Oguz said photons have roughly 10,000 times lower loss and 10,000 times larger bandwidth than electrons, which is why nearly all AI data-center and intercontinental transmission already uses optics. He argued optical matrix operations can scale energy use linearly with input/output size rather than quadratically, since weights can passively modulate light without active energy.

He explained that optical computing isn't yet widespread because converting between digital/electronic and optical/analog domains is costly, calibration requires active energy, and nonlinear activations are hard to implement optically. His EPFL/Google project built a diffusion-based image generator (for small datasets like MNIST and Fashion-MNIST) where fixed optical layers perform denoising steps via light propagation, reusing the same physical weights across hundreds of steps. He reported the system showed a power-law scaling behavior similar to digital neural networks and a "significant energy performance advantage" for generating a single image versus a comparable GPU model, though he flagged that this assumed fixed, micro-fabricated weights rather than the prototype's spatial light modulator setup.

"I believe the opportunity in optical computing is through designing hardware and algorithms together."
— Ilker Oguz

Audience questions on optical computing's limits

In the Q&A, an audience member who had worked at Lightmatter said the company pivoted away from multispectral compute toward HBM interconnects after finding ADC/DAC resolution issues hurt prediction accuracy in the field. Oguz agreed that advantages erode when large data volumes must be transmitted relative to parameter count, and said application-specific designs make more sense than general-purpose optical hardware for now.

On scaling, Oguz said manufacturing can reuse much of the existing electronics tooling, but larger light wavelengths mean larger feature sizes, and finding cheap ways to implement nonlinearity remains the main obstacle to building deep optical networks. On optical storage, he said holographic memories and phase-change materials have been tried but none yet rival electronic memory for cheap, fast read/write.

Neuromorphic computing's status, according to Alok Vasudev

Alok Vasudev, a partner at a $1.5 billion venture fund, said neuromorphic computing borrows organizing principles from the brain — intertwined memory and computation, spike-based communication, and structural plasticity — but that modern AI (backprop, transformer attention, training at scale) has already diverged from biological mechanisms. He said the brain represents an "ultimate model" of hardware-software co-design.

He grouped current neuromorphic efforts into three categories: energy-efficiency-driven work co-mingling memory and compute (citing startup Dmatrix and IBM), edge/sensing-driven work on spiking networks and on-device learning (citing Intel's research chip flying a drone), and exploration of new exotic devices like resistive memories and memristors. His overall assessment: — Alok Vasudev. Asked for the most promising direction, he pointed to memory-compute co-packaging efforts like Dmatrix as closer to market, calling more brain-faithful approaches still early-stage and dependent on founders able to raise enough capital to pursue them.

Audience debate on noise, determinism and materials

One attendee asked about the role of noise and reproducibility in computation, noting the brain doesn't aim for bit-level reproducibility the way digital silicon does. Vasudev pointed to thermodynamic computing and historical Boltzmann machine research as attempts to put noise to productive use. Chaubard added that zero-order optimization requires reproducible noise from a fixed seed to remain learnable, and said a neuroscientist he consulted, Shaul Druckmann, told him this reproducibility is "really the issue" for biological plausibility of such methods.

Other questions touched on silicon versus carbon substrates (Vasudev said the field isn't overly tied to one material and cited an unrelated team working on diamond substrates) and on packaging efficiency, where Vasudev said he personally expects optics to address interconnect-driven energy costs tied to data-center capacitance.

Teaching brain cells to play Doom, according to Sean Cole

Sean Cole, who worked with Cortical Labs, described treating cultured human brain cells as an input-output dynamical system rather than a matrix-multiplication device. Game state (health, ammo, screen image) is encoded into electrical stimuli delivered across a 59-channel multi-electrode array; resulting spikes are decoded into in-game actions. He contrasted this with Cortical Labs' earlier Pong demonstration, where hand-designed encoding/decoding sufficed for a simple two-action game, saying Doom's larger action space (54 joint actions) made hand-mapping across only 59 channels impractical.

The team used a PPO-based reinforcement learning loop, training an encoder via stochastic (beta-sampled) stimulation since gradients cannot propagate through the electrode array, and a deliberately undersized linear decoder to prevent the silicon system from "cheating" by learning to play the game itself rather than relying on the cells. Feedback was delivered as synchronous stimulation for good actions and asynchronous stimulation for bad ones, scaled by a critic's prediction error ("surprise"), drawing on Karl Friston's free energy principle. Cole said the system's two components never reach equilibrium, with entropy actively penalized rather than rewarded, unlike standard RL practice, because the cells themselves supply sufficient exploratory entropy.

"I think that human intelligence is not optimal for compute. We're optimized for survival... but we're not optimized for compute."
— Sean Cole

Asked how scalable the approach is, Cole said the two open problems are reaching frontier-level intelligence on biological cells and then distributing that intelligence to many users, adding plainly that on the second problem, "we fundamentally don't know how to do this yet."

Written by AI from the video's transcript. It can compress, misattribute or miss context — the original video is the source. Not investment advice.

More from Y Combinator

Y Combinator

Founders say general-purpose LLMs, not robot-specific models, may crack robotics

Waddle Labs' Vincent and Hamming, and RoboCurve's Jay, argue frontier LLMs like Astra can control robots via code and tool calls with little robot-specific training.

Y Combinator

Eight Proven Tactics to Transform Your Outbound Sales From Zero to Revenue

For early-stage founders staring at reply rates hovering near zero, the outbound sales challenge can feel insurmountable. Yet the path from cold outreach to closed revenue doesn't require expensive tools or elaborate automation—it demands a methodical, founder-led approach that treats each interaction as a learning opportunity.

Y Combinator

From Scaffolding to Self-Improvement: How Agent Harnesses Became the Unlock for AGI-Level Performance

Just a month ago, a Reddit thread dismissed prompt engineering as "not a research problem." Another commenter called it "scaffolding" unworthy of serious academic attention. Yet agent harnesses—the frameworks that orchestrate how models interact with tools, memory, and execution environments—have delivered 18% performance improvements between iterations and have become the difference between solving ARC-AGI and not.

Y Combinator

The Great Migration: Why 85% of Fortune 500s Are Running Open AI Models

A quiet revolution is underway in enterprise AI adoption. According to Jeffrey Morgan, co-founder and CEO of Ollama—the platform used by 9 million developers and 85% of the Fortune 500 —businesses are rapidly pivoting from closed, proprietary models to open-source alternatives. The driver? It's not just cost. It's control, customization, and increasingly, geopolitical pragmatism .

Y Combinator

The Real Story of YC at 21 Years: What Makes Founders Formidable

Y Combinator recently completed its 47th batch — marking 21 years of the world's most influential startup accelerator. Despite persistent claims that "YC has jumped the shark," the fundamentals of building successful startups remain remarkably consistent across two decades.

Y Combinator

The $5 Flight: How Hart Aerospace Built the World's Largest Electric Aircraft

In a remote hangar in Platsburg, New York, the world's largest electric aircraft lifted off for the first time—a moment seven years in the making that began with a 3D printed model small enough to hold in one hand. What flew that day wasn't just an aircraft. It was a direct challenge to the fundamental economics of short-haul aviation, and a bet that the jet engine's 40-year reign over regional routes is coming to an end.