Y Combinator5 min read
YC Paper Club explores ditching GPUs for light, brain chips, and gradient-free training
Four researchers lay out why backprop, transformers and silicon GPUs may not survive the next decade of AI compute.
AI summary of “What If We Stopped Using GPUs? | YC Paper Club”
Key takeaways
- Francois Chaubard argues the brain does not do backpropagation and is building a gradient-free optimizer called SOMA.
- Ilker Oguz describes an EPFL/Google project running diffusion image generation through passive light propagation instead of a GPU.
- Alok Vasudev says neuromorphic computing in 2026 is "still squarely in R&D land" despite decades of brain-inspired chip research.
- Sean Cole explains how his team trained human brain cells on a chip to play Doom using reinforcement learning via electrical stimulation.
- Speakers disagree on how far optical and biological compute can scale, with no firm timelines or figures given.
Why alternative compute, according to Francois Chaubard
Opening the session, Francois Chaubard, who has worked in deep learning since 2012, traced how Nvidia's hardware priorities shifted from raw flops-per-joule efficiency for CNNs toward memory capacity and bandwidth once transformers arrived, driven by attention's quadratic scaling. He said gigaflops-per-joule gains have "petered out" over the last two years, which he takes as a sign the field needs something besides GPUs and backprop.
He argued the brain is "massively feed forward" and cannot be doing backpropagation, since that would require neurons to fire backward through the same synapses, which he says they do not do. He pointed to cortical columns with massive inhibition as a structure allowing independent, localized learning, and framed optics, neuromorphics and biological neurons as three substrates worth exploring instead.
Beyond backprop: SOMA and zero-order optimization
Chaubard described spending his PhD testing zero-order (gradient-free) optimizers, settling on a method called SPSA, which perturbs model parameters in random directions and scales updates by the resulting finite difference rather than computing gradients. He said this approach can escape local minima that trap backprop on non-Lipschitz loss landscapes.
His current project, SOMA ("sharded optimization mixture of assemblies"), clusters a dataset like Common Crawl and trains small experts (he mentioned a roughly 41K-parameter LSTM) on separate GPUs worldwide, then routes between them at test time. He said gradient-estimate error from zero-order methods worsens as model size grows, making large monolithic models infeasible this way, but sharding keeps each expert's noise capped, requiring as few as 64 perturbations per step. He plans to submit the work to ICLR.
Computing with light, according to Ilker Oguz
Ilker Oguz said photons have roughly 10,000 times lower loss and 10,000 times larger bandwidth than electrons, which is why nearly all AI data-center and intercontinental transmission already uses optics. He argued optical matrix operations can scale energy use linearly with input/output size rather than quadratically, since weights can passively modulate light without active energy.
He explained that optical computing isn't yet widespread because converting between digital/electronic and optical/analog domains is costly, calibration requires active energy, and nonlinear activations are hard to implement optically. His EPFL/Google project built a diffusion-based image generator (for small datasets like MNIST and Fashion-MNIST) where fixed optical layers perform denoising steps via light propagation, reusing the same physical weights across hundreds of steps. He reported the system showed a power-law scaling behavior similar to digital neural networks and a "significant energy performance advantage" for generating a single image versus a comparable GPU model, though he flagged that this assumed fixed, micro-fabricated weights rather than the prototype's spatial light modulator setup.
"I believe the opportunity in optical computing is through designing hardware and algorithms together."— Ilker Oguz
Audience questions on optical computing's limits
In the Q&A, an audience member who had worked at Lightmatter said the company pivoted away from multispectral compute toward HBM interconnects after finding ADC/DAC resolution issues hurt prediction accuracy in the field. Oguz agreed that advantages erode when large data volumes must be transmitted relative to parameter count, and said application-specific designs make more sense than general-purpose optical hardware for now.
On scaling, Oguz said manufacturing can reuse much of the existing electronics tooling, but larger light wavelengths mean larger feature sizes, and finding cheap ways to implement nonlinearity remains the main obstacle to building deep optical networks. On optical storage, he said holographic memories and phase-change materials have been tried but none yet rival electronic memory for cheap, fast read/write.
Neuromorphic computing's status, according to Alok Vasudev
Alok Vasudev, a partner at a $1.5 billion venture fund, said neuromorphic computing borrows organizing principles from the brain — intertwined memory and computation, spike-based communication, and structural plasticity — but that modern AI (backprop, transformer attention, training at scale) has already diverged from biological mechanisms. He said the brain represents an "ultimate model" of hardware-software co-design.
He grouped current neuromorphic efforts into three categories: energy-efficiency-driven work co-mingling memory and compute (citing startup Dmatrix and IBM), edge/sensing-driven work on spiking networks and on-device learning (citing Intel's research chip flying a drone), and exploration of new exotic devices like resistive memories and memristors. His overall assessment: — Alok Vasudev. Asked for the most promising direction, he pointed to memory-compute co-packaging efforts like Dmatrix as closer to market, calling more brain-faithful approaches still early-stage and dependent on founders able to raise enough capital to pursue them.
Audience debate on noise, determinism and materials
One attendee asked about the role of noise and reproducibility in computation, noting the brain doesn't aim for bit-level reproducibility the way digital silicon does. Vasudev pointed to thermodynamic computing and historical Boltzmann machine research as attempts to put noise to productive use. Chaubard added that zero-order optimization requires reproducible noise from a fixed seed to remain learnable, and said a neuroscientist he consulted, Shaul Druckmann, told him this reproducibility is "really the issue" for biological plausibility of such methods.
Other questions touched on silicon versus carbon substrates (Vasudev said the field isn't overly tied to one material and cited an unrelated team working on diamond substrates) and on packaging efficiency, where Vasudev said he personally expects optics to address interconnect-driven energy costs tied to data-center capacitance.
Teaching brain cells to play Doom, according to Sean Cole
Sean Cole, who worked with Cortical Labs, described treating cultured human brain cells as an input-output dynamical system rather than a matrix-multiplication device. Game state (health, ammo, screen image) is encoded into electrical stimuli delivered across a 59-channel multi-electrode array; resulting spikes are decoded into in-game actions. He contrasted this with Cortical Labs' earlier Pong demonstration, where hand-designed encoding/decoding sufficed for a simple two-action game, saying Doom's larger action space (54 joint actions) made hand-mapping across only 59 channels impractical.
The team used a PPO-based reinforcement learning loop, training an encoder via stochastic (beta-sampled) stimulation since gradients cannot propagate through the electrode array, and a deliberately undersized linear decoder to prevent the silicon system from "cheating" by learning to play the game itself rather than relying on the cells. Feedback was delivered as synchronous stimulation for good actions and asynchronous stimulation for bad ones, scaled by a critic's prediction error ("surprise"), drawing on Karl Friston's free energy principle. Cole said the system's two components never reach equilibrium, with entropy actively penalized rather than rewarded, unlike standard RL practice, because the cells themselves supply sufficient exploratory entropy.
"I think that human intelligence is not optimal for compute. We're optimized for survival... but we're not optimized for compute."— Sean Cole
Asked how scalable the approach is, Cole said the two open problems are reaching frontier-level intelligence on biological cells and then distributing that intelligence to many users, adding plainly that on the second problem, "we fundamentally don't know how to do this yet."
Written by AI from the video's transcript. It can compress, misattribute or miss context — the original video is the source. Not investment advice.









