In an industry captivated by large language models, one contrarian thesis is attracting serious capital and serious ambition: world models. Alex Lebrun, the French serial entrepreneur behind Wit.ai (acquired by Facebook), shared the strategic, technical, and philosophical choices behind Ami Labs — a startup that raised €1.2 billion at seed to build foundational AI differently.
This isn't just about raising record capital. It's about rethinking how machines learn, challenging the LLM orthodoxy, and betting on a future where AI develops common sense through direct sensory experience — not just text.
🧠 What Are World Models, and Why Do They Matter?
Large language models learn by reading text — essentially consuming humanity's written records. They are, as Lebrun puts it, "like someone who has never left the room they were born in, but read all the books every day for centuries." The result? Impressive linguistic fluency, but no grounded understanding of the physical world.
World models take a fundamentally different approach. Instead of learning through language, they learn directly from sensory data: video, audio, touch, and interaction. The analogy? A baby learning about the world by observing, touching, and experimenting — not by reading Wikipedia.
"A world model is a model where we don't cheat. We are learning directly like humans, like animals, from the real world."
This approach promises machines with common sense and the ability to operate safely in open, unpredictable environments — something current robotics and VLAs (vision-language-action models) struggle with.
🤖 The Robotics Use Case: From Narrow Automation to General Purpose
Today's robots are brittle. They excel at narrow, repetitive tasks in controlled settings but fail spectacularly in open environments. Lebrun cited viral examples: robots doing karate demos that nearly strike children, or dancing units that destroy everything around them.
Current robotics relies on hardware evolution without corresponding advances in AI cognition. World models aim to change that by providing robots with spatial reasoning, physics intuition, and adaptive behavior — the cognitive substrate needed for household assistants, outdoor autonomy, and flexible industrial work.
Why not just use VLAs? Vision-language-action models are, in Lebrun's words, "a very bad hack where you have a hammer (an LLM) and try to see everything as a nail." They are:
- Inaccurate for real-time physical tasks
- Too slow (high latency kills robotics applications)
- Prohibitively expensive for 24/7 operation
World models, by contrast, are designed to be lightweight at inference once trained, with fewer parameters and faster response times — essential for embodied AI.
💡 LLMs vs. World Models: Complementary, Not Competitive
Lebrun is clear: this isn't about replacing LLMs. Language models excel at low-dimensional, discrete symbol manipulation — language, math, code. For these domains, they're unbeatable.
But the world is high-dimensional, noisy, and continuous. For tasks requiring spatial reasoning, temporal prediction, and sensory grounding, world models are expected to be "vastly superior."
"LLMs don't lead to artificial general intelligence. Most people disagreed a year ago. Now I think most people agree."
The shift in consensus is subtle but significant. As the AGI narrative around LLMs recedes, attention is turning to architectures that can handle the messy, embodied challenges of the real world.
🏗️ The Ingredients: Talent, Data, Compute — and Risk
Building a foundational model requires three critical inputs:
- Talent: Deep expertise in a nascent field where most researchers are already committed to LLM work
- Data: Vast amounts of video, audio, and eventually robotic sensory data
- Compute: Thousands of GPUs — secured not just with money, but with strategic planning in a supply-constrained market
Even with €1.2 billion in the bank, securing compute remains a bottleneck. Capital is necessary but not sufficient in the current GPU landscape.
Yet Lebrun insists the hardest part isn't the money — it's the risk tolerance. Startups can take bets that large corporations cannot. Meta, Google, and others had access to the same transformer research that powered ChatGPT, but didn't productize it aggressively.
"The one thing you have that big companies don't is the ability to take crazy risks. If you don't take any risk, you're dead."
🎯 The Real Cost of a €1.2B Seed: Expectations, Not Dilution
Raising the largest seed round in European history wasn't the hard part. Lebrun had done it before — multiple exits, deep credibility, and a co-founder in Yann LeCun, one of AI's founding figures.
But the capital came with a hidden tax: expectations.
"The real cost of this €1.2 billion was not dilution. The real cost is expectations... If you raise a billion and people don't see anything coming out of it for two years, then it's very hard to survive."
The team is operating on an internal timeline, though they're tight-lipped on public milestones. The pressure isn't just financial — it's reputational and existential.
🌍 Why Paris, Not SF?
Ami Labs is headquartered in Paris, with additional hubs in New York, Montreal, and Singapore. Notably absent? San Francisco.
It's a deliberate choice. Requiring engineers to relocate — especially from SF to New York — signals commitment. It filters for builders who are all-in, not tourists testing the waters.
"It's very important at this early stage to have people who join with real commitment, not just because they want to try."
Still, the team maintains deep SF connections, traveling frequently and staying embedded in the Bay Area ecosystem. The strategy? Capture the best of both worlds without the distraction or dilution of a crowded competitive landscape.
📚 Lessons from Four Startups: Start Narrow, Think Big
Lebrun's entrepreneurial journey spans two decades and four companies:
- Virtual (2002): A chatbot company launched 20 years too early
- Wit.ai (2013): Conversational AI acquired by Facebook in 2015
- Nabla: An AI-powered healthcare assistant
- Ami Labs (2024): Foundational world models
His advice to founders? Start narrow, but be ambitious within that narrow scope.
"Choose a very narrow problem — one thing for one industry. But in this narrow tunnel, have a very big vision. Don't be afraid to share your long-term ambition."
Trying to solve everything for everyone as a small startup is a recipe for failure. But solving one thing exceptionally well — with a credible roadmap to something transformative — is how you earn attention, capital, and momentum.
🚀 The 5-10 Year Vision: Robots with Common Sense
If Ami Labs succeeds, the world will look different. Not unrecognizable, but augmented:
- Helpful robots in homes and workplaces — capable, adaptive, and safe
- Jobs transformed, not eliminated — less dangerous, less repetitive, more human-centric
- Machines with common sense — able to reason about the physical world in real time
"Most jobs will still be here, but not as dangerous or hard to do as they are today."
It's a future where AI doesn't just talk about the world — it understands it.
✅ Key Takeaways
- World models learn from sensory experience, not text — offering a path to embodied intelligence and common sense
- LLMs and world models are complementary, not competitive — each excels in different domains
- Raising massive capital brings massive expectations — the real cost isn't dilution, it's delivery pressure
- Risk tolerance is a startup's edge — big companies can't afford to be as bold
- Start narrow, think big — solve one problem exceptionally well, but articulate a transformative long-term vision
- Building outside SF can be strategic — if you stay connected and filter for committed talent
As the AI landscape matures, the next frontier isn't just about more data or bigger models — it's about different architectures that ground intelligence in the real world. Ami Labs is betting €1.2 billion that the future belongs to machines that don't just read about cats — but have actually met one.