Hello, Agent:

The real bottleneck in scaling AI isn't compute with Bassam Tabbara at Upbound

August 25, 2026

Show Notes:

In this episode, we talk with Bassam Tabbara, Founder and CEO at Upbound, Founder of Crossplane, and Founder of Modelplane, about what it takes to run AI inference at scale.

Bassam walks us through his path from writing BASIC on a ZX Spectrum in Lebanon to building the early automation behind Hotmail at Microsoft, founding Crossplane, and then Modelplane: a new open-source fleet orchestrator built for running any model, on any engine, on any hardware.

Key Takeaways:

(00:00) Alex introduces Bassam Tabbara, inventor of Rook and creator of Crossplane, here to discuss his newest open-source release, ModelPlane.

(00:47) Bassam's control-plane origin story: running half a million computers at Akamai with a 24-person NOC taught him reconciliation loops are the only way to scale.

(01:50) Owning your own intelligence isn't just about cost — it's about training models on your company's private data and controlling who has access to it.

(03:20) ModelPlane was inspired by Upbound's enterprise customers, who are already running inference on their own Kubernetes clusters — a workload that doesn't bin-pack and forces you to span clusters, providers, and deployment modes.

(04:05) GPU scarcity changes the game: capacity reservations, H100s reselling above list price, and wildly varying costs across regions and providers.

(05:30) The AI stack is tipping horizontal — more choice at the model, serving-engine, infrastructure, and accelerator layers, with tighter coupling between models and engines coming.

(07:20) The chip-manufacturer example: Global 5000 enterprises with deep intellectual property have no real alternative but to fine-tune open-weight models on their own business processes.

(09:45) Fleet-wide inference was the unsolved problem in open source — vLLM and SGLang solved the cluster and node scope, so Upbound built ModelPlane to run the right model on the right engine on the right hardware, globally.

(11:35) ModelPlane is literally built on Crossplane, extending the Linux → Kubernetes standardization lineage to inference — and everyone running inference at scale (labs, neoclouds, hyperscalers) has already built this privately.

(14:35) Separation of concerns: platform teams deploy ModelPlane with cloud credentials, cost policy, and governance, while ML teams declare model deployments as app resources alongside their agents.

(17:00) The API-vs-self-host trade-off: frontier APIs win on ease today, but CIOs need real choices around cost, sovereignty, and compliance — the world won't be uni-model, uni-lab.

(19:55) Bassam's car analogy: Ferraris (frontier APIs) will exist, but BMWs and Toyotas (self-hosted open weights) will serve most segments — open weights tipped the market horizontal, and choice brings fragmentation.

(22:40) ModelPlane is just getting started at modelplane.ai — Bassam invites the open-source community to contribute, backed by Crossplane's CNCF graduation track record.

(23:52) Alex closes the episode on sovereign intelligence and teases next week's guest: Catherine, CEO of Kernel AI, on computer-use agents.

Thanks for listening to "Hello Agent!: The podcast at the intersection of data & agents." Remember to subscribe so you don't miss an episode.

Transcript

Red panda wearing headphones speaking into a microphone with an orange patterned background.

Learn when each episode drops

Subscribe and never miss a Redpanda 'Hello, Agent' podcast. We hate spam and will never sell your contact information.

Other episodes

View all episodes
Jeremy Edberg
C-Suite Advisor
@
DBOS, Inc.

Durable execution, reliability engineering, and the future of agentic AI with Jeremy Edberg at DBOS

Jeremy Edberg explains how durable execution lets AI agents save state, replay, and resume after failure—filling a key infrastructure gap.

Play episode
Text Link
Dominik Tornow
Founder & CEO
@
Resonate

Building interruption-tolerant agents with durable execution with Dominik Tornow at Resonate HQ

Dominik Tornow explores why durable execution is critical for reliable AI agents and how durable promises simplify building long-running, interruption-tolerant multi

Play episode
Text Link
Nicolas Dupont
Founder & CEO
@
Cyborg

Designing secure architectures for AI agent-driven workflows

Nicolas Dupont explores the world's first confidential vector database and how to deploy RAG agents securely on regulated, sensitive data.

Play episode
Text Link
Red panda wearing headphones speaking into a microphone with an orange patterned background.

Stay up-to-date with the latest 'Hello, Agent' episodes

Learn how industry pioneers are building, deploying, and scaling enterprise AI agents. Sign up to get new episodes in your inbox.