Kubernetes as a Common Platform

This week we examine the emergence of Kubernetes as a common infrastructure platform. We contrast cloud assumptions with bare-metal reality and dig deep on supportability, repairability, and the operational challenges of networking, storage, GPUs, and other specialized systems. We also get into brittle vendor tooling, version changes, and how these can make remediation difficult. I think you’ll like this one!

Transcript: https://otter.ai/u/cA80HK4iVJWflodPqIYf9mW3Mec?utm_source=copy_url

Kubecon SC25 Debrief

In this episode, we debrief several industry events I went to last year, including Supercomputing, KubeCon, Stack, the AI Infrastructure Show, and the Red Hat AI Infrastructure Summit. We dive deep into some observations from the shows and what they tell us about the gaps and fractures in how we are working to build AI infrastructure. We focus on how observability is being used for evaluation, tuning, performance issues, GPU dropouts, and cluster management, while anomaly detection and root cause analysis remain less common, and we note that networking is still underserved. We also get into the shift from building clusters to observing and fixing them after deployment, especially for agentic systems, and we end by highlighting the need for observability across application, identity, networking, and infrastructure layers.

Transcript: https://otter.ai/u/y6FNvERJRe_8qnmAgVlmvd6kwb8?utm_source=copy_url

AWS Outage

In this episode, we discuss the October 2025 Amazon outage. The conversation took place during the outage, and though it’s been a few months now, the insights and discussions are still very interesting. We trace how a DynamoDB and DNS-related failure cascaded through core AWS services and had a larger blast radius than expected. We also look at whether the outage was accidental or malicious and compare it to previous large cloud outages caused by internal errors or cascading failures. Some really interesting ideas come up around redundancy, failover, local infrastructure, and how data-centered business models change priorities around accountability, compliance, and valuation.

Transcript: https://otter.ai/u/0SHTGqt3cmSEDX5v8YLsK7eDyIE?utm_source=copy_url

Back After a Break

In this episode, we discuss the rising cost of using AI and how usage-based pricing, model changes, and capacity limits are affecting daily work as AI moves from experimentation into operational use. We also talk about multi-model workflows, hybrid infrastructure, and examples of using hosted models alongside open models locally for tasks such as writing and named entity resolution. We get into the need for enterprises to run their own AI infrastructure, including questions around GPU pooling, routing, reservation, data sovereignty, and service levels.

MCP Agents and Context

In this episode, we continue our journey even deeper into how agentic vibe coding and other AI-based automation. This time we focus on Model Control Protocol (MCP) and its application in our bare metal automation solution, Digital Rebar. We examine deterministic versus stochastic AI approaches and the importance of reliable system integration without competing with other agentic systems. We highlight MCP’s role in streamlining interactions across data sources, with a focus on practical applications in finance and infrastructure resilience. The episode ends with a preview of future conversations on user experience transformation in infrastructure operations. Enjoy!

Transcript here: otter.ai/u/LmtQ9QAc79izN0PacE…=transcript&tab=chat

Vibe Coding Mapping [TechOps]

In this episode, we continue our Vibe Coding experiment. Now that we’ve figured out how to interface with MaaS, this time we wrestle with mapping and how different systems interact with each other. We’re joined by Greg Althaus, RackN CTO, who reviews the project and asks some really great questions. We talk about our decision to restart the experiment, taking the lessons we’ve learned to the newer software available. Enjoy!

Transcript: https://otter.ai/u/NohM8_D7DfbUu9iZc1MSpxi00kg?utm_source=copy_url 

Rob Weinhold: The Art of Crisis Leadership [Cloud 2030 Book Club]

In this episode, we talk about Rob Weinhold’s book, “The Art of Crisis Leadership.” We explore the vital principle of “owning your narrative” in crisis management, and we share some personal stories related to the themes in the book. We analyze the differences between personal and organizational crises, emphasizing storytelling, transparency, and trust as keys to effective leadership. Even if you haven’t read the book, there’s a lot to get out of this great conversation.

Transcript: otter.ai/u/acAsrwSOpObslvjSVg…?utm_source=copy_url

Vibe Coding for Ops [TechOps]

In this episode, we do some live vibe coding– using AI to write code. We share tips and tricks on having the best vibe coding experience and avoiding some common pitfalls. You’ll get to hear what we do, how we discover what the steps are, just how easy it is to interact with the system, to set up a basic environment. We also start to explore the limitations of vibe coding. We encourage you to listen along and try on your own!

Transcript: otter.ai/u/CqKdtWZWYb3AdPtcb-…?utm_source=copy_url

Model Context Protocol Exploration

Today we continue our exploration of vibe coding by digging into the Model Context Protocol, or MCP. We look at how MCPs connect chatbots to backend systems, why natural language matters for complex queries, and what it takes to build smarter, more adaptable interfaces. The discussion covers practical strategies for refining and automating these systems using API docs, making this a solid deep dive into the future of human-to-machine interaction.

Transcript: otter.ai/u/5zn42OkdumP-HXIdi5…?utm_source=copy_url