AWS Outage

In this episode, we discuss the October 2025 Amazon outage. The conversation took place during the outage, and though it’s been a few months now, the insights and discussions are still very interesting. We trace how a DynamoDB and DNS-related failure cascaded through core AWS services and had a larger blast radius than expected. We also look at whether the outage was accidental or malicious and compare it to previous large cloud outages caused by internal errors or cascading failures. Some really interesting ideas come up around redundancy, failover, local infrastructure, and how data-centered business models change priorities around accountability, compliance, and valuation.

Transcript: https://otter.ai/u/0SHTGqt3cmSEDX5v8YLsK7eDyIE?utm_source=copy_url

Is 2025 even harder than we expected?

We review 2025 predictions today and dig into why I think this year is going to be both boring and terrifying for a lot of enterprise IT leaders. That, of course, spans Amazon, Reinvent storage, VMware, AI, and Agentic AI – we run the gamut on what is coming and why this is actually going to be a very challenging year.

Transcript: otter.ai/u/H6UvLC-r2zmBO9A5jf…?utm_source=copy_url

Reference: zenoh.io by ZettaScale

Cloud2030predictionsstoragecontainerscloudAWSvmwareai

DeepThink AI and Kubernetes

We springboard from DeepThinking AI and have a robust conversation about what impact DeepThink is having on the industry. We also discuss where we see things going into the dilemma of people building AI infrastructure and working to do that quickly, robustly and with strong governance. This is necessary to ensure that they can quickly update and manage that AI infrastructure that they’re spending so much money to build, and this leads into a broader conversation about virtualization, containers and open shift.

Recorded Jan 30, 2025

Transcript: otter.ai/u/79JxdYOiXUoSS44pYP…?utm_source=copy_url

Reference: www.perplexity.ai/search/provide-a…TQ6LJG_X0SlB5g#8

KubeVirt in the Enterprise

This is one of those fun conversations where we’re really diving not just into the tech but the enterprise consumption of the tech and how people are thinking about it. How does technology like Kubernetes evolve and get used in ways that the community is not thinking about and find a whole new path for adoption and commercialization?

If this is going on in your organization, we want to hear from you. We want you to be part of the conversation, because this is a really important transition point for the industry, for people questioning their VMware consumption, and for people looking to expand their Kubernetes footprints.

Transcript: otter.ai/u/nEzH4t1JDEa51fvQFc…?utm_source=copy_url

Software Defined Edge

We revisit edge infrastructure and the motivations behind building and managing edge infrastructure with an unusual take. In this case, we ask ourselves if all of these edge devices are becoming more software defined or becoming more standardized, off the shelf component tree. And will that change how we look at managing and running edge infrastructure? Will we shift compute and operations processes into these ever smarter devices? The answer is going to surprise you.

Transcript: otter.ai/u/tGIcIC1bijvaW4OkJN…?utm_source=copy_url

Training Small LLMs

In this episode, we dive deep into the emerging world of building and training small language models. We’ll discuss the benefits, risks, and challenges companies face as they work to create more targeted and efficient AI models. From managing hardware and power requirements to ensuring data privacy and governance, we’ll cover the key considerations for enterprises looking to leverage the power of small language models. Join us as we unpack this fascinating topic and consider the implications for the future of AI and infrastructure operations.

Transcript otter.ai/u/xJ5T-x70WUFQ55ZAsRQr57q6zwE
Reference: www.composabl.com/

Process: Good, Bad And Ugly

This podcast episode explores the challenges of process improvement in IT operations, using examples from data centers, automotive, and cybersecurity.

The discussion covers the slow evolution of secure boot, the difficulties cloud providers face in translating their processes to the broader market, and the emergence of vehicle-to-anything ecosystems. The group delves into the need for standardization and security in vehicle ecosystems, as well as the policy management and automation challenges enterprises face.

The conversation also examines the balance of trust in technology versus human expertise, particularly around the use of AI and the risks of generative AI. The CrowdStrike incident is analyzed, with debate around the responsibility of CrowdStrike, Microsoft, and Delta’s operational controls. The impact on cyber insurance and the need for broader risk management approaches are also discussed, highlighting the interconnectedness of process improvement and risk management, and the call for greater industry collaboration to address these challenges.

Transcript: otter.ai/u/93JhNjmekqf0ttX21g…?utm_source=copy_url

Crowd Strike vs Operations Responsibility

This episode explores the intersection of infrastructure automation and security through the lens of the Crowd Strike outage. We’ll discuss the tension between maintaining stable, reliable data center infrastructure and the need to embrace change and innovation.

Recent events like the CrowdStrike outage demonstrate the paradox that infrastructure teams face. We’ll dive into the importance of having multiple control planes and standardized processes that can adapt to rapid industry changes.

Transcript: otter.ai/u/Wos9IOPfpSGPOYNT-A4muQccA5w

Cloud2030Crowd StrikeOperationsWindowsOutageCloud