AWS Outage

In this episode, we discuss the October 2025 Amazon outage. The conversation took place during the outage, and though it’s been a few months now, the insights and discussions are still very interesting. We trace how a DynamoDB and DNS-related failure cascaded through core AWS services and had a larger blast radius than expected. We also look at whether the outage was accidental or malicious and compare it to previous large cloud outages caused by internal errors or cascading failures. Some really interesting ideas come up around redundancy, failover, local infrastructure, and how data-centered business models change priorities around accountability, compliance, and valuation.

Transcript: https://otter.ai/u/0SHTGqt3cmSEDX5v8YLsK7eDyIE?utm_source=copy_url

Back After a Break

In this episode, we discuss the rising cost of using AI and how usage-based pricing, model changes, and capacity limits are affecting daily work as AI moves from experimentation into operational use. We also talk about multi-model workflows, hybrid infrastructure, and examples of using hosted models alongside open models locally for tasks such as writing and named entity resolution. We get into the need for enterprises to run their own AI infrastructure, including questions around GPU pooling, routing, reservation, data sovereignty, and service levels.

Vibe Coding Mapping [TechOps]

In this episode, we continue our Vibe Coding experiment. Now that we’ve figured out how to interface with MaaS, this time we wrestle with mapping and how different systems interact with each other. We’re joined by Greg Althaus, RackN CTO, who reviews the project and asks some really great questions. We talk about our decision to restart the experiment, taking the lessons we’ve learned to the newer software available. Enjoy!

Transcript: https://otter.ai/u/NohM8_D7DfbUu9iZc1MSpxi00kg?utm_source=copy_url 

TechOps Scaling Challenges

In this episode, we talk about scale and the hard realities of system failure in large tech operations. We explore why rare failures become common at scale, and what it takes to build systems that can handle that pressure. From predictive diagnostics to component redundancy, we share practical insights on keeping high-performance and AI infrastructure resilient. This is not theory, it is grounded in real-world lessons from managing complex environments and learning how to plan, isolate, and adapt when things go wrong.

Transcript: otter.ai/u/X8JYiADfPPLEfQ-gge…?utm_source=copy_url

The Opportunity for OpenShift Infrastructure

Today we tackle the generational infrastructure shift that’s keeping IT leaders awake at night: OpenShift virtualization adoption. We dig deep into why organizations are struggling to migrate from traditional VM-focused infrastructure to Kubernetes-managed infrastructure. We explore the real hurdles blocking this transition and unpack the strategic positioning that matters when you’re moving to container-orchestrated infrastructure. This isn’t about dumping everything into Kubernetes and calling it done, we examine what it really takes to use Kubernetes as your infrastructure abstraction layer while navigating the operational realities that make or break these migrations.

Transcript: otter.ai/u/IY2Y0a4aFN99ILg9da…?utm_source=copy_url

HA Troubleshooting [Tech Ops]

This episode of the TechOps series goes into high availability troubleshooting. Not just high availability, not just troubleshooting, but actually talking through what it takes to manage and maintain and fix HA systems. This is part of a longer discussion we’ve been having and so there’s some really interesting ideas in the middle of these discussions that I hope will shape your thinking as you build high availability systems, diagnostics and troubleshooting for people who are in high availability very complex environments.

Transcript: otter.ai/u/wM__4w1YIzZnhVdgLu…?utm_source=copy_url

References:
status.openai.com/incidents/ctrsv3lwd797\

High Availability Technology in DRP [TechOps]

Today we dive into RackN high availability technology and what we did to build consensus based raft HA capabilities directly into Digital Rebar. This is one of those episodes where we are talking specifically and only about Digital Rebar, so it is a vendored conversation from that perspective.

If you are building HA systems, or are interested in how HA systems work, this is a great session to learn firsthand from our experience!

Transcript: otter.ai/u/9lA9djczp5GkJbj12k…?utm_source=copy_url

Why is adding LLM into an App so hard?

We talk about current events, the acquisition of data stacks and the closing of the HashiCorp acquisition by IBM. Later, we dive into the productivity of AI and what’s going on – are companies really getting the benefits that they expect from AI chat bot integrations and what the challenges are?

We touch base on a little bit of something more infrastructure focused, where I give a preview of work I’ve been doing on separating Kubernetes virtualization from Kubernetes development use cases, which is something that we will be talking about more in the future.

References:
www.windowscentral.com/software-apps…ind-a-paywall
www.ibm.com/new/announcements/i…ise-ai-applications
www.youtube.com/watch?v=Ioc3r70HNLM
www.linkedin.com/posts/dhinchclif…9498138624-jR2R/
20250227

Software Defined Edge

We revisit edge infrastructure and the motivations behind building and managing edge infrastructure with an unusual take. In this case, we ask ourselves if all of these edge devices are becoming more software defined or becoming more standardized, off the shelf component tree. And will that change how we look at managing and running edge infrastructure? Will we shift compute and operations processes into these ever smarter devices? The answer is going to surprise you.

Transcript: otter.ai/u/tGIcIC1bijvaW4OkJN…?utm_source=copy_url