Vibe 3-Step [TechOps]

In this episode we work through a Vibe Coding Operations session using Claude AI inside Digital Rebar as an analysis tool. We describe a three-phase flow of plan, execute, and analyze for inspecting a machine, gathering command output, and reviewing the results. We also talk about troubleshooting an installer problem, differences in user account flag behavior, and improvements in Claude’s debugging across sessions.

Transcript: https://otter.ai/u/-wk5QfPFPgYOxm_dOvTSIezDwTs?utm_source=copy_url

When Vibe Fails [TechOps]

This episode we discuss a failure in vibe coding with Claude and Digital Rebar. We try to build a CloudMD-driven setup for running Claude in a container and analyzing inputs, and we run into some problems when the generated design becomes too complex and leads to wrong assumptions. We discuss the broader lesson that too much context can reduce reliability, and that an incremental design approach may work better before extending the idea into FlexiFlow, templates, and other AI-assisted infrastructure tasks.

Transcript: otter.ai/u/aW8p_bf235g4Eg8nr-…?utm_source=copy_url

Vibe Digital Rebar Install [TechOps]

In this episode, we continue our exploration of vibe coding, this time we build a prompt for installing Digital Rebar. We test and refine the prompt across multiple runs, focusing on readiness checks, firewall ports, SSH keys, license file handling, and keeping the model within the provided instructions.We also look at where the install keeps failing, including NAT gateway setup, DNS configuration, key selection, and bootstrap completion, and use those failures to make the instructions more explicit. Finally, we discuss using prompts as install instructions for an agent system and how the same approach could apply to other complex setups like as OpenShift.

Vibe Coding Project Restart [TechOps]

This week, we continue our Vibe Coding, and use the lessons and prompt instructions we’ve learned from previous sessions to restart the project using the latest version of Vibe Code. We walk through the process of restarting and see the significant difference made by the new tech! We also talk through some alternatives to using Visual Studio. Another great session! The video is available on the2030.cloud.

Transcript: https://otter.ai/u/HowJk3d6_RPrRCsE2TOTl1rLQHU?utm_source=copy_url

AI UX Building

In this episode of Cloud2030, we explore AI’s transformative impact on user experience (UX) and the relevance of stochastic and deterministic systems in designing effective AI interfaces. We give our predictions about the decline of traditional form-based UX in favor of conversational interfaces that allow users to interact with AI more naturally.
We go into how AI can enhance user experiences by dynamically gathering information, while also addressing security concerns related to sensitive data. It’s a great conversation, hope you enjoy!

Transcript: otter.ai/u/_KLPdnZ6UvfApikAAt…?utm_source=copy_url

Kubernetes as a Common Platform

This week we examine the emergence of Kubernetes as a common infrastructure platform. We contrast cloud assumptions with bare-metal reality and dig deep on supportability, repairability, and the operational challenges of networking, storage, GPUs, and other specialized systems. We also get into brittle vendor tooling, version changes, and how these can make remediation difficult. I think you’ll like this one!

Transcript: https://otter.ai/u/cA80HK4iVJWflodPqIYf9mW3Mec?utm_source=copy_url

Kubecon SC25 Debrief

In this episode, we debrief several industry events I went to last year, including Supercomputing, KubeCon, Stack, the AI Infrastructure Show, and the Red Hat AI Infrastructure Summit. We dive deep into some observations from the shows and what they tell us about the gaps and fractures in how we are working to build AI infrastructure. We focus on how observability is being used for evaluation, tuning, performance issues, GPU dropouts, and cluster management, while anomaly detection and root cause analysis remain less common, and we note that networking is still underserved. We also get into the shift from building clusters to observing and fixing them after deployment, especially for agentic systems, and we end by highlighting the need for observability across application, identity, networking, and infrastructure layers.

Transcript: https://otter.ai/u/y6FNvERJRe_8qnmAgVlmvd6kwb8?utm_source=copy_url

AWS Outage

In this episode, we discuss the October 2025 Amazon outage. The conversation took place during the outage, and though it’s been a few months now, the insights and discussions are still very interesting. We trace how a DynamoDB and DNS-related failure cascaded through core AWS services and had a larger blast radius than expected. We also look at whether the outage was accidental or malicious and compare it to previous large cloud outages caused by internal errors or cascading failures. Some really interesting ideas come up around redundancy, failover, local infrastructure, and how data-centered business models change priorities around accountability, compliance, and valuation.

Transcript: https://otter.ai/u/0SHTGqt3cmSEDX5v8YLsK7eDyIE?utm_source=copy_url

Back After a Break

In this episode, we discuss the rising cost of using AI and how usage-based pricing, model changes, and capacity limits are affecting daily work as AI moves from experimentation into operational use. We also talk about multi-model workflows, hybrid infrastructure, and examples of using hosted models alongside open models locally for tasks such as writing and named entity resolution. We get into the need for enterprises to run their own AI infrastructure, including questions around GPU pooling, routing, reservation, data sovereignty, and service levels.

MCP Agents and Context

In this episode, we continue our journey even deeper into how agentic vibe coding and other AI-based automation. This time we focus on Model Control Protocol (MCP) and its application in our bare metal automation solution, Digital Rebar. We examine deterministic versus stochastic AI approaches and the importance of reliable system integration without competing with other agentic systems. We highlight MCP’s role in streamlining interactions across data sources, with a focus on practical applications in finance and infrastructure resilience. The episode ends with a preview of future conversations on user experience transformation in infrastructure operations. Enjoy!

Transcript here: otter.ai/u/LmtQ9QAc79izN0PacE…=transcript&tab=chat