Goldilocks Platforms [w James Urquhart]

A Goldilocks’ balance challenges us to trade off prescriptive and flexible platforms. James Urquhart shares his experiences with Cloud Foundry, VMware, and Amazon about trying to find the right balance between building it yourself versus a prescriptive service approach.

We’ve decided that there needs to be a middle zone with enough opportunity for customization, as well as enough pre-set, prescriptive methods to create sustainability.

In this episode, we talk about that balance and how different processes have done it in industry.

Transcript: otter.ai/u/OQBfCHldtYjUpqjKdkN3KjzLiR0
Image: www.pexels.com/photo/brown-teddy…h-outside-207891/

Rob’s Hot Take:

In the Cloud 2030 Podcast Lunch and Learn on March 22nd, Rob Hirschfeld explores the Goldilocks problem, focusing on the challenge of striking the right balance between prescriptive platforms and open toolboxes. He emphasizes the difficulty of handling the diverse and heterogeneous nature of various systems while aiming for reuse, reduction of toil, and collaboration. Hirschfeld points out the nuanced nature of variations within the 80-20 rule, stressing the importance of considering how toolboxy or prescriptive a platform should be based on individual needs. To delve deeper into this thought-provoking discussion with insights from James Urquhart, check out the full episode on March 22nd at the2030.cloud and become part of these engaging conversations.

Complexity vs Value [& Okta hack]

The Okta hack highlights the value versus complexity trade off. In today’s episode, we ask if the complexity of using single sign on is the right move in this context. We also think about how to deal with these interconnected systems that have high degrees of complexity.

We also discussed API design, and whether or not we should have more rigid or flexible APIs. You can’t remove complexity from the system, but you can hide it. The structure of APIs will push complexity into either the users’ realm or the operators’ realm.

Transcript: otter.ai/u/cftY6wlMTzAceT2EiHF4u4u0dpE
Image: www.pexels.com/photo/photo-of-an…ng-money-7884134/

Rob’s Hot Take:

In the Cloud 2030 Podcast on March 24th, Rob Hirschfeld delves into the complex relationship between complexity, value, and immutability in system design, particularly focusing on API interfaces. He emphasizes the trade-offs involved in exposing options to users, providing flexibility but potentially increasing complexity. The discussion highlights the practicality of using immutability and templates to control API complexity, acknowledging the challenges of finding the right balance and the importance of transparency in decision-making. To explore these insights further, listen to the full episode on March 24th at the2030.cloud and participate in these open conversations.

Improving Automation Safety

Making automation safe is essential to making it usable at scale. How do we make automation safe? We found a lot of great insights drawing from space craft design, aircraft, aircraft design and other systems where safety is super important.

Automation is a force multiplier. If we don’t factor in safety when we build it,then we could create a lot of harm in systems from wasteful spending to actual injury. These designs have very real implications.

Transcript: otter.ai/u/p9w4aKOqm3rpHhbDtRTaLgN3GIA
Image: www.pexels.com/photo/toddler-usi…-on-road-1642055/

Rob’s Hot Take:

In the Cloud 2030 Podcast on March 15th, Rob Hirschfeld underscores the critical importance of automation safety in system design. Emphasizing the need for thorough testing, he discusses how safety, especially in complex systems like airplanes and spacecraft, requires continuous testing and monitoring. The conversation delves into the significance of not just completing tasks but also exercising and testing systems in various scenarios to ensure their safety. To explore these insights further, listen to the full episode on March 15th at the2030.cloud and participate in the ongoing discussions.

Expanding GitOps Beyond K8s

GitOps is a really important way of collaborating and communicating about infrastructure.

But can GitOps escape from Kubernetes? While we did talk about Kubernetes too, we mainly talked about what it takes to implement GitOps outside of Kubernetes. We considered building a GitOps architecture and then having people understand and use it. We also cover the fundamental parts of GitOps like having a reconciler and a bunch of tools that drive clusters.

Transcript: otter.ai/u/oq4D06Sd_rtUvXBVXC0Wx3KA2sQ
Image: www.pexels.com/photo/people-with…popcorns-7234318/

Rob’s Hot Take:

In the March 8th DevOps Lunch and Learn session on GitOps, Rob Hirschfeld emphasizes the crucial role of immutability in operations. The concept of specifying a fixed state, configuration set, or resource transforms how automation, infrastructure building, and system maintenance are approached. The investment in immutable components enhances change resilience, making it easier to adapt and keep up with changes while ensuring stability. Join the ongoing conversations and roundtables at the2030.cloud to contribute to discussions on these transformative concepts.

Migrating Long Term Applications

How should we think about migrating legacy workloads to new infrastructure and modernize them?

The group addresses this question methodically incuding how databases get linked, how they get used, how they get migrated, how important it is to maintain languages and what it would take to migrate in language. In the end, we look back on that conversation apply lessons learned to what we are building today,

This is absolutely essential because new designs will become tomorrow’s legacy! We’ll be struggling to migrate those in 10 or 15 years too. So everything we can learn helps prevent that cycle.

Transcript: otter.ai/u/sHB8507KjZlZPBMToBUCEKjPVQY
Photo: www.pexels.com/photo/man-and-wom…tainside-8968077/

Rob’s Hot Take:

Hello, I’m Rob Hirschfeld, CEO and co-founder of RackN, providing a hot take on the January 25th discussion about migrating legacy applications to the cloud. While the topic may seem limited, the reality is that today’s legacy was once a cutting-edge application, underscoring the importance of designing with future migrations in mind. The key challenges identified in the conversation were complexity and coupling, emphasizing the need for clean, referenceable APIs to facilitate smoother migrations. To delve into these insights further, listen to the full episode on January 25th and join the ongoing discussions at the2030.cloud.

Can Machines Update Themselves?

We know that humans have trouble keeping systems updated, but… how can we address the challenge of knowing which updates are required and, critically, if the updates with break other systems? Even knowing if they worked is a really thorny problem!

In this episode, we focus on actions about what’s going on and why this problem has persisted in industry for so long. Starting from the news of the day about CentOS 8 mirrors being taken down. That’s exactly the type of challenge we are facing when we think about where updates and repos are coming from.

Transcript: otter.ai/u/rRMIT6kkTTtyWrzdBnuq63nvKuE
Photo: www.pexels.com/photo/a-man-using…quipment-5996696/

Rob’s Hot Take:

Rob Hirschfeld, CEO and co-founder of RackN, discusses the challenges of system maintenance and lifecycle in the Cloud 2030 podcast. He emphasizes the difficulty of keeping systems up to date and understanding dependencies, leading to a lack of confidence in system updates due to the fear of breaking or degrading them. Hirschfeld advocates for a change in the industry to prioritize test and verification practices, enabling more effective and confident system updates.

Can DevOps Be More Collaborative / MSFT & Activision

We have a lot of questions about improving collaboration in organizations:
How do we deal with change in organizations
How can we get organizations to work together better?
How do we encourage collaboration around the automation spaces that we’re trying to build in DevOps.

In our discussion, a lot came back to something as simple as version control!

We also discuss how we handle coupling between systems. In order to collaborate, we have to couple systems. But if we couple them, we create complexity.

This podcast includes our warm up conversation about Microsoft acquiring Activision because that is ALSO about how you integrate to organizations and business plans! This was news of the day and I think you’ll be very interested in our take on it.

Transcript: otter.ai/u/itNrmoL9MgG980D8CZdo9XWd-xI
Image: www.pexels.com/photo/a-woman-pla…ith-kids-7176471/

Machine Learning in Operations

Today’s episode is about how to trust machine learning in operations. This is a really serious issue because the attraction of machine learning is strong, but does not translate into operations.

Why doesn’t it translate? Because operations is a closed loop process where we constantly get feedback and have to adapt and adjust. That makes it difficult to train models and hope that they work. This discussion gets into why that’s the case and what we can do about it.

Then we explore scenarios for machine learning and AI in operations.

Transcript: otter.ai/u/UBjf5IVnKvebTfQgW1xlneZdjOU
Photo: www.pexels.com/photo/high-angle-…of-robot-2599244/

What is Platform Engineering?

What is platform engineering? And why is it necessary and how to make it work compared to DevOps.

In this conversation, we really hit on the challenges of creating automation teams for building automation in scalable ways. Frustratingly, we never really came up with a particularly good answer to “what is a platform team” and why you should care. Strangely, your organization is probably building one.

Transcript otter.ai/u/zJeQbqXIyD8kZUxfKQdvQAfQGog
Image: www.pexels.com/photo/building-co…chnology-9617733/

Rob’s Hot Take:

Rob Hirschfeld, CEO and co-founder of RackN and host of the Cloud 2030 Podcast, reflects on the November 9th DevOps Lunch and Learn session focused on platform engineering. He highlights the challenge of executing platform engineering initiatives despite the straightforward concept of improving automation and tooling at an architectural level. Hirschfeld emphasizes the importance of defining success metrics, empowering teams to enforce standards, and adopting consistent, repeatable patterns and practices to advance the industry’s maturity. He encourages listeners to explore the insightful discussion at the2030.cloud for a deeper understanding of platform engineering’s significance.

Modular Automation via Pipelines and Digital Twins

Today’s episode, we talked about the challenge of making Modular Automation. We broke down why that is so hard and really dug into ways in which we can increase the modularity and reuse of automation.

That led us to talking about infrastructure pipelines, infrastructure, reuse, and sharing state via digital twins in infrastructure.

All of this comes together in really fascinating ways!

Transcript: otter.ai/u/O1KFCsmEpMpV-hNZjTzJ_ewxaUA
Photo by Alena Darmel from Pexels [ID 7751055]