Launching The Akuity Agentic Control Plane Learn More →

Launching The Akuity Agentic Control Plane Learn More →

Fleet Management In Practice: Four Real-World Use Cases

Hong Wang

Fleet Management In Practice: Four Real-World Use Cases
Fleet Management In Practice: Four Real-World Use Cases

By Hong Wang, CEO of Akuity and co-creator of Argo CD and Kargo

This is part 3 of our fleet management series. If you're new to the series, start with Part 1: The Fleet Is the New Cluster and Part 2: The Fleet Management Gap.

TL;DR

  • Many platform teams now deliver to fleets of clusters, customer environments, and stores, but still use delivery tools built for a handful of environments.

  • Common use cases show what that costs, from patching a CVE across shared clusters to keeping regulated rollouts auditable.

  • The same challenges run through each case, and they share a root cause: today's tools let teams control each target or see the whole rollout, but not both.

  • Fleet management, coming soon to Kargo Enterprise, gives teams both.

Platform teams used to deliver to a few environments: dev, staging, production. Today, many deliver to a fleet: dozens of clusters, hundreds of customer environments, or thousands of stores and sites that all need the same change.

I co-created Argo CD nearly a decade ago and Kargo after it, and I've spent the years since working with platform teams that run them at scale. Across those teams, the same practice comes up under different names: waves, rings, partitions, canaries, risk tiers. Each is a way of moving a change through the fleet one group at a time, and many teams still do it with homegrown scripts and manual steps.

According to the 2026 Argo CD user survey, 62% of respondents at large companies promote changes between environments with custom CI scripts.

Part 1 of this series made the case that the fleet has replaced the cluster, and Part 2 showed where today's tools stop short of managing a fleet. This third post walks through common fleet management use cases, drawn from platform teams across industries, and the challenges they share.


Figure 1: The same change, before and now. Before, it moved through three environments in turn. Now it has to move through hundreds or thousands of targets in waves, starting with a small pilot group.

Figure 1: The same change, before and now. Before, it moved through three environments in turn. Now it has to move through hundreds or thousands of targets in waves, starting with a small pilot group.

What are common fleet management use cases?

Common fleet management use cases include patching a CVE across shared clusters, releasing to hundreds of customer clusters, rolling out to thousands of store clusters, and keeping regulated rollouts auditable. In each case, the change is routine, but the challenge is moving it safely through the fleet.

#1 Patching a CVE across shared clusters

A platform engineer needs to patch cert-manager, the add-on that issues and renews TLS certificates, across 100 clusters to close a vulnerability. The usual approach is to update the version in a central config file and let it roll out to every cluster at once. There's no pilot or canary group and no health check afterward, so a bad rollout surfaces hours or days later, when certificates start expiring.

Teams we've talked to fall into two camps. Some build control themselves: they send each patch through release branches, manual gates, and manager approvals, then wait for each cluster's maintenance window and track progress across tickets, chat threads, and Git history. Others want controls their tooling doesn't give them: skipping clusters in a maintenance window to return to later, and pausing or resuming a rollout at any time.

The cost: Until the patch lands everywhere, the vulnerability stays open on every cluster the rollout hasn't reached, and the team has no quick way to see which clusters those are.

#2 Shipping a release to hundreds of customer clusters

A SaaS company runs a dedicated Kubernetes cluster for each of its 300 enterprise customers, and every cluster runs the same 100 apps. For a routine release, the team has two options. Release to all 300 clusters at once and lose sight of which customer's cluster broke. Or give every app on every cluster its own pipeline: 100 apps across 300 clusters is 30,000 pipelines, each with its own approval, and no single place to see where the release stands.

Both show up in practice. One provider sends a single repository tag to every customer cluster at once, with no pilot group and no feedback until the release is everywhere. It wants feedback before a change merges, and releases that move one customer at a time. Another gives every customer its own maintenance window, so it has to decide the rollout order customer by customer.

The cost: If the release goes to all 300 clusters at once, a failure the team could have caught on one customer's cluster reaches every customer. If it goes out per pipeline, every release needs 30,000 approvals, too many to work as a gate.

#3 Rolling out to thousands of store clusters

A retail chain runs a small cluster in each of its thousands of stores and needs to roll out a new point-of-sale build without a bug reaching the busiest stores first. The platform team sets the rollout order by hand and keeps it in a spreadsheet or hardcodes it into the pipeline, so the order goes stale whenever a store opens or closes.

Teams running clusters at physical sites told us much the same. One rolls out in batches ordered by risk, from lower-stakes locations to the most critical. With no health visibility between batches, an engineer watches each batch until services restart cleanly, then starts the next. Another wants ring-based releases that start with its lowest-risk sites.

The cost: Every rollout needs an engineer watching from the first store to the last, and every opening or closure means updating the rollout order by hand. The highest-stakes part of the rollout comes last: the busiest stores, the ones the business can least afford to disrupt, after hours of manual watching.

#4 Keeping regulated rollouts approved and auditable

A compliance lead needs every production change, across clusters in several countries and regions, to have a sign-off and an entry in the audit log. Teams designed the approval process for a few environments, so they either collect an approval for every cluster or can't prove which clusters got a change, when, and who approved it. At air-gapped sites, someone also has to load the update by hand before the rollout can start.

Regulated teams we've spoken with feel the strain in three ways. One team plans to name, for every cluster, who must approve a change to it. Another, running a large central Argo CD estate, wants ordered groups within a single production stage, Jira approval gates, and CVE checks before promotion, and calls its current auditability weak. A third can reach its production environment only through a one-way data transfer link, so someone loads every update by hand before a rollout can start.

The cost: Every cluster added is another approval to collect or another gap in the audit trail, and every air-gapped site adds manual work before a rollout can begin.

What challenges do these use cases share?

The same challenges come up in each use case, and they share a root cause: today's tools let teams control each target or see the whole rollout, but not both.

Put the whole fleet in one step, and you get one view of the rollout, but no way to order it, pause it, or stop it at the first bad target. Give each target its own step, and you get that control back, but the rollout splinters into hundreds or thousands of pipelines with no single view across them. Each challenge below is one side of that trade-off, or a workaround for it.

  1. Rollout order lives outside the delivery tool: Teams roll out changes in groups, but they decide those groups and their order outside the delivery tool, in scripts, spreadsheets, or an engineer's judgment. As the fleet changes, the target list drifts out of date.

  2. Failures aren't contained: A release that goes everywhere at once can't stop at the first bad target. A release that goes site by site can stall when one site fails. And without rollback as a single operation, there's no clean way back from either.

  3. There's no single view of the rollout: Tools built around one cluster or one application at a time can't show which version runs where, which targets are healthy, or what's blocked. Teams piece that answer together by hand.

  4. Governance doesn't scale with the fleet: Teams designed approvals, ticketing, and security gates for a handful of environments. Run once per target, they multiply as the fleet grows, and air-gapped sites add manual work before a rollout can start. At fleet scale, governance needs to run once per rollout, leave one record of every target reached, and work across air-gapped sites.

These challenges map to a building block we outlined in The Fleet Is the New Cluster: The Infrastructure Shift Platform Engineers Can't Ignore

Fleet management is coming soon to Kargo Enterprise

Argo CD keeps clusters in sync with the desired state. Kargo, our continuous promotion tool, moves releases through environments like dev, staging, and production, with gates and verification at each step.

Fleet management is the next layer: promotion that works across hundreds or thousands of targets, not only a few environments, so teams no longer have to choose between controlling each target and seeing the whole rollout. Fleet management is coming to Kargo Enterprise, with early access starting in December.

Talk with an Akuity solutions architect to learn more.

Hong Wang is the CEO of Akuity and a co-creator of Argo CD and Kargo.

Ready to simplify delivery with Akuity?

Deploy, promote, and operate applications reliably, powered by OSS you trust and Intelligence you control.

Ready to simplify delivery with Akuity?

Deploy, promote, and operate applications reliably, powered by OSS you trust and Intelligence you control.

Ready to simplify delivery with Akuity?

Deploy, promote, and operate applications reliably, powered by OSS you trust and Intelligence you control.

Sign Up for Akuity Updates

Practical guidance on MTTR reduction, GitOps at scale, and safe automation, with product updates from the Argo CD and Kargo team.

@2026 Akuity Inc. All rights reserved.

Akuity Inc. 440 N. Wolfe Road, Sunnyvale, CA 94085-3869 US +1-510-771-7837

SOC2 Type 2 Compliant

Sign Up for Akuity Updates

Practical guidance on MTTR reduction, GitOps at scale, and safe automation, with product updates from the Argo CD and Kargo team.

@2026 Akuity Inc. All rights reserved.

Akuity Inc. 440 N. Wolfe Road, Sunnyvale, CA 94085-3869 US +1-510-771-7837

SOC2 Type 2 Compliant

Sign Up for Akuity Updates

Practical guidance on MTTR reduction, GitOps at scale, and safe automation, with product updates from the Argo CD and Kargo team.

@2026 Akuity Inc. All rights reserved.

Akuity Inc. 440 N. Wolfe Road, Sunnyvale, CA 94085-3869 US +1-510-771-7837

SOC2 Type 2 Compliant