Fleet Management In Practice: Four Real-World Use Cases
Hong Wang
By Hong Wang, CEO of Akuity and co-creator of Argo CD
This is part 3 of our fleet management series. If you're new to the series, start with Part 1: The Fleet Is the New Cluster and Part 2: The Fleet Management Gap.
TL;DR
Many platform teams now deliver to fleets of clusters, customer environments, and stores, and still move changes through them by hand.
Common use cases show what that costs: patching a CVE across shared clusters, releasing to hundreds of customer clusters, rolling out to thousands of store clusters, and keeping regulated rollouts approved and auditable.
The same challenges run through all of them: rollout order is managed by hand, failures aren't contained, there's no single view of the rollout, and governance doesn't scale with the fleet.
The root cause: today's tools force a choice between controlling each target and seeing the whole fleet.
Platform teams used to deliver to a handful of environments: dev, staging, production. Today, many deliver to a fleet: dozens of clusters, hundreds of customer environments, or thousands of stores and sites that all need the same change.
In our conversations with these teams, the same practice comes up under different names: waves, rings, partitions, canaries, risk tiers. Each is a way of moving a change through the fleet one group at a time, and many teams still do it with homegrown scripts and manual steps.
In the Argo CD 2026 user survey, custom CI scripts remain the most common way teams promote changes between environments, used by 62% of respondents at companies with more than 500 people.
Part 1 of this series made the case that the fleet has replaced the cluster, and Part 2 showed where today's tools stop short of managing a fleet. This post is for platform teams delivering to a fleet, from a few dozen clusters to thousands of targets. It walks through common fleet management use cases, drawn from conversations with software teams at companies across industries.

Figure 1: The same change, before and now. Before, it moved through three environments in turn. Now it has to move through hundreds or thousands of targets in waves, starting with a small pilot group.
What are common fleet management use cases?
Four fleet management use cases show the pattern: patching a CVE across shared clusters, releasing to hundreds of customer clusters, rolling out to thousands of store clusters, and keeping regulated rollouts auditable. In each case, the change is routine but the challenge is moving it safely through the fleet.
#1 Patching a CVE across shared clusters
A platform engineer needs to patch cert-manager, the add-on that issues and renews TLS certificates, across 100 clusters to close a vulnerability. The usual approach is to update the version in a central config file and let it roll out to every cluster at once. There's no pilot or canary stage (a small first group) and no health check afterward, so a bad rollout surfaces hours or days later, when certificates start expiring.
In our conversations with platform teams, this plays out in two ways. Some teams build control themselves: they send each patch through release branches, manual gates, and manager approvals, then wait for each cluster's maintenance window and track progress across tickets, chat threads, and Git history.
Other teams want controls their tooling doesn't give them: skipping clusters in a maintenance window and coming back to them later, pausing and resuming a rollout at any time.
The cost: Until the patch lands everywhere, the vulnerability stays open on every cluster the rollout hasn't reached, and no one can say which clusters those are.
#2 Shipping a release to hundreds of customer clusters
A SaaS company runs a dedicated Kubernetes cluster for each of its 300 enterprise customers, and every cluster runs the same 100 apps. Shipping a routine release leaves two options. Release to all 300 clusters at once and lose sight of which customer's cluster broke. Or give every app on every cluster its own pipeline: 100 apps across 300 clusters is 30,000 pipelines, each with its own approval, and no single place to see where the release stands.
Teams we've talked to run into both options. One provider sends a single repository tag to every customer cluster at once, with no first group and no feedback until the release is everywhere. It wants feedback before a change merges and a release that moves one customer at a time. Another gives every customer its own maintenance window, so it has to decide the rollout order customer by customer.
The cost: Released all at once, a failure that could have been caught on one customer's cluster reaches all 300. Released per pipeline, every release means 30,000 approvals, which is not a meaningful gate in reality.
#3 Rolling out to thousands of store clusters
A retail chain runs a small cluster in each of its thousands of stores and needs to roll out a new point-of-sale build without a bug reaching the busiest stores first. The platform team sets the rollout order by hand and keeps it in a spreadsheet or hardcoded into the pipeline, so the order goes stale whenever a store opens or closes.
Teams running clusters at physical sites told us much the same. One rolls out in batches ordered by risk, from lower-stakes locations to the most critical. With no health visibility between batches, an engineer waits for services to restart cleanly before starting the next batch. Another wants ring-based releases that start with its lowest-risk sites.
The cost: Every rollout needs an engineer watching from the first store to the last, and every opening or closure means editing the pipeline. The stores at the end of the line, the busiest ones, are the ones the business can least afford to disrupt.
#4 Keeping regulated rollouts approved and auditable
A compliance lead needs every production change, across clusters in several countries and regions, signed off and recorded in an audit log. The approval process was built for a few environments, so teams either collect an approval for every cluster or can't prove which clusters got a change, when, and who approved it. At isolated sites, someone also has to load the update by hand before the rollout can start.
In our conversations with regulated teams, the strain shows up three ways. One team plans to name, for every cluster, who must approve a change to it. Another, running a large central Argo CD estate, wants ordered groups within a single production stage, Jira approval gates, and CVE checks before promotion, and calls its current auditability weak. A third can reach its production environment only through a one-way data transfer link.
The cost: Every cluster added is another approval to collect or another gap in the audit trail, and every isolated site adds manual work before a rollout can begin.
The challenges all of these use cases share
The same four challenges come up in each use case, and they share a root cause: today's tools force a choice between controlling each target and seeing the whole fleet.
Put the whole fleet in one stage, and you get one view of the rollout, but no way to order it, pause it, or stop it at the first bad target. Give each target its own stage and you get that control back, but the rollout splinters into hundreds or thousands of pipelines that no one can see as a whole. Each challenge below is one side of that trade-off, or a workaround for it.
Rollout order is managed by hand: Teams roll out changes in groups, but they decide those groups and their order outside the delivery tool, in scripts, spreadsheets, or an engineer's judgment. As the fleet changes, the target list drifts out of date.
Failures aren't contained: A release that goes everywhere at once can't stop at the first bad target. A release that goes site by site can stall when one site fails. And without rollback as a single operation, there's no clean way back from either.
There's no single view of the rollout: Tools built around one cluster or one application at a time can't show which version runs where, which targets are healthy, or what's blocked. Teams piece that answer together by hand.
Governance doesn't scale with the fleet: Approvals, ticketing, and security gates were built for a handful of environments. Run once per target, they multiply as the fleet grows, and isolated sites add manual work before a rollout can start. At fleet scale, governance needs to run once per rollout, leave one record of every target reached, and work across air-gapped sites.
These challenges map to a building block we outlined in The Fleet Is the New Cluster: The Infrastructure Shift Platform Engineers Can't Ignore
Fleet management is coming soon to Kargo Enterprise
We created Argo CD to keep clusters in sync with the desired state. We then built Kargo, our continuous promotion tool, to move releases through environments like dev, staging, and production, with gates and verification at each step.
Fleet management is the next layer: promotion that works across hundreds or thousands of targets, not only a few environments, so teams no longer have to choose between controlling each target and seeing the whole fleet. Fleet management is coming to Kargo Enterprise, with early access starting in December. Stay tuned.
Talk with an Akuity solutions architect to learn more. Contact us today.

