Launching The Akuity Agentic Control Plane Learn More →

Launching The Akuity Agentic Control Plane Learn More →

How to Scale Argo CD for Kubernetes Workloads

Blake Pettersson


By Blake Pettersson, Argo CD maintainer at Akuity. Originally published August 15, 2023. Updated October, 2026.

Scaling Argo CD starts with finding what's slowing it down. A repo-server waiting on manifest generation needs a different fix than a throttled Kubernetes API, so confirm the constraint with metrics, then apply the least costly approach that addresses it.

Many teams plan Argo CD capacity around Application count, since it's the number Argo CD surfaces most prominently. But resource churn, manifest-generation cost, repository activity, and Kubernetes API latency determine the load on each component, so two installations with the same Application count can slow down for different reasons.

Teams scaling Kubernetes tend to consolidate more clusters onto fewer Argo CD instances, and each new cluster adds load to the same components. In the 2026 Argo CD user survey, scaling and performance ranked as the top challenge teams report.

If you're a platform engineer or SRE and your production Argo CD installation has slowed under growing load, this guide is for you. I've been an Argo CD maintainer for over four years, and in that time I've helped Akuity customers and open source Argo CD users work through scaling problems in production. The approaches below are the ones I've seen work, in the order I'd try them.

The defaults, metrics, and configuration guidance below use the Argo CD v3.5 release line.

What you'll learn

  • Where each Argo CD component hits its limits. Each component slows for different reasons, from resource churn to large Application lists.

  • How to pinpoint the bottleneck before you change settings. Symptoms point to a likely component, and a small set of metrics confirms which one.

  • What to check before you add capacity. Slow dependencies and resource-starved components limit what extra replicas and concurrency can achieve.

  • Why scaling starts with reducing what Argo CD has to process. Resource exclusions, webhooks, and manifest-generate-paths cut load without adding resources.

  • When to use each of the five approaches to scaling Argo CD. Tuning, scaling, sharding, splitting, and re-architecting each address a different constraint.

What determines Argo CD scalability?

Argo CD scalability depends on how much load reaches each of its core components, and on how quickly the systems they depend on respond. Each component handles a different part of the workload, so knowing what each one does is the first step in telling which one is slowing down.

Application Controller. The Application Controller compares desired manifests with live state in each destination cluster and executes syncs. Its load grows with the number, size, and churn of managed resources, with reconciliation frequency, and with the size of the live state cache for each destination cluster.

repo-server. The repo-server clones Git repositories and generates manifests with Helm, Kustomize, or config management plugins. Repository size, file count, and Git activity drive fetch work. Generation consumes CPU, memory, and temporary disk, and the cost varies by tool. When several Applications share a repository, one commit can trigger regeneration for each of them.

ApplicationSet Controller. The ApplicationSet Controller runs generators that create and update Applications. Generator complexity, polling frequency, and SCM API latency and rate limits drive its load. The number of Applications it creates or updates per reconciliation sets the load it passes to the Application Controller.

API server (argocd-server). The API server handles requests from the UI, CLI, and API clients. Application-list and resource-tree requests drive its load independently of reconciliation, so a slow Application list can coexist with a healthy Application Controller.

Redis. Redis caches data the other components read, so its load follows their activity.

Maintainer's note: The repo-server is usually the first component to become a bottleneck as installations grow. The factor teams underestimate most is churn. A handful of resource types that change every few seconds can keep the controller busier than thousands of Applications that never change.

Each of these components also relies on systems outside Argo CD. Destination Kubernetes APIs, Git hosts, and SCM providers set limits of their own. A destination API with high latency or throttling delays reconciliation for that cluster, and Git latency or SCM rate limits slow the components that depend on them. Adding Argo CD capacity won't get you past a slow dependency.

Which Argo CD component or dependency is causing the slowdown?

When Argo CD slows down, it's hard to pinpoint which component is responsible. The symptom you see often points somewhere other than the cause. A slow repo-server can make sync state lag in the Application Controller, and a throttled Kubernetes API can make a healthy controller look saturated. The table below maps each symptom to the most likely component or dependency and the metrics that confirm it, so you can narrow the search before you change settings.

The table doesn't list target numbers, because what counts as slow depends on your installation. A repo-server wait time that's normal for one team can signal trouble for another, depending on how many Applications, clusters, and repositories each runs. The useful comparison is against your own normal, so record these metrics while Argo CD is healthy. When a slowdown starts, the change from that baseline tells you where to look.

Table: Argo CD bottleneck diagnosis by symptom


Symptom

Likely bottleneck

What to measure

Health or sync state lags

Application Controller

argocd_app_reconcile versus Kubernetes request and rate-limiter latency

One cluster is much slower than others

Destination Kubernetes API or network

argocd_kubectl_request_duration_seconds, argocd_kubectl_rate_limiter_duration_seconds, and argocd_kubectl_request_retries_total by cluster

Application appears late or not at all

ApplicationSet Controller

argocd_appset_reconcile, argocd_appset_owned_applications, and SCM latency and rate limits

Manifest refresh is slow

repo-server

argocd_git_request_duration_seconds, argocd_repo_pending_request_total, and argocd_repo_parallelism_wait_duration_seconds

Application list is slow but reconciliation is healthy

API server (argocd-server)

argocd-server latency, CPU, response size

Cache latency rises across components

Redis

argocd_redis_request_duration and argocd_redis_request_total

Application Controller

The question to answer here is whether the controller itself is slow or whether it's waiting on the Kubernetes API. argocd_app_reconcile shows how long each reconciliation takes in total. Three Kubernetes client metrics (request duration, rate-limiter duration, and retries) show how much of that time the controller spent waiting on the API. If most of the time goes to waiting, adding controller capacity won't help, and the fix is on the API side.

Maintainer's note: A “slow” controller often isn't slow at all. It's waiting on a throttled Kubernetes API client. I regularly see teams raise processor counts when the fix was client QPS and burst, or a resource exclusion that stops the controller from watching resources it doesn't manage.

An AWS scalability test backs this up. The test ran 10,000 Applications with small manifests, which kept repo-server work limited. Raising the controller's status and operation processors from 25/50 to 100/200 left sync times at around 40 minutes, while raising Kubernetes client QPS and burst cut them to roughly 11 minutes.

Destination Kubernetes API

If one cluster lags while the others stay healthy, the likely cause is that cluster's Kubernetes API or the network between it and Argo CD, not the controller. To confirm, break the same three Kubernetes client metrics (request duration, rate-limiter duration, and retries) out by cluster, and compare the slow one against the rest. Sharding won't help here. It spreads clusters across more controllers, but the controller still reaches that cluster over the same slow connection.

ApplicationSet Controller

If an Application that an ApplicationSet should have created is missing or shows up late, start with the ApplicationSet Controller. If the Application exists but its manifests, health, or sync state lag, the problem is further downstream, so look at the repo-server and Application Controller instead.

Two metrics show what the ApplicationSet Controller is doing. argocd_appset_reconcile tracks how long each reconciliation takes, and argocd_appset_owned_applications counts the Applications it manages. If your generators pull from GitHub, Argo CD v3.5 also offers GitHub API requests and rate-limit metrics, which show whether GitHub is throttling the controller. These are off by default, so you'll need to enable them.

repo-server

If manifests take a long time to refresh after a commit, the repo-server is the likely bottleneck. It can be slow for three different reasons, and each has its own metric:

  • Waiting for a generation slot. The repo-server limits how many manifests it generates at once. argocd_repo_parallelism_wait_duration_seconds shows how long requests wait for a free slot. High values mean more requests are arriving than the repo-server can process in parallel.

  • Waiting for a repository lock. Some requests need exclusive access to a repository, so they take turns. argocd_repo_pending_request_total counts the requests waiting for that access. High values mean requests are lining up behind each other.

  • Waiting on Git. argocd_git_request_duration_seconds shows how long Git fetches take. High values point to a slow Git host or network rather than to the repo-server.

Each reason calls for a different fix, covered in How to choose the right approach to scaling Argo CD.

API server (argocd-server)

If the Application list in the UI loads slowly while syncs and health checks keep up, the API server is the likely cause. Check its latency, CPU, and response size. The API server doesn't hold state, so you can add replicas to spread requests across more pods. Extra replicas won't shrink a large Application list, though. Argo CD v3.5 sends the full list to the browser and splits it into pages there, so a large installation means large responses and more work in the browser. The Akuity Platform enables server-side pagination by default, which returns the list in smaller pieces.

Redis

If cache latency rises across several components at once, check Redis. Two metrics show what's happening: argocd_redis_request_duration tracks how long cache requests take, and argocd_redis_request_total counts them. A rise in both can mean Redis is short on capacity, or that the other components are sending it more requests than usual. To tell which, compare the Redis metrics with reconciliation activity and resource-event volume. If those rose at the same time, Redis is absorbing extra load from elsewhere, and the fix belongs with the component generating it.

How to choose the right approach to scaling Argo CD

The right approach depends on the constraint your metrics identify. There are five common approaches to scaling Argo CD, ranging from configuration changes inside a single instance to changing how Argo CD runs across your clusters. Each takes more time and operational work than the one before it, so start at the top of the list and weigh the expected gain against the cost at each step.

  1. Tune the Argo CD instance when the bottleneck is configuration, avoidable reconciliation, or a component short on CPU, memory, or disk.

  2. Scale Argo CD components when a component stays saturated after tuning and its dependencies still have spare capacity.

  3. Shard the Argo CD Application Controller when controller load stays high across several destination clusters.

  4. Split Argo CD into separate instances when ownership, isolation, or network paths call for independent control planes.

  5. Re-architect the Argo CD topology when latency or coordination problems return after the first four approaches, or when instances and shards have become the overhead.

1. Tune the Argo CD instance

Tuning covers three checks. Work through them before you add capacity, since more concurrency or replicas on a component that's short on resources, or that's doing work it doesn't need to do, raises cost without fixing the slowdown.

Confirm the dependency has room. If the Kubernetes API is throttling requests, Git is slow, a manifest tool is expensive, or the SCM provider is rate limiting, extra Argo CD capacity won't speed things up. Check the dependency metrics from the diagnosis table first.

Give starved components the resources they need. Sustained CPU throttling, memory pressure or OOMKills (pods killed for exceeding their memory limit), and full temporary disk slow a component regardless of its settings. Fix these before you run more work concurrently.

Remove work Argo CD doesn't need to do.

  • Resource exclusions stop Argo CD from watching resource types it doesn't manage, which reduces controller load.

  • Ignored resource updates stop the controller from reprocessing changes that don't affect sync, such as status-only updates on noisy CRDs.

  • Webhooks trigger a refresh when a commit lands, so Argo CD picks up changes without waiting for the next polling interval, and you can lengthen that interval. Refresh jitter spreads timed refreshes out so they don't all run at once.

  • For monorepos, the argocd.argoproj.io/manifest-generate-paths annotation limits regeneration to the Applications whose files changed. It works when its configured paths include every file that affects the Application's manifests. Shallow clones reduce clone work by limiting Git history, but they disable the non-webhook Git-history comparison the annotation relies on. Webhook filtering still works with shallow clones.

Maintainer’s note: If I had to pick one change with the biggest payoff: resource exclusions. In my experience, excluding high-churn resource types Argo CD doesn't manage, and ignoring status-only updates on noisy CRDs, often cuts controller load more than added capacity would.

2. Scale Argo CD components

When a component stays saturated after tuning and its dependencies still have capacity, add concurrency or replicas to that component.

repo-server. If requests are waiting for a generation slot and the pod has spare CPU and memory, raise --parallelismlimit so it generates more manifests at once. If requests still exceed capacity, add replicas. Extra replicas won't make a slow Git fetch or an expensive generator faster, and they won't clear repository-lock waits. In Argo CD v3.5, Helm generation for Applications that share a repository directory runs concurrently by default, but check plugin and repository-lock constraints before you raise concurrency.

ApplicationSet Controller. Check generator polling intervals, webhook delivery, and SCM provider limits before you raise applicationsetcontroller.concurrent.reconciliations.max. Faster generation creates and updates more Applications, which adds load on the Application Controller, so watch its reconciliation latency after the change.

Application Controller. Argo CD v3.5 defaults to 20 status processors and 10 operation processors (controller.status.processors and controller.operation.processors). Raise them when the controller has spare resources and the destination APIs can handle more concurrent requests. Compare Kubernetes rate-limiter latency before and after. If it rises, the API has become the constraint, as the AWS test showed.

API server. When the API server is saturated, add replicas to spread requests across more pods.

3. Shard the Argo CD Application Controller

Sharding splits destination clusters across multiple Application Controller replicas. Each cluster, and the Applications that target it, stays with one shard. Shard when controller load stays high across several clusters after you've resolved resource starvation and API bottlenecks.

Maintainer's note: The most common mistake I see is sharding before fixing the underlying churn, which spreads the same waste across more replicas. The other is expecting sharding to help a single oversized cluster. Argo CD assigns shards per cluster, so that cluster's load lands on one controller regardless.

When the topology itself is the constraint

Approaches 4 and 5 change how Argo CD is deployed rather than how it's configured. They matter more as organizations connect more clusters to fewer instances: between 2025 and 2026, the share of Argo CD instances managing a single cluster fell from 25% to 19%. Three signals tell you tuning, scaling, and sharding have run their course:

  • Shard rebalancing keeps coming back. If you rebalance shards each time you add a cluster, the way clusters are assigned to controllers is the constraint, not the controllers' capacity.

  • Latency follows distance. If reconciliation is slowest for the clusters farthest from Argo CD, the constraint is network distance to remote Kubernetes APIs, not component saturation.

  • One failure would stop unrelated teams. If a single instance going down would halt deployments for teams that don't otherwise work together, the risk is the shared failure domain, and more capacity won't reduce it.

4. Split Argo CD into separate instances

Splitting gives teams or environments their own Argo CD instance. It fits when ownership, isolation, or blast radius drives the need, or when some clusters sit on network paths too slow or unreliable to manage from a central instance. Deutsche Telekom runs one Argo CD instance for non-production and one for production across roughly 500 microservice repositories, a split driven by security guidelines rather than capacity.

Each instance you add brings its own configuration, credentials, upgrades, observability, and incident response. Weigh that operating cost alongside the performance and isolation gains.

Maintainer's note: Teams rarely split control planes for performance alone. Security and blast radius usually drive that decision, and any performance gain is a welcome side effect. I'd be cautious about splitting purely to fix a slowdown. Work through tuning, scaling, and sharding first, before you reach for that hammer.

5. Re-architect the Argo CD topology

Re-architecting changes where Argo CD's components run. It fits when remote API latency or coordination problems return after tuning, sharding, and splitting, or when the instances and shards you already run have become the overhead.

Figure: Argo CD’s Open-source topologies

Figure: Argo CD’s Open-source topologies

Four topologies handle these constraints differently, and the Argo CD architecture comparison covers each in more depth:

  • Centralized (Single Argo CD instance). All components run in one cluster and reach destination clusters over the network. The tuning guidance above assumes this topology.

  • Per-cluster. Each cluster runs its own instance. This removes remote API latency and creates hard failure boundaries, but you operate and upgrade an instance per cluster.

  • Hybrid. The API and UI stay central, while reconciliation runs closer to the clusters.

  • Agent-based. One control plane stays central, and the Application Controller moves into each managed cluster. This removes the remote API path without multiplying the instances you operate.


Figure: Akuity’s Argo CD Agent-Based Architecture

Figure: Akuity’s Argo CD Agent-Based Architecture

Agent-based addresses both triggers for re-architecting at once. It removes the remote API path, and it avoids the instance sprawl that per-cluster and split setups create. The open-source Argo CD Agent, an argoproj-labs project, is still early. It hasn't reached a 1.0 release, and failover support arrived as a beta feature in April 2026.

Where to start scaling Argo CD

Scaling Argo CD comes down to finding the constraint, removing load Argo CD doesn't need to carry, and adding capacity only where the metrics point. In my experience, many of the slowdowns teams bring to me come down to churn or a throttled API client rather than a shortage of replicas. Each approach in this guide works, but each takes time and operational effort, and that effort grows as your installation does.

For organizations that want to move fast and keep platform teams focused on their own roadmap, the Akuity Platform takes on that effort. Akuity operates the Argo CD control plane and has re-architected Argo CD around an agent-based model that removes remote Kubernetes API latency, avoids instance sprawl, and keeps each cluster's API endpoint private.

Organizations run the Akuity Platform at scale today, with 1 million + Applications across 4,000+ clusters under management. Mastery moved from several self-managed Argo CD instances with sharded controllers to the Akuity Platform and now runs 40,000+ Applications across 50+ clusters.

Book a demo or start a free trial to see how the Akuity Platform runs your workloads.

Frequently asked questions about scaling Argo CD

How many Applications can one Argo CD instance handle?

Argo CD has no fixed Application-count limit. The number of resources each Application manages, how often they change, how expensive the manifests are to generate, and how fast Kubernetes APIs and Git respond determine the load behind that count. To estimate capacity, test a workload that resembles yours and measure reconciliation time, manifest generation time, and API latency.

When should the Application Controller be sharded?

Shard when controller load stays high across several destination clusters after you've fixed resource shortages and Kubernetes API bottlenecks. Argo CD assigns each destination cluster to one controller replica, so the Applications targeting a cluster stay together. If only one cluster is slow, sharding won't help. Investigate that cluster's Kubernetes API and network instead.

Do more repo-server replicas fix slow manifest generation?

Sometimes. Extra replicas help when requests are waiting for a free generation slot and the work can run in parallel. They won't help with repository-lock waits or slow Git fetches, which need their own diagnosis, and an expensive generator still does the same work for each request.

What is agent-based Argo CD?

Agent-based Argo CD keeps one central control plane for the API and UI, and runs the Application Controller inside each managed cluster. Each controller reconciles against its own cluster's local Kubernetes API, which removes the latency of reaching remote clusters over the network without the overhead of running a full Argo CD instance per cluster. The Akuity Platform (Akuity Deploy) runs this model in production today. The open source Argo CD Agent, an argoproj-labs project, has not yet reached a 1.0 release.

Blake Pettersson is an Argo CD maintainer at Akuity. He helped lead the work that made OCI registries a first-class source in Argo CD, bringing GitOps to disconnected and regulated environments, and he works with Akuity customers on running Argo CD at scale.

Ready to simplify delivery with Akuity?

Deploy, promote, and operate applications reliably, powered by OSS you trust and Intelligence you control.

Ready to simplify delivery with Akuity?

Deploy, promote, and operate applications reliably, powered by OSS you trust and Intelligence you control.

Ready to simplify delivery with Akuity?

Deploy, promote, and operate applications reliably, powered by OSS you trust and Intelligence you control.

Sign Up for Akuity Updates

Practical guidance on MTTR reduction, GitOps at scale, and safe automation, with product updates from the Argo CD and Kargo team.

@2026 Akuity Inc. All rights reserved.

Akuity Inc. 440 N. Wolfe Road, Sunnyvale, CA 94085-3869 US +1-510-771-7837

SOC2 Type 2 Compliant

Sign Up for Akuity Updates

Practical guidance on MTTR reduction, GitOps at scale, and safe automation, with product updates from the Argo CD and Kargo team.

@2026 Akuity Inc. All rights reserved.

Akuity Inc. 440 N. Wolfe Road, Sunnyvale, CA 94085-3869 US +1-510-771-7837

SOC2 Type 2 Compliant

Sign Up for Akuity Updates

Practical guidance on MTTR reduction, GitOps at scale, and safe automation, with product updates from the Argo CD and Kargo team.

@2026 Akuity Inc. All rights reserved.

Akuity Inc. 440 N. Wolfe Road, Sunnyvale, CA 94085-3869 US +1-510-771-7837

SOC2 Type 2 Compliant