A Kubernetes Engineer's Hands-On Guide to Argo CD
In Kubernetes operations, many teams begin their automation journey with push-based CI/CD: after building a container image, the CI runner directly executes kubectl apply or helm upgrade inside the pipeline to push changes into the cluster.
As systems scale, significant pain points emerge: CI runners must hold high-privilege cluster credentials, configuration drift from manual kubectl edit operations goes untracked, and deleting YAML manifests often leaves orphaned resources lingering in the cluster.
Pull-based GitOps fundamentally inverts this workflow. An in-cluster controller actively compares the target desired state declared in Git repositories against the live operational state of the cluster, automatically reconciling any drift. The CI pipeline’s responsibility is streamlined to building images and committing changes to Git, completely eliminating direct access to the Kubernetes API.
NOTE
Further reading: A Practical Guide to GitOps for Kubernetes Users
“Git is the single source of truth; the cluster is merely a live projection of Git.”
Based on the Argo CD v3.5.3 stable release, this guide skips basic Kubernetes concepts to dive straight into Argo CD’s core architecture, state synchronization and self-healing strategies, diff tuning pitfalls, and production-grade security boundaries.
Argo CD Core Architecture and Component Roles
The Three Core Components
Argo CD’s core architecture relies on three primary components working in concert:
- API Server (
argocd-server): Exposes gRPC/REST API endpoints for the Web UI, CLI, and external integrations. Handles application lifecycle management, RBAC, credential storage (via Kubernetes Secrets), and inbound Git webhook events. - Repository Server (
argocd-repo-server): Maintains a local cache of Git repositories. Given a repo URL, target revision, directory path, and parameters, it renders Helm templates or runs Kustomize builds to generate raw Kubernetes manifests. - Application Controller (
argocd-application-controller): A Kubernetes custom controller that continuously compares live cluster state against the Git target state, calculates diffs, drives reconciliation, and executes lifecycle hooks across PreSync, Sync, and PostSync phases.
+------------------------------------+
| Developer / CI |
+------------------------------------+
|
| commit
v
+------------------------------------+
| Git Config Repo |
+------------------------------------+
|
| polling (120s+jitter)
| or Git Webhook
v
+------------------------------------+
| Kubernetes Cluster |
| |
| +--------------------------+ |
| | argocd-repo-server | |
| | (render manifests) | |
| +--------------------------+ |
| | |
| v |
| +--------------------------+ |
| | application-controller | |
| | (compare Live/Target) | |
| +--------------------------+ |
| | |
| v |
| +--------------------------+ |
| | Live K8s Objects | |
| +--------------------------+ |
+------------------------------------+
In addition, standard installation manifests include argocd-applicationset-controller for generating multi-application setups, argocd-redis for caching rendered manifests, and auxiliary components for SSO and notifications.
Sync Status vs. Health Status
When assessing application state, it is critical to understand that Sync Status and Health Status are orthogonal dimensions (see Resource Health Assessment):
- Sync Status (
Synced/OutOfSync): Indicates whether cluster object definitions strictly match what is declared in Git. - Health Status (
Healthy/Progressing/Degraded): Indicates whether cluster resources like Pods and Services are functioning properly at runtime.
A service stuck in CrashLoopBackOff due to a corrupted container image can easily be Synced yet Degraded in Argo CD.
Core Custom Resources: Application and AppProject Security Boundaries
Declarative management in Argo CD is governed by two core Custom Resource Definitions (CRDs):
- Application defines what gets deployed where—each Application maps a Git source path to a destination cluster and namespace, serving as the fundamental deployment unit.
- AppProject provides the security perimeter for Applications—it restricts which Git repositories Applications can pull from, which clusters and namespaces they can target, and which Kubernetes resource kinds they are permitted to manage.
Their relationship is one of strict containment: every Application must belong to exactly one AppProject, which defines its operating scope.
+----------------------------------------------------+
| AppProject (e.g., team-web) |
| - Allowed Source Repos |
| - Destination Clusters & Namespaces |
| - Resource Whitelist / Blacklist |
| - Project-level RBAC Roles |
| |
| +--------------------------------------------+ |
| | Application (e.g., guestbook-prod) | |
| | - spec.source (repo, path, targetRevision) | |
| | - spec.destination (server, namespace) | |
| | - spec.syncPolicy (automated, retry) | |
| +--------------------------------------------+ |
+----------------------------------------------------+
Application Specification and Best Practices
The Application CRD declares the source and destination for a single deployment:
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
name: guestbook-staging
namespace: argocd
finalizers:
- resources-finalizer.argocd.argoproj.io
spec:
project: team-web
source:
repoURL: https://github.com/example-org/k8s-config.git
targetRevision: main
path: apps/guestbook/overlays/staging
destination:
server: https://kubernetes.default.svc
namespace: guestbook-staging
When configuring Applications, keep three critical rules in mind:
- Namespace Constraints: By default,
ApplicationandAppProjectmanifests must be created in the namespace where Argo CD is installed (typicallyargocd). - Mutually Exclusive Destination Fields: In
spec.destination, specify eitherserverorname—defining both triggers a schema parsing error. - Cascading Deletion with Finalizers: If
resources-finalizer.argocd.argoproj.iois omitted, deleting anApplicationleaves its underlying Deployments and Services untouched in the cluster. Adding this finalizer ensures cascading deletion.
AppProject Multi-Tenant Security Boundaries
AppProject is the foundational isolation boundary when sharing an Argo CD instance across multiple teams. Applications without an explicit project assignment fall into the default project, which permits any source repository, any destination cluster, and all resource kinds by default (see AppProject Multi-Tenancy)—a critical security vulnerability in production environments.
apiVersion: argoproj.io/v1alpha1
kind: AppProject
metadata:
name: team-web
namespace: argocd
spec:
description: Web Team Project Boundary
sourceRepos:
- https://github.com/example-org/k8s-config.git
destinations:
- server: https://kubernetes.default.svc
namespace: guestbook-*
clusterResourceWhitelist:
- group: ''
kind: Namespace
namespaceResourceBlacklist:
- group: ''
kind: ResourceQuota
- group: networking.k8s.io
kind: NetworkPolicy
Security Pitfall: If an AppProject’s destinations whitelist allows deploying to the argocd namespace, workloads under that project could modify Argo CD’s own configuration and escalate privileges across the entire platform. Keep the deployable namespace whitelist as restrictive as possible.
Sync Policies, Self-Healing, and Adoption Stages
spec.syncPolicy controls the automated behavior of the controller when it detects state drift (see Automated Sync and Self-Healing Policies).
spec:
syncPolicy:
automated:
enabled: true
prune: true
selfHeal: true
allowEmpty: false
syncOptions:
- CreateNamespace=true
- PruneLast=true
retry:
limit: 5
backoff:
duration: 5s
factor: 2
maxDuration: 3m
Core Automated Sync Switches
automated.enabled: Enables automated synchronization. When set tofalse, none of the other options take effect.prune: Disabled by default. When enabled, resources removed from Git are pruned (deleted) from the cluster.allowEmpty: Defaults tofalse, preventing accidental cluster-wide resource purges if a misconfigured path causes Git manifests to render as empty.selfHeal: Disabled by default. When enabled, manual drift introduced viakubectl editor direct cluster modifications is automatically reconciled back to the desired Git state within seconds.retry: Retries failed sync attempts using an exponential backoff algorithm. Note: without retry configured, a failed sync will not re-trigger automatically for the same commit.
Recommended Phased Rollout for Production
In a GitOps workflow, human gates shift to Git: CI opens a PR against the Config Repo, which is reviewed before merging.
Once the PR is merged, Argo CD detects the difference. If automated sync (automated: true) is enabled, changes apply immediately; with manual sync, an operator must explicitly trigger a sync via the UI or CLI.
When adopting GitOps, avoid enabling full automation in production on day one. A four-phase rollout ensures smooth onboarding:
- Phase 1: Manual Verification: Keep syncs manual for the first 1–2 weeks. Let operations teams observe diffs, get comfortable with
OutOfSyncalerts, and fix static manifest declarations. - Phase 2: Automated Non-Production: Enable
automatedandselfHealin development environments, reinforcing the cultural rule that “all hotfixes must be committed to Git.” - Phase 3: Prune Activation: Enable
prune: truein non-production once repo structures stabilize and naming conventions are verified. - Phase 4: Production Trade-offs: Enable automated sync in production once rigorous PR review processes are established; for scheduled maintenance windows, configure
Sync Windowsto restrict deployment timing.
Precision Orchestration: Sync Waves and Hooks
When deployments require strict sequencing—such as running database migrations before updating web services or ensuring CRDs exist before their Custom Resources (CRs)—Argo CD relies on Sync Phases (Hooks) and Sync Waves (see Sync Phases and Sync Waves).
Phase: PreSync
+-----------------------------+
| wave -1: DB Migration Job |
| (hook: PreSync) |
+-----------------------------+
|
v
Wait for Success
|
v
Phase: Sync (Main Apply)
+-----------------------------+
| wave 0: ConfigMap / Secret |
+-----------------------------+
|
v
+-----------------------------+
| wave 1: App Deployment |
+-----------------------------+
Hook Phases and Deletion Policies
Adding the argocd.argoproj.io/hook annotation to resource metadata binds that resource to a specific lifecycle phase:
| Hook Phase | Execution Timing | Typical Use Cases and Behavior |
|---|---|---|
PreSync | Before main manifests are applied | Runs DB migrations; sync halts immediately on failure to block the deployment |
Sync | Concurrently with main manifests | Applies standard resources; further ordered using sync-wave |
PostSync | After all resources are applied and Healthy | Triggers deployment notifications, cache warming, or integration tests |
SyncFail | When a sync operation fails | Runs cleanup tasks or triggers alerting workflows |
Hook cleanup is governed by argocd.argoproj.io/hook-delete-policy. The default value is BeforeHookCreation, which removes previous hook resources before launching a new one. The Job from the current run is kept until the next sync, success or failure alike, so failed migration Jobs stay in the cluster for log inspection and debugging. With HookSucceeded, the Job is deleted once it succeeds.
Sync Waves and Ordering Rules
argocd.argoproj.io/sync-wave supports integers (including negative values), where lower numbers execute first.
The overall sync sequence follows: Phase → Wave (ascending) → Resource Kind (Namespaces before other resources) → Resource Name. Argo CD waits until all resources in the current wave reach Healthy before proceeding to the next wave. During pruning, the order is reversed (higher waves are pruned first).
Practical Design Guidelines:
- Migration scripts executed in
PreSyncmust be idempotent and backward-compatible. Between PreSync completion and new Pod readiness, old Pods continue serving live traffic. Furthermore, reverting application code will not automatically roll back database schemas. - If a deployment includes both CRDs and CRs, isolate CRDs in a negative wave and configure
SkipDryRunOnMissingResource=trueto prevent dry-run validation failures before schemas are registered.
Example: Running DB Migrations Before Web Service Updates
Suppose a web application requires database schema migrations before deploying a new release. The orchestration is straightforward: place the Migration Job in PreSync to guarantee successful execution before main manifests apply, and position the Web Deployment in Sync wave 1 so it updates only after foundational resources (wave 0) are ready.
# db-migrate.yaml — PreSync Hook
apiVersion: batch/v1
kind: Job
metadata:
name: db-migrate
annotations:
argocd.argoproj.io/hook: PreSync
argocd.argoproj.io/sync-wave: "-1"
argocd.argoproj.io/hook-delete-policy: BeforeHookCreation
spec:
template:
spec:
containers:
- name: migrate
image: example-org/web-app:v2.3.0
command: ["python", "manage.py", "migrate", "--no-input"]
restartPolicy: Never
backoffLimit: 0
---
# web-deployment.yaml — Main Sync Resource
apiVersion: apps/v1
kind: Deployment
metadata:
name: web-app
annotations:
argocd.argoproj.io/sync-wave: "1"
spec:
replicas: 3
template:
spec:
containers:
- name: web
image: example-org/web-app:v2.3.0
The execution flow proceeds as follows: PreSync wave -1 (Migration Job) → Wait for Job success → Sync wave 0 (ConfigMap / Secret) → Sync wave 1 (Web Deployment). If the Migration Job fails, synchronization halts immediately: the existing Pods keep serving, and the new code never touches an incompatible schema.
Image Updates and Git Repository Structure
Under GitOps, after CI builds a new container image, it must update the Git repository via a well-defined automation strategy.
Separating Config Repos from Source Repos
The established best practice is to isolate Kubernetes manifests in a dedicated Config Repository (also referred to as a GitOps repo, deploy repo, or infra repo), separate from the application Source Repository (App Repo).
This separation prevents configuration changes from triggering expensive application rebuilds and cleanly separates developer commit permissions from production deployment privileges.
For container image tagging: never use latest in production. Manifests at a given Git revision must remain immutable; using latest destroys the deterministic link between Git history and cluster runtime state.
NOTE
Further reading: OCI 101: The Common Language of the Container World
Image Update Strategies Compared
| Mechanism | How It Works | Advantages | Limitations & Trade-offs |
|---|---|---|---|
| CI Commits to Config Repo | CI scripts modify the image tag in overlays and commit or open a PR directly | Transparent workflow, zero extra operational overhead, ideal for most teams | Requires CI credentials for the Config Repo; must handle concurrent push conflicts |
| Argo CD Image Updater | A standalone controller monitors container registries and writes new tags back to Git | CI does not require access to the Config Repo | Requires configuring git write-back mode; officially discouraged for critical production workloads |
Most teams should start with CI committing to the Config Repo—direct commits in dev, automated PRs for review in prod. Dedicated progressive delivery platforms (like Kargo) can be evaluated when environment promotion rules become extremely complex.
Recommended Repository Structure
Below is the standard Config Repo pattern in the Argo CD community—using Kustomize base and overlays to delineate environments, with Application and AppProject manifests centralized in an argocd/ directory:
k8s-config/
+-- apps/
| +-- guestbook/
| +-- base/
| | +-- kustomization.yaml
| | +-- deployment.yaml
| | +-- service.yaml
| +-- overlays/
| +-- dev/
| | +-- kustomization.yaml
| +-- prod/
| +-- kustomization.yaml
+-- argocd/
+-- projects/
| +-- team-web.yaml
+-- applications/
+-- guestbook-dev.yaml
+-- guestbook-prod.yaml
Key Architectural Takeaways:
- Directory-based over branch-based environments: Track all environments on a single
mainbranch using Kustomize overlays. Branch-based environments inevitably devolve into cherry-picking and merge conflicts. - Version-controlled Application declarations: Placing Application and AppProject manifests in
argocd/makes platform configuration auditable and reproducible. As the number of applications grows,ApplicationSetwith the Git Directory Generator can automatically scan directories to generate Applications.
Rollback Strategies: The Idiomatic Path vs. Emergency Controls
In a GitOps architecture, the standard operating procedure for rollbacks differs fundamentally from traditional workflows.
Idiomatic Rollbacks: git revert
Because Git is the single source of truth, the idiomatic way to roll back is to run git revert on the problematic commit in the Config Repo, followed by merging the PR into main:
git revert <faulty-commit-sha>
# Open PR, review, and merge
This approach preserves a complete audit trail and integrates seamlessly with automated sync without creating divergence between the cluster and Git.
UI/CLI Rollback Limitations and Emergency Procedures
Argo CD’s UI Rollback works completely differently from git revert—it temporarily instructs the controller to sync to an older Git revision without modifying Git itself.
If automated sync or selfHeal is active, the controller will quickly pull the cluster back to Git HEAD (the faulty version), overwriting the rollback. UI rollback is an emergency stopgap, not a permanent fix.
Official documentation explicitly notes that applications with automated sync enabled cannot be rolled back directly via the UI/CLI.
WARNING
UI Rollback is only an emergency stopgap: with automated or selfHeal enabled, the controller pulls the cluster back to Git HEAD within seconds and overwrites the rollback. Always pause automated sync before using it.
If manual intervention via the UI is required during an incident, follow this four-step emergency procedure:
- Disable Automated Sync: Run
argocd app set <APP> --sync-policy noneto halt automated reconciliation. - Execute Emergency Rollback: Trigger rollback via UI or
argocd app rollback <APP>to restore a known-good revision. - Commit the Git Revert: Afterward, be sure to commit a
git revertin the Git repository so the Git target state and the cluster’s operational state re-align. - Re-enable Automated Sync: Once synchronized, re-enable automated sync to maintain GitOps governance.
Diff Calculation and Deep OutOfSync Tuning
A common frustration when adopting Argo CD is an application reporting OutOfSync immediately after a successful sync, or the controller entering an endless reconcile loop with another operator.
Common Sources of State Drift
- HPA Replicas Management: A Deployment declares
replicas: 3, while an HPA dynamically scales it to 6. Argo CD detects the difference and marks itOutOfSync. IfselfHealis active, the controller reverts it to 3, only for the HPA to scale it back to 6. Official Best Practice: Remove thereplicasfield from Deployment manifests when managed by an HPA. - Admission Webhooks and Sidecar Injection: Init containers injected by service meshes or
caBundlefields populated by cert-manager create comparison discrepancies. - Value Normalization: a value declared as
1000min YAML is stored by the API server as1, or3072Mias3Gi, causing literal comparison mismatches.
ignoreDifferences and Critical Settings
The rule of thumb for handling discrepancies: fix manifest specifications first; only use ignoreDifferences for unavoidable mutations by third-party controllers:
spec:
ignoreDifferences:
- group: apps
kind: Deployment
name: guestbook
jsonPointers:
- /spec/replicas
- group: apps
kind: Deployment
jqPathExpressions:
- .spec.template.spec.initContainers[] | select(.name == "istio-init")
syncPolicy:
syncOptions:
- RespectIgnoreDifferences=true
WARNING
ignoreDifferences by default only affects diff calculations, not what gets applied during sync: the controller will still attempt to apply Git values. You must explicitly set syncOptions: [RespectIgnoreDifferences=true] to skip designated fields during sync.
Server-Side Diff and Server-Side Apply
Argo CD historically used Legacy 3-way diffing. Since v3.1.0, Server-Side Diff is officially Stable.
Server-Side Diff performs a dry-run Server-Side Apply (SSA) against the API Server, comparing returned predictions with live state.
Its core advantage: admission webhooks and schema defaults participate natively in the diff calculation, and SSA’s fieldManager ownership tracking greatly reduces the false positives caused by webhook field mutations.
In practice, Server-Side Diff is commonly paired with Server-Side Apply:
apiVersion: argoproj.io/v1alpha1
kind: Application
metadata:
annotations:
argocd.argoproj.io/compare-options: ServerSideDiff=true
spec:
syncPolicy:
syncOptions:
- ServerSideApply=true
Secret Management Best Practices
In GitOps, “all configuration in Git” never means storing plaintext Kubernetes Secrets (base64 is encoding, not encryption). The official Secret Management Guide divides practices into two camps:
[Git Repo: Encrypted CR / External Ref]
|
| (Argo CD syncs CR as is, no plaintext exposure)
v
[K8s Cluster: Operator / ESO pulls secret from Vault/KMS]
|
v
[K8s Secret Object Created In-Cluster]
Pattern 1: In-Cluster Decryption (Strongly Recommended)
Git stores only encrypted Custom Resources or external key references, while operators inside the destination cluster generate the actual Secret objects:
| Tool | Git Manifest Contents | Ideal Use Cases |
|---|---|---|
| External Secrets Operator (ESO) | ExternalSecret (references external key vaults) | Centralized key rotation across cloud vaults (AWS Secrets Manager / GCP Secret Manager / HashiCorp Vault) |
| Sealed Secrets | SealedSecret (encrypted with cluster public key) | Lightweight local decryption without external secret managers |
| Secrets Store CSI Driver | SecretProviderClass (volume mounts) | Sensitive credentials that should avoid being stored as Secret objects in etcd |
Architectural Advantage: Argo CD never touches plaintext secrets, fully decoupling secret rotation from application sync lifecycles.
Pattern 2: Repo-Server Injection at Render Time (Potential Anti-Pattern)
This pattern injects secrets at manifest render time via Config Management Plugins (e.g., argocd-vault-plugin). Risks: Rendered plaintext manifests reside in the argocd-redis cache and can be read via the gRPC API. Unless required by legacy constraints, new projects should use in-cluster decryption exclusively.
Production Pitfalls and Troubleshooting Checklist
When running Argo CD in production, keep this troubleshooting checklist handy:
- Application deleted but resources remain in cluster: Verify whether
metadata.finalizersincludesresources-finalizer.argocd.argoproj.io. - Delayed deployments after pushing to Git: Argo CD polls Git every 120 seconds plus random jitter (up to 180s). Configure Git webhooks for real-time responsiveness.
- CR returns
the server could not find the requested resource: If the CRD is generated dynamically during the same sync by a third-party controller, addargocd.argoproj.io/sync-options: SkipDryRunOnMissingResource=trueto the CR. - Application stuck in
Progressingstatus: Certain Ingress controllers or StatefulSets using theOnDeletestrategy do not populate expected status fields. Adjust controller configuration or define custom Resource Health Lua scripts. Manifest generation error (cached)message: Failed rendering attempts are cached in Redis to prevent retry storms. Click Hard Refresh in the UI (or runargocd app get <APP> --hard-refresh) and inspectargocd-repo-serverPod logs.helm lsshows no deployed charts: Argo CD useshelm templateto render static YAML and applies resources directly; it does not maintain Helm release records in the cluster. This is expected by design.
Conclusion: Making the Cluster a True Projection of Git
The promise of Argo CD is straightforward: keep cluster state continuously aligned with Git.
Achieving this requires establishing the discipline of committing every change to Git before progressively turning on features like selfHeal and prune. Rushing to enable full automation too early risks eroding team trust after the first accidental deletion.
Once the workflow is dialed in, production deployments become quiet, uneventful affairs triggered by a merged pull request—boring and predictable, exactly as infrastructure should be.
NOTE
Further reading: A Practical Introduction to Continuous Delivery