Unreleased documentation. Choose your installed release in the version menu. Features described here may be absent from that release.
High availability¶
The chart starts two controller replicas and two HAProxy replicas by default. Spread each pair across nodes so a node failure leaves a working controller and proxy; see anti-affinity. Each controller needs the full controller resource budget.
What happens when a controller fails¶
Leader election uses a Kubernetes Lease to choose the replica that deploys configuration. Every replica watches resources, renders changes, and serves admission requests. Followers keep their caches and render state ready for takeover.
If the leader fails, a follower acquires the lease and renders the current state before deploying. A voluntary handoff releases the lease without waiting for it to expire. HAProxy continues serving its existing configuration during election.
Configuration¶
Add settings to your complete Helm values file
and apply it with helm upgrade.
Leader election defaults¶
Leader election is enabled by default in the Helm chart:
# values.yaml (chart defaults)
controller:
replicaCount: 2 # Run 2 replicas for HA
config:
controller:
leaderElection:
enabled: true
leaseName: "" # Defaults to the Helm release fullname
leaseDuration: 30s # Lease-expiry delay before takeover attempts
renewDeadline: 20s # Leader retries renewal for this long
retryPeriod: 5s # Interval between renewal attempts
Note
The defaults allow brief API-server and CPU stalls without surrendering leadership. Shorter lease and renewal periods can reduce failover time but make the controller more sensitive to those stalls.
Disable leader election¶
A single replica can keep leader election enabled. If you choose to disable it,
keep controller.replicaCount: 1 and autoscaling off so two controllers can't
deploy independently:
Timing parameters¶
The timing parameters control failover speed and tolerance:
| Parameter | Chart default | Purpose | Recommendations |
|---|---|---|---|
leaseDuration |
30s |
Lease-expiry delay before takeover attempts | Increase for flaky networks (60s+) |
renewDeadline |
20s |
How long leader retries before giving up | Must be < leaseDuration |
retryPeriod |
5s |
Interval between leader renewal attempts | Should be < renewDeadline |
After a crash, a follower waits leaseDuration from its last observation of a
lease renewal before attempting election. Retry jitter and API latency add to
that delay; leaseDuration + retryPeriod isn't a hard upper bound. A voluntary
handoff releases the lease without waiting for expiry. HAProxy keeps serving its
current configuration while a new leader is elected.
Deployment¶
Standard high-availability Deployment¶
The default Helm installation already enables leader election and starts two controller replicas.
Scaling¶
Set the replica count in your Helm values and apply them through your normal deployment workflow:
Keep at least two replicas for failover. Adding replicas increases the total resource reservation; it doesn't reduce the work each replica performs.
Autoscaling¶
The chart ships an optional HorizontalPodAutoscaler:
controller:
autoscaling:
enabled: true
minReplicas: 2 # keep at least 2 for failover
maxReplicas: 10
targetCPUUtilizationPercentage: 80
Adding controller replicas increases admission capacity and maintains more warm
standbys. Each replica renders, but only the leader deploys; replicas don't divide
one rendering workload between them. To increase traffic capacity, scale HAProxy
with haproxy.keda or haproxy.replicaCount.
The chart grants the Lease permissions needed for leader election.
Monitoring leadership¶
Check current leader¶
The Lease resource is named after the chart fullname — haptic for helm install haptic …, but <release>-haptic when the release name doesn't contain haptic. Override by setting controller.config.controller.leaderElection.leaseName.
# List leases in the release namespace
kubectl get lease -n haptic
# View the lease for the haptic release
kubectl get lease -n haptic haptic -o yaml
# Output shows current leader:
# spec:
# holderIdentity: haptic-controller-7d9f8b4c6d-abc12
View leadership status in logs¶
# Leader logs show:
kubectl logs -n haptic deployment/haptic-controller | grep -E "leader|election"
# Example output:
# level=INFO msg="Leader election started" identity=pod-abc12 lease=<release>
# level=INFO msg="Became leader: pod-abc12" identity=pod-abc12
Prometheus metrics¶
Monitor leader election via metrics endpoint:
In another terminal:
For a Prometheus setup that adds the Kubernetes namespace label, query the
default installation:
# Current leader (should be 1 across all replicas)
sum(haptic_leader_election_is_leader{namespace="haptic"})
# Identify which pod is leader
haptic_leader_election_is_leader{namespace="haptic"} == 1
# Leadership transition rate (should be low)
rate(haptic_leader_election_transitions_total{namespace="haptic"}[1h])
Troubleshooting¶
Start with fleet diagnostics to identify the controller or HAProxy pod that needs attention:
| Symptom | Check next |
|---|---|
| No leader and no new configurations deploying | Read controller logs for Lease permission errors, API connectivity failures, or missing pod identity. The chart supplies permissions and identity by default. |
| More than one reported leader | Restrict the metric query to one release. Its controller replicas must use the same Lease name and namespace. Separate releases each have a leader. |
| Frequent leadership changes | Check CPU throttling, memory pressure, node health, and API latency. Adjust lease timings only after identifying why renewal fails. |
| One leader, but HAProxy isn't updating | Check validation and deployment findings in haptic doctor; inspect agent connectivity and minDeploymentInterval if structural changes are waiting. |
Read recent logs from all controller replicas in the default release:
kubectl logs --namespace haptic \
--selector app.kubernetes.io/instance=haptic,app.kubernetes.io/component=controller \
--container controller --tail=100
For resource pressure, use the sizing guide. For rejected configuration or failed deployment, follow troubleshooting.
Best practices¶
Replica count¶
Keep at least two controller replicas when you need failover. A single replica is sufficient when you can tolerate a controller outage; HAProxy continues serving its last configuration during that outage.
The chart creates a PodDisruptionBudget with minAvailable: 1 when
controller.replicaCount > 1. It limits voluntary disruption, such as a node
drain; it doesn't prevent an unexpected node failure. Configure it with:
Resource allocation¶
Every replica renders changes and holds resource and render caches. The leader also deploys configuration and writes status. Give all replicas enough CPU and memory to handle peak load after election. Use the resource sizing guide for starting requests and limits. Apply the same budget to every replica.
Anti-affinity¶
Prefer separate nodes for the controller replicas and for the HAProxy replicas.
This example targets the haptic release:
controller:
podSpec:
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app.kubernetes.io/name: haptic
app.kubernetes.io/instance: haptic
app.kubernetes.io/component: controller
topologyKey: kubernetes.io/hostname
haproxy:
podSpec:
affinity:
podAntiAffinity:
preferredDuringSchedulingIgnoredDuringExecution:
- weight: 100
podAffinityTerm:
labelSelector:
matchLabels:
app.kubernetes.io/name: haptic
app.kubernetes.io/instance: haptic
app.kubernetes.io/component: loadbalancer
topologyKey: kubernetes.io/hostname
This is a scheduling preference, so it doesn't guarantee separate nodes when
capacity is limited. After rollout, check the NODE column and confirm that
each pair spans more than one node:
Monitoring and alerts¶
The bundled PrometheusRule includes HAProxyControllerNoLeader. Enable it
through the monitoring setup.
Migration from single-replica¶
Keep your complete settings in haptic-values.yaml. The chart already grants the Lease permissions needed for leader election.
-
If you previously disabled leader election, enable it while keeping one replica, then apply the values and wait for the rollout before scaling:
-
Set
controller.replicaCount: 2in the same values file and apply it: -
Wait for both controller replicas to become ready:
-
Confirm that the Lease names an active leader:
See also¶
- Monitoring Guide - Prometheus metrics and alerting
- Debugging Guide - Runtime introspection and troubleshooting
- Security Guide - RBAC and security best practices
- Performance Guide - Resource sizing and optimization
- Troubleshooting Guide - General troubleshooting